BetaEntity Annotation Prototype
← Back to institutions

Annotated abstract

O-021 A hybrid large language model-deterministic rules engine for blinded CPT code assignment in neurointerventional surgery

neurintsurg · 2026-07-19 · canonical JSON source

1 visible annotations · policy: published · automated confidence ≥ 75.00%

Document resource

Introduction/Purpose Accurate CPT coding is critical for compliant billing and revenue integrity. In neurointervention, complexity arises from multiple coordinated codes spanning diagnostic catheterization, intervention, and radiologic supervision and interpretation within a single case, all governed by NCCI edits and MUE limits. We evaluated a two-stage pipeline in which a large language model (LLM) extracts structured information from operative-note text, followed by a deterministic rules engine that applies fixed billing logic, and tested whether the rules layer could be refined independently of the model.Materials and Methods A GPT-4.1 model extracted structured data from free-text operative notes; deterministic Python rules then applied catheterization hierarchy, intervention, and supervision-and-interpretation logic, NCCI bundling, and MUE limits across 56 codes, compared against CPT code assignments from a certified professional coder. The pipeline was blinded and tested on consecutive cases from 2024-2025 across three cranial neurointerventional cohorts from two attendings: cohort 1 (88 cases, attending A), cohort 2 (21 cases, attending B), and cohort 3 (120 cases, attending B, after targeted rules-layer refinement). Five systematic rules-layer errors identified in cohorts 1-2 were corrected between cohorts 2 and 3, without modifying model prompts, schemas, or weights. Primary outcomes were micro-averaged precision, recall, and F1.Results In cohort 1, precision was 0.92, recall 0.87, and F1 0.90, with exact code-set concordance in 57%. Under a new attending's template and case mix (cohort 2), F1 declined to 0.85, consistent with a generalization penalty rather. After targeted refinement, cohort 3 F1 rose to 0.91 (precision 0.89, recall 0.93), with exact concordance in 62%; per-code F1 improved for every targeted code (e.g., 36225, 0.50 to 0.88). Across all 229 cases, F1 was 0.90 ( Figure 1). Each case processed in seconds at negligible cost.Conclusion A hybrid LLM-deterministic pipeline achieved high agreement with a certified professional coder (F1 0.90-0.91), with exact code-set concordance in over 60% of validation cases. A modest generalization penalty on new documentation was recoverable through refinement of the deterministic rules layer alone, without model recalibration or re-prompting. Paired with specialist oversight, this auditable architecture may be a feasible component of clinical coding workflows. Further validation across a broader range of operators, dictation styles, and case mixes is warranted.Disclosures M. Longo: None. N. Mummareddy: None. G. Koutsouras: None. A. Bhamidipati: None. K. Raygor: None.Abstract O-021 Figure 1