Document resource
Cerebral aneurysm treatment methods vary depending on aneurysm morphology, patient factors, and surgeon preference. Large language models (LLMs) offer a potential platform for clinical decision-support, including pre-operative planning, but need fine-tuning to perform context-specific tasks within the neuroendovascular space. This study evaluates the baseline alignment of frontier LLMs with expert neurointerventionalists, determines model dependence on text versus images, and informs a framework to optimize foundation models for neurointervention.We randomly selected 100 aneurysms from the Cleveland Clinic Brain Aneurysm Archive (2015-2025), including primary and recurrent cases; 17 excluded due to inadequate imaging. Three neurointerventional surgeons independently reviewed cases in randomized order with washout periods across three conditions: (1) text only (clinical history and radiographic findings), (2) images only (angiographic working views), and (3) combined. Reviewers selected one treatment from eight standardized endovascular and surgical options (figure 1). The actual technique was treated as an additional independent opinion. The same cases were presented to Claude Opus 4.6 and Sonnet 4.6 (zero temperature, independent sessions per case per condition). Inter-rater agreement was assessed using Gwet’s AC1 and pairwise concordance in R.At least two raters agreed on the preferred technique in 95.2% of text-only and 90.4% of image-only cases (4-rater AC1 = 0.390 text; 0.263 images). On text, the LLMs matched at least two of the four raters in 54.2% (Opus) and 53.0% (Sonnet) of cases; on images this fell to 22.9% (Opus) and 25.3% (Sonnet).Sonnet matched the operative report at rates comparable to individual surgeons (48.2% vs. 43.4/44.6%/45.8%), though with evident category-level biases, while Opus matched 38.6%. Both models demonstrated a ‘vision gap,’ falling to 15.7% concordance on images. Both LLMs fell outside all three surgeon opinions in 24.1% (Opus) and 27.7% (Sonnet) of text-only cases, rising to 55.4% for both models on images, with failure modes including coil-alone default bias, MCA bifurcation clipping bias, under-recommendation of balloon-assisted coil embolization and flow disruption, and inability to reason about retreatment contexts (figure 1).Cerebral aneurysm treatment selection exhibits substantial expert variability. Frontier LLMs approach physician-level concordance on clinical text but demonstrate failure to leverage angiographic images for treatment planning. LLM under-recommendation of balloon-assisted coiling and flow disruption may reflect temporal shifts in training data relative to evolving practice patterns. These findings indicate that fine-tuning must prioritize image-based training, and that class imbalance in low-frequency techniques may require enrichment strategies. Future work will evaluate combined text-and-image inputs and literature-augmented prompting.Disclosures A. Mahmood: None. M. Nagy: None. J. Luo: None.Abstract P-027 Figure 1(A) Treatment selection frequency by rater (text-only). (B) Treatment selection frequency by rater (image-only), illustrating balloon-assisted coil embolization omission, coil-alone default, and microsurgical clipping bias in LLM recommendations. Unanimous agreement clustered around well-established paradigms, while complete disagreement occurred predominantly in bifurcation aneurysms and cases with branch vessel incorporation at the neck. (C/D) Expert consensus distribution across text-only (C) and image-only (D) conditions. (E/F) Number of experts matched by Opus across text-only (E) and image-only (F) conditions. Both expert consensus and LLM alignment degrade in image-only conditions. Gwet’s AC1: 3-rater text = 0.405, images = 0.281; 4-rater text = 0.390, images = 0.263