Document resource
Background To assess the diagnostic performance of large language models (LLMs) for contrast-enhanced ultrasound (CEUS) interpretation of untreated focal liver lesions and to examine the impact of adding clinical data, disclosing candidate diagnoses, and using English vs. Chinese prompts.Methods We retrospectively analyzed 706 untreated focal liver lesions. For each lesion, 31 structured ultrasound/CEUS features (e.g., echogenicity, margin, enhancement pattern, and washout time) and relevant clinical variables (e.g., liver enzymes, tumor markers, hepatitis B markers) were provided to multiple LLMs (including DeepSeek and GPT-series) under four conditions: imaging-only without candidate disclosure; imaging-only with candidate disclosure; imaging+clinical without candidate disclosure; and imaging+clinical with candidate disclosure. Each setting was tested in English and Chinese and repeated three times; mean rates were reported. Outcomes included benign/malignant classification and specific diagnosis (both with ‘indeterminate’ allowed), with primary emphasis on disease-level diagnostic accuracy.Results Overall, LLMs performed better for benign/malignant classification than for specific diagnosis. Using English prompts, mean classification accuracy ranged from 0.639–0.906 and disease-level accuracy from 0.441–0.695; using Chinese prompts, classification accuracy ranged from 0.629–0.870 and disease-level accuracy from 0.468–0.671. Notably, our best disease-level diagnostic performance to date was achieved with GPT-5.2 Thinking, using an English prompt that included clinical context and a broad candidate diagnostic range, which aligns with our primary endpoint focused on specific diagnosis accuracy. Indeterminate outputs varied across models and input strategies (classification: 0.035–0.280; disease-level: 0.017–0.372), indicating that clinical data and candidate disclosure can meaningfully change model behavior.Conclusions LLMs show practical utility for CEUS-based assessment of liver lesions—especially for malignancy screening and structured decision support—although performance depends on prompting strategy and language.