BetaEntity Annotation Prototype
← Back to interventions

Annotated abstract

3552 Large language models provide accurate feedback for neurology case-based learning

bmjno · 2025-10-23 · canonical JSON source

2 visible annotations · policy: published · automated confidence ≥ 75.00%

Document resource

Background/Objectives Artificial intelligence, including large language models (LLMs), has significant potential to provide clinical learners with effective feedback. There is currently limited evidence in this area, which is needed due to risks associated with ‘hallucinations’. The aim of this study was to examine how LLM feedback on a case-based learning exercise compared to that of human experts.Methods Five distinct case-based learning interactions were undertaken by four student investigators to generate 20 cases for feedback. Feedback was provided by an LLM and three experts (one American neurologist, one Australian neurologist, and one medical education expert). These items of feedback were then evaluated using metrics including previously published feedback evaluation tools – the QuAL and EFeCT scores. The QuAL and EFeCT scores were completed by two further neurologists and the student investigators who undertook the cases.Results LLMs and expert feedback were similar in terms of word length and number of sentences. In the feedback provided, the AI commented on 20/20 (100%) aspects of the key learning points, compared to 39/60 (65%) for the human experts. Components of the history, examination, and investigations were referred to with similar frequency. Both expert and student evaluation of the feedback demonstrated higher scores for the LLM than for the human experts on both the QuAL and EFeCT scores (P <0.001). There were no medical inaccuracies in the LLM feedback.Conclusion In this study, LLM feedback on case-based interactions was at least similar in terms of content to a panel of experts, and possibly superior.