Document resource
Background/Objectives Artificial intelligence (AI), including large language models (LLM), has been proposed as a potential strategy to augment case-based learning (CBL). However, LLM use for educational purposes is mired in concerns regarding ‘hallucinations’. This concern about generating faulty information is frequently discussed in the literature but has been relatively understudied in the CBL context.Methods This study employed a cross-sectional analysis of the ability of an LLM to respond to medical student questions, as may occur in a CBL scenario. Five descriptions of patient cases were prepared, each with a different neurological presenting complaint. OpenAI’s GPT-4o-mini was used to emulate the patient in each interaction through a custom online interface. Three medical student investigators interrogated the LLM cases through free-text questioning regarding history, examination, and investigation results to arrive at a diagnosis. All student-investigator questions and AI-generated responses were evaluated by medical officers.Results There was a total of 857 question-response pairs generated following the interrogation of the five LLM cases. In response to student-generated questions, the LLM adhered to the provided case in 832/857 (97.1%) of responses. There were 25/857 (2.9%) responses in which the LLM provided information beyond that which was in the provided case, one of which was considered inconsistent with the case. Therefore, overall, 856/857 (99.9%) of LLM responses were appropriate.Conclusion LLM-generated CBL scenarios have significant potential to provide high-quality, interactive medical education. Further research and testing could facilitate the process of integrating this technology into everyday use.