BetaEntity Annotation Prototype
← Back to diseases

Annotated abstract

8371 AI is not coming for your job (yet): especially, if that includes writing useful patient information leaflets

archdischild · 2025-10-06 · canonical JSON source

18 visible annotations · policy: published · automated confidence ≥ 75.00%

Document resource

Objective We have been promised that AI will make our lives easier. We assessed three open AI platforms’ generation of parental information leaflets for paediatric conditions.Method We performed a structured qualitative analysis using the ‘Suitability Assessment of Materials (SAM) framework’ for health-related educational resources. SAM consists of 6 criteria: ‘content’(information to help solve problems); ‘literacy demand’(common words are used); ‘graphics’(simple appropriate drawings); ‘layout and typography’(type size 12 point etc); ‘learning stimulation and motivation’(complex topics are subdivided into small parts); and ‘cultural appropriateness’(examples present the culture in positive ways’).We randomly sourced NHS information leaflets for parents in 13 paediatric conditions (gastroenteritis; jaundice; asthma; eczema; scarlet fever; UTI; meningitis; cleft lip and palate; fever; abdominal pain; headache; sepsis; and chickenpox) from hospitals in England regions, Scotland and Wales; no leaflets from Northern Ireland were found. Using a random number generator, 2 leaflets from England (North), 2 from England (South), 1 from Wales, and 1 from Scotland were sourced for each condition, finding 73 NHS leaflets. Using the prompt: ‘Create an information leaflet for a parent of a child with [insert condition]’,’ChatGPT™’,’Google Gemini™’, and ‘Microsoft Copilot AI™’ were each asked to produce a leaflet for each condition making 39 AI leaflets.The SAM scores for NHS’s and the 3 AI’s leaflets were compared.Results We found that only 1 in 5 Trusts/Health Boards produced a leaflet available for each condition.Overall NHS leaflets scored better with a mean (95% confidence intervals of the mean) 77.7%(75.6–79.9). The overall score for AI leaflets was 67.4%(63.8–70.8)(p<0.0001).The NHS leaflets scored statistically significantly higher than AI for:’Content’ 88.4%(85.8%-91.0%) versus 75.0%(72.3–77.7)(p<0.05);’Graphics’ 65.7%(60.7–69.7) versus 46.7%(43.0–50.4)(p<0.05);’Cultural appropriateness’ 57.0%(51.3–63.7) versus 37.0%(31.0–43.0)(p<0.05).The AI leaflets scored statistically significantly higher for ‘Literacy demand’86.1%(83.5–88.7) versus NHS 78.0%(74.5–81.5)(p<0.05).There were no differences for ‘Layout and typography’ and ‘Learning stimulation and motivation’.ChatGPT™ performed statistically significantly better overall than Copilot™ and Gemini™: 77.1%(71.6–82.6) versus 62.1%(59.5–64.6)(p=0.0004) versus 62.9%(57.2–68.6)(p=0.003). There was no difference in scores between Copilot™ and Gemini™There was no difference in scores between ChatGPT™ overall and NHS leaflets, but the ‘Content’ score for ChatGPT™ was lower 71.5%(64.1%-79.1%)(p=0.0001).Conclusion Open AI cannot yet write parent information leaflets. Even if you used ChatGPT™ significant changes are needed for the required NHS standard.AI leaflets lacked detail and graphics; most missed at least one crucial informational element. None of the AIs generated or incorporated pictures independently.AI scored higher for literacy demand, due to consistent use of short sentences, bullet points, and basic vocabulary.More detailed AI prompts might result in creation of better leaflets. Scoring could not be blinded, but the SAM framework should have mitigated that.As only 1 in 5 Trusts/Health Boards produced a leaflet, it might be in the best interest of patient safety and improving educational health services, to adapt another centre’s leaflets rather than rely on AI. Future studies need to include user assessment.