BetaEntity Annotation Prototype
← Back to treatments

Annotated abstract

70 Robust, ultra-efficient extraction of neurological concepts from free-flowing clinical notes

jnnp · 2025-11-26 · canonical JSON source

8 visible annotations · policy: published · automated confidence ≥ 75.00%

Document resource

Introduction Extracting large-scale structured tables manually from unstructured free-flowing text for clinical, intelligence and research analytic insight is impractical. Natural language processing (NLP) with emerging machine learning tools can facilitate structured data extraction.Methods Across three regional NHS England Secure Data Environments (SDEs) (Lancashire & South Cumbria, London and Wessex) we are developing and training NLP models (MedCAT, Medical Concept Annotation Tool) to achieve automated extraction of multiple sclerosis (MS) diagnostic subtypes and disease modifying treatment (DMT) information using SNOMED CT standards. We will then map to the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM), a standardised data framework now adopted by the SDEs.Results We have manually annotated 391 consultation notes for MS patients at King’s College Hospital and the National Hospital for Neurology and Neurosurgery. We will increase the number of annotations for model training, deploy and validate these models against manually extracted data at University Hospital Southampton and the Royal Preston Hospital and then expand the scope of the work, including OMOP mapping.Conclusion Our approach will pave the way for ultra-efficient, privacy-conscious analyses of real-world data at scale in a standardised manner across SDEs.hedley.emsley@lancaster.ac.uk