BetaEntity Annotation Prototype
← Back to institutions

Annotated abstract

Predicting risk of early-onset sepsis in low-resource neonatal units using routine healthcare data: development and evaluation of multivariable statistical and machine learning models

bmjpo · 2025-09-29 · canonical JSON source

3 visible annotations · policy: published · automated confidence ≥ 75.00%

Document resource

Background Neonatal sepsis is a major cause of morbidity and mortality in low-resource settings and accurate, context-appropriate diagnostic methods are urgently needed to improve clinical outcomes.Methods We used data collected using Neotree, an open source digital health intervention tool, from neonates admitted to Sally Mugabe Central Hospital in Harare between February 2021 and September 2024 to model a composite outcome variable comprised of senior clinician-assigned diagnosis at discharge or cause of death and blood culture test results. Three statistical and machine learning algorithms were developed, tuned where appropriate using cross-validation and evaluated.Results In total, 917 cases of early-onset neonatal sepsis were identified among the 18 345 neonates in our study sample, comprising 664 cases of clinician diagnosis and 253 positive blood culture results. With area under the receiver operating characteristic curve as a metric, LightGBM, a machine learning gradient-boosted tree classifier, performed marginally better (0.712; 95% CI 0.673 to 0.75) than logistic regression (0.687; 95% CI 0.646 to 0.728) on a held-out evaluation dataset. A simple and easily interpretable machine learning model, the k-neighbours classifier, offered comparable performance (0.699; 95% CI 0.662 to 0.736).Conclusions This study explored the potential advantages of using machine learning in the triage of neonates at risk of sepsis in low-resource settings where gold-standard blood culture test results are often unavailable. While the differences in performance metrics were not statistically significant, the machine learning approaches in our study offer other advantages including more intuitive predictions and the ability to handle missing data without imputation.