BetaEntity Annotation Prototype
← Back to interventions

Annotated full text

Five years after CONSORT-AI, not much has changed: a call to action for artificial intelligence research in oncology

bmjonc · 2025-08-24 · canonical JSON source

9 visible annotations · policy: published · automated confidence ≥ 75.00%

Document resource

Since its introduction in 1996, the Consolidated Standards of Reporting Trials (CONSORT) statement has been widely used to facilitate transparent reporting of randomised controlled trials (RCTs), including those in oncology.1 Due to the unique methodological and implementation challenges associated with artificial intelligence (AI)-based interventions, the CONSORT group subsequently released the CONSORT-AI extension in 2020.2 This extension introduced 14 AI-specific items to strengthen the clarity, reproducibility and reliability of trials involving AI tools. Five years following the introduction of CONSORT-AI, has reporting quality improved for AI oncology trials? The systematic review by Chen et al suggests otherwise, revealing a troubling decline in overall adherence to established reporting guidelines (ie, CONSORT 2010 and CONSORT-AI).3 Adherence to CONSORT and CONSORT-AI peaked in 2022 at 87% and 96%, respectively, only to drop to 71% and 79% by 2024 as the number of published trials continued to rise. This decline is unlikely to be explained by delayed uptake of CONSORT-AI alone. More plausibly, it reflects a systemic failure: researchers are either unaware of these guidelines or are choosing not to implement them, and journals appear unwilling or unable to enforce them. Not surprisingly, trials judged to be at high risk of bias were also those least likely to comply with reporting standards. These findings suggest an association between rigorous study conduct and transparent reporting. AI studies in both model development and RCT settings that are found to be at high risk of bias are also often less compliant with reporting guidelines, raising concerns about the transparency, reproducibility and clinical applicability of their results.4 This issue is especially relevant given the inherent ‘black box’ nature of AI predictive models and the increasingly assertive outputs generated by large language models, which may mask underlying methodological weaknesses. Closing these gaps will require commitment from all levels of research. Journals and editors should mandate adherence to reporting guidelines such as CONSORT-AI as a prerequisite for publication of AI-focused RCTs.2 Protocol repositories should recommend the use of similar guidelines such as Standard Protocol Items: Recommendations for Interventional Trials - Artificial Intelligence (SPIRIT-AI) from the outset to promote transparency in trial design.5 Authors, peer reviewers and editorial boards must also be equipped and encouraged to apply the most relevant AI-specific guideline at each stage of the AI development pathway (figure 1).2 4–14 Without the collective accountability and engagement from all stakeholders, reporting quality may continue to decline, limiting the credibility and quality of AI research in oncology.Figure 1Relevant artificial intelligence (AI)-specific guidelines at each stage of the AI development pathway. Guidelines that are currently under development are not included. CHEERS-AI, Consolidated Health Economic Evaluation Reporting Standards for Interventions that use Artificial Intelligence; CONSORT-AI, Consolidated Standards of Reporting Trials-Artificial Intelligence; DEAL, Development, Evaluation, and Assessment of Large Language Models; DECIDE-AI, Developmental and Exploratory Clinical Investigations of DEcision support systems driven by Artificial Intelligence; FUTURE-AI, Fair, Universal, Traceable, Usable, Robust, Explainable-Artificial Intelligence; PROBAST+AI, Prediction model Risk Of Bias ASsessment Tool-Artificial Intelligence; SPIRIT-AI, Standard Protocol Items: Recommendations for Interventional Trials-Artificial Intelligence; STANDING, Standards for Data Diversity, Inclusivity, and Generalisability; TRIPOD+AI, Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis-Artificial Intelligence; TRIPOD-LLM, Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis-Large Language Model While Chen et al provide valuable insights into the reporting quality and risk of bias in AI RCTs within oncology, several areas remain unexplored. Over 80% of the included trials focused on AI interventions for screening, diagnosis or monitoring. While these applications are clinically relevant, their impact on patient outcomes ultimately depends on whether AI-generated alerts or predictions—when available—reach the appropriate end users, who must also have sufficient resources, training and trust in the AI system to act on them.15 In addition, this review did not examine the relationship between reporting quality, risk of bias and trial outcomes, such as whether the AI intervention yielded positive or negative results, or the magnitude of its effect—factors that may influence the interpretation and clinical applicability of study findings. Lastly, this review did not include any information on adverse events or unintended consequences attributable to AI interventions, an increasingly important consideration given the potential for harm especially when these tools are deployed in real-world clinical settings.16 However, this likely reflects a lack of reporting within the included trials, rather than a limitation of the review itself. These gaps highlight the need for future work to go beyond reporting and methodological evaluations and address the broader implications and safety of AI interventions in healthcare. In conclusion, the promise of AI in oncology depends not only on technological innovation but also on the quality of evidence supporting its use. As illustrated by Chen et al, adherence to established AI reporting guidelines is declining at a time when it is more critical than ever. Poor reporting compromises the ability of stakeholders, including clinicians, patients and policymakers, to make informed decisions about AI integration into care pathways. Tackling these challenges must begin with a call to action to strengthen adherence to current reporting standards. Future work must also examine whether AI interventions are truly improving patient outcomes and whether harms are being systematically identified and reported. Only through such a comprehensive approach can AI be responsibly and effectively translated into meaningful improvements in cancer care.