Document resource
WHAT IS ALREADY KNOWN ON THIS TOPIC Artificial intelligence (AI) interventions in oncology are evaluated through randomised controlled trials (RCTs), but the completeness of reporting remains unclear, potentially impacting reproducibility and clinical translation.CONSORT (Consolidated Standards of Reporting Trials) 2010 and CONSORT-AI 2020 are consensus statements that aim to enhance reporting standardisation and transparency, yet adherence to these guidelines has not been systematically assessed.WHAT THIS STUDY ADDS This study identifies moderate overall adherence to CONSORT 2010 and CONSORT-AI 2020 guidelines in AI oncology RCTs, with better reporting of non-AI-specific items compared with AI-specific ones following the release of CONSORT-AI 2020.Specific gaps in reporting of methodology, harms and reproducibility were noted, highlighting areas for improvement.HOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE OR POLICY Promoting RCT adherence to CONSORT reporting guidelines through journal mandates, clearer guideline definitions and editorial oversight can improve reporting practices.Addressing critical gaps ensures AI interventions are rigorously evaluated and transparently reported, fostering trust and enabling responsible clinical adoption.Introduction The intersection of clinical oncology and digital medicine, an emergent field that leverages sensors, software and algorithms to measure and intervene in service of human health, has exploded in popularity due to the growth of big data in cancer science and promising potential to support clinical decision-making. 1 2 By learning patterns from large-scale, multimodal cancer datasets,3 artificial intelligence (AI)-enabled tools have demonstrated the ability to apply these learnt patterns to automate simple, administrative tasks or assist with specialised clinician-led tasks across several areas in oncology.4–6 Notably, the integration of AI tools in clinical decision-making poses the promising potential to predict and contribute to improved patient outcomes with clinician oversight.7 Despite the enthusiasm for the integration of AI in oncology, there remains a pressing need for rigorous evaluation of the feasibility, acceptability, safety and efficacy of these tools before routine clinical adoption.Robust evidence on the clinical impact of AI tools in high-quality randomised clinical trials (RCTs) is required to prospectively validate AI applications in clinical settings since RCTs are regarded as the gold standard for generating high-quality evidence in medicine.8 Due to its importance in evidence-based medicine, the integrity of an RCT’s findings depends heavily on the quality of its design, conduct and reporting. Poorly conducted or inadequately reported RCTs can compromise the validity and reproducibility of their conclusions, potentially leading to harmful clinical practice.9 Moreover, Software as a Medical Device (SaMD), including many AI interventions in oncology, must undergo formal regulatory review to obtain market authorisation. AI-based tools are classified as high-risk medical devices when used in clinical decision support and must demonstrate clinical evidence of safety and efficacy and post-market surveillance, both of which require rigorous clinical evaluation in RCTs and complete, transparent reporting.To address these challenges, reporting guidelines have been developed to enhance transparency and reproducibility in clinical research. The CONSORT (Consolidated Standards of Reporting Trials) 2010 consensus statement intended to design a minimum reporting standard for reporting of RCTs.10 Recognising the unique considerations of AI-based interventions and their increasing prevalence in RCTs, the CONSORT-AI extension was introduced in 2020 to provide specific guidance on reporting trials involving AI.11 The CONSORT guideline and CONSORT-AI extension aim to ensure that key aspects of AI systems, such as their development, validation and implementation, are comprehensively described in RCT reports. Importantly, standardised reporting of both trial results and trial design factors is useful to address limitations that lead to study incompletion or failure.12 Given the potential and increasing utilisation of AI tools designed to take advantage of big data in oncology, there remains uncertainty about the rigour and completeness of reporting of RCTs evaluating these interventions in oncology. This systematic review investigates the concordance of RCTs for AI interventions in oncology using CONSORT-AI and summarises the landscape of studies to characterise the current state of this growing research field.Materials and methods We queried OVID Medline and EMBASE on 22 October 2024 using oncology (‘neoplasm’, ‘cancer’, ‘tumor’, ‘onco’), clinical trial (‘clinical trial’, ‘controlled trial’, ‘randomized trial’, ‘rct’) and artificial intelligence (‘artificial intelligence’, ‘machine learning’, ‘deep learning’, ‘neural network’) search terms in consultation with a librarian. English, primary, peer-reviewed research articles published in the year 2000 to the date of the search were included. Non-English articles, editorials, conference abstracts, secondary reviews and pre-prints were excluded. The full search strategy is detailed in online supplemental data 1. This systematic review followed the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) reporting guidelines and was prospectively registered in PROSPERO (CRD42024606232).SP110.1136/bmjonc-2025-000733.supp1Supplementary data The literature search yielded 4004 articles, and after deduplication of 1138 articles, 2866 unique articles remained for abstract screening (figure 1). Abstract screening was conducted in duplicate by a team of three reviewers (KA, RS, JFA) and resolved in discussion with a fourth reviewer (DC). Abstracts of articles were screened for inclusion if the article described a primary, randomised clinical trial involving an artificial intelligence intervention in at least one arm of the trial and was related to oncology. Articles were identified as evaluating an artificial intelligence intervention if the term ‘artificial intelligence’ was used to describe the trial intervention in the intervention arm. Articles were secondary research articles, or applied in only non-oncology settings were excluded. Abstract screening yielded 110 articles for full-text review. Full-text screening was conducted in duplicate by a team of three reviewers (KA, RS, JFA) and resolved in discussion with a fourth reviewer (DC) following the same inclusion and exclusion criteria as the abstract screening procedure. Full-text screening yielded 57 articles for inclusion in this study. Reference lists of the 57 included articles were screened to confirm that all relevant research sources were included.Figure 1PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) flow diagram of the study screening procedure. RCT, randomised controlled trial.Key study attributes from the full text of each article were extracted in duplicate (KA, RS, JFA), including the purpose of the AI intervention, the intended user population, age groups of trial participants, the country that hosted the trial, the clinical specialty, the sample size and the number of recruitment sites, as well as the qualitative summaries of the methodology and primary outcomes. Risk of Bias 2.0 (RoB 2) assessment was completed followed by CONSORT 2010 and CONSORT-AI 2020 guideline assessment by a team of three reviewers (KA, RS, JFA). CONSORT 2010 and CONSORT-AI 2020 guideline assessment was based on the trial report only. Conflicts were resolved through consensus discussion with a fourth reviewer (DC). Reporting of adherence to CONSORT reporting guidelines in the full text of each article was extracted. Journal submission guidelines for each article were assessed to extract where adherence to CONSORT reporting guidelines was recommended.Analysis of CONSORT guideline concordance was assessed by the percentage of RCTs that reported each item. Subgroup analyses based on types of CONSORT checklist items (all items, AI items, non-AI items) and characteristics of the included trials (study-reporting guideline adherence, journal-recommended guideline adherence, overall risk of bias) were conducted. Given the release of CONSORT-AI in September 2020 and the time lag between the release of the consensus statement and widespread adoption, we considered articles published in 2021 onwards to be post-CONSORT-AI and articles published in or before 2020 to be pre-CONSORT-AI for the purpose of subgroup analyses. Within each trial, CONSORT guideline concordance was assessed by the percentage of items that were reported (CONSORT n=25, CONSORT-AI n=14). The primary summary metric was median concordance, given the inclusion of a subset of outlier trials with high risk of bias and low reporting concordance that would skew mean concordance. Mann-Whitney U test with Bonferroni correction was used to evaluate for differences in the CONSORT concordance distribution between studies published before and after the release of CONSORT-AI. Fisher exact hypothesis testing was used to compare the difference in item-specific concordance between trials published before and after the release of CONSORT-AI. Subgroup comparisons of concordance based on self-reported guideline, journal-recommended guideline and risk of bias were evaluated qualitatively without hypothesis testing due to small sample sizes of subgroups. Full data extraction of study-level characteristics, item-wise reporting guideline adherence and risk of bias assessments is reported in online supplemental data 2. Statistical analysis was conducted with the Python V.3.8.9 programming language and the scipy V.1.14.1 statistical analysis software package.SP210.1136/bmjonc-2025-000733.supp2Supplementary data Results Study characteristics Our study included 57 randomised clinical trials evaluating AI interventions in oncology ( online supplemental table 1).13–69 The majority of trials (n=49, 86%) were published in the year 2021 onwards following the release of the CONSORT-AI extension in 2020. The studies primarily reported AI interventions focused on screening (n=31, 54%) followed by diagnosis (n=11, 19%) and monitoring (n=6, 11%) that were intended for use by clinicians (n=49, 88%) rather than patients (n=7, 13%). AI interventions were primarily applied for endoscopy assistance (n=33, 58%), diagnostic support (n=7, 12%) and treatment planning (n=4, 7%). Trials commonly evaluated data from adult patients with cancer (n=55, 95%), with a minority of trials evaluating data from adult and paediatric patients with cancer (n=2, 6%). Notably, the most common countries that led AI clinical trials in oncology included China (n=19, 33%), USA (n=8, 14%) and Canada (n=6, 11%) (online supplemental figure 1). The median number of participants was 502 (IQR 120–1261) and the median number of recruitment sites was 1 (IQR 1–4.5). Out of the 57 trials included in this review, 47 (82%) trials reported registration in an established trial registry.SP710.1136/bmjonc-2025-000733.supp7Supplementary data SP510.1136/bmjonc-2025-000733.supp5Supplementary data Overall CONSORT and CONSORT-AI concordance Across all 57 trials included in this review, median concordance across 37 non-AI-specific CONSORT 2010 items and 14 AI-specific CONSORT-AI 2020 items was 82% ( table 1). Median concordance was 93% for CONSORT-AI 2020 items and 81% CONSORT 2010 items. Trials published before the release of CONSORT-AI (on or before 2020, n=8) had higher overall CONSORT concordance compared with trials published after the release of CONSORT-AI (after 2020, n=49) (online supplemental figure 2 (online supplemental figure 2A)). Likewise, we also observed the trials published before CONSORT-AI had higher CONSORT 2010 concordance compared with trials published after CONSORT-AI (online supplemental figure 2C (Figure 2C)). Regardless of trial publication date, similar concordance with CONSORT-AI was observed among trials published before the release of CONSORT-AI or after the release of CONSORT-AI (online supplemental figure 2B (Figure 2B)).SP610.1136/bmjonc-2025-000733.supp6Supplementary data SP310.1136/bmjonc-2025-000733.supp3Supplementary data Table 1Overall CONSORT concordance stratified by trial publication date before the release of CONSORT-AI (on or before 2020, n=8) or after the release of CONSORT-AI (after 2020, n=49)CategoryAll trials<=2020 (%)2021–2024 (%)Mann-Whitney p valueConcordance with CONSORT82 (71–88)92 (87–92.7)82 (71–86)0.011Concordance with non-AI CONSORT81 (73–84)92 (85–92)81 (73–84)0.006Concordance with AI CONSORT93 (79–93)93 (91–95)93 (79–93)0.23Mann-Whitney test was used to compare guideline concordance between trials published before and after the release of CONSORT-AI. Median and IQR percentage values of study concordance with reporting guidelines are shown.AI, artificial intelligence; CONSORT, Consolidated Standards of Reporting Trials.Among trials published after the release of CONSORT-AI in 2020 (n=49), we observed that recent trials published in 2024 had lower median overall concordance with CONSORT, non-AI CONSORT items and AI-specific CONSORT-AI items compared with trials published in 2021 immediately after CONSORT-AI’s release. Similar trends comparing trials published in 2022 and 2023 to trials published in 2024 were observed (online supplemental table 2).SP810.1136/bmjonc-2025-000733.supp8Supplementary data CONSORT item-specific concordance Among CONSORT-specific items for all 57 trials included in this review, the items with the highest concordance were often associated with the description of the background, such as Items 1b, 2 a, 2b and 5 each with 100% concordance, as well as description of the limitations and future directions, such as Items 20, 21 and 22 with at least 96% concordance across all 57 trials ( table 2). Besides the optional items indicated if relevant (Items 3b, 6b, 7b, 11b), the items with the lowest concordance were often associated with the study methodology and results, such as who was enrolled and allocated participants (Item 10), reason for trial conclusion (Item 14b), presentation of absolute and relative effect sizes for binary outcomes (Item 17b), and reporting of important harms and unintended effects in each group (Item 19).Table 2CONSORT 2010 item-specific concordance stratified by trial publication date before the release of CONSORT-AI (on or before 2020, n=8) or after the release of CONSORT-AI (after 2020, n=49)CONSORT checklistPre-CONSORT concordance (%)Post-CONSORT concordance (%)Total concordance (%)Fisher exact p value1a. Identification as a randomised trial in the title100929311b. Structured summary of trial design, methods, results and conclusions10010010012a. Scientific background and explanation of rationale10010010012b. Specific objectives or hypotheses10010010013a. Description of trial design (such as parallel, factorial) including allocation ratio10088890.583b. Important changes to methods after trial commencement (such as eligibility criteria), with reasons75818<0.014a. Eligibility criteria for participants100969614b. Settings and locations where the data were collected10096961The interventions for each group with sufficient details to allow replication, including how and when they were actually administered10010010016a. Completely defined pre-specified primary and secondary outcome measures, including how and when they were assessed8896950.376b. Any changes to trial outcomes after the trial commenced, with reasons75818<0.017a. How sample size was determined88868617b. When applicable, explanation of any interim analyses and stopping guidelines13121218a. Method used to generate the random allocation sequence88787918b. Type of randomisation; details of any restriction (such as blocking and block size)8871740.67Mechanism used to implement the random allocation sequence (such as sequentially numbered containers), describing any steps taken to conceal the sequence until interventions were assigned8871740.67Who generated the random allocation sequence, who enrolled participants, and who assigned participants to interventions7557600.4511a. If done, who was blinded after assignment to interventions (for example, participants, care providers, those assessing outcomes) and how10071750.1811b. If relevant, description of the similarity of interventions2520210.6712a. Statistical methods used to compare groups for primary and secondary outcomes1009899112b. Methods for additional analyses, such as subgroup analyses and adjusted analyses756970113a. For each group, the numbers of participants who were randomly assigned, received intended treatment and were analysed for the primary outcome8892910.5513b. For each group, losses and exclusions after randomisation, together with reasons888686114a. Dates defining the periods of recruitment and follow-up888888114b. Why the trial ended or was stopped012110.58A table showing baseline demographic and clinical characteristics for each group10088890.58For each group, number of participants (denominator) included in each analysis and whether the analysis was by original assigned groups1009495117a. For each primary and secondary outcome, results for each group and the estimated effect size and its precision (such as 95% CI)10082840.3317b. For binary outcomes, presentation of both absolute and relative effect sizes is recommended8841470.021Results of any other analyses performed, including subgroup analyses and adjusted analyses, distinguishing prespecified from exploratory8857610.13All important harms or unintended effects in each group7553560.44Trial limitations, addressing sources of potential bias, imprecision and, if relevant, multiplicity of analyses8898960.26Generalisability (external validity, applicability) of the trial findings10098981Interpretation consistent with results, balancing benefits and harms, and considering other relevant evidence10098981Registration number and name of trial registry8890891Where the full trial protocol can be accessed, if available8884841Sources of funding and other support8888881Fisher exact test was used to compare item-specific concordance between trials published before and after the release of CONSORT-AI. Median values of item-specific concordance are shown.AI, artificial intelligence; CONSORT, Consolidated Standards of Reporting Trials.CONSORT-AI item-specific concordance Among CONSORT-AI-specific items for all 57 trials included in this review, the items with the highest concordance were often associated with the purpose of the AI intervention, such as Items 1ab-i, 1ab-ii, 2a-i, as well as description of the input and output of the AI intervention, such as Items 5iv, 5v and 5vi ( table 3). The CONSORT-AI-specific items with the lowest total concordance were often associated with the details necessary for reproducibility using the AI intervention, such as inclusion and exclusion criteria at the level of the input data (Item 4a-ii), version of the AI algorithm (Item 5i), handling of low-quality input data (Item 5iii), analyses of performance error (Item 19i), and data and code access (Item 25i). We observed that no CONSORT-AI items were differentially concordant when between trials published before the release of CONSORT-AI (on or before 2020, n=8) or after the release of CONSORT-AI (after 2020, n=49), we observed that no CONSORT-AI items were differentially concordant.Table 3CONSORT-AI 2020 item-specific concordance stratified by trial publication date before the release of CONSORT-AI (on or before 2020, n=8) or after the release of CONSORT-AI (after 2020, n=49)CONSORT checklistPre-CONSORT concordance (%)Post-CONSORT concordance (%)Total concordance (%)Fisher exact p value1ab-i. Indicate that the intervention involves artificial intelligence/machine learning in the title and/or abstract and specify the type of model100949511ab-ii. State the intended use of the AI intervention within the trial in the title and/or abstract10010010012a-i. explain the intended use of the AI intervention in the context of the clinical pathway, including its purpose and its intended users (eg, healthcare professionals, patients, public)10010010014a-i. State the inclusion and exclusion criteria at the level of participants8894930.464a-ii. State the inclusion and exclusion criteria at the level of the input data10073770.184b-i. Describe how the AI intervention was integrated into the trial setting, including any onsite or offsite requirements10010010015i. State which version of the AI algorithm was used88787915ii. Describe how the input data were acquired and selected for the AI intervention100969615iii. Describe how poor quality or unavailable input data were assessed and handled10073770.1775iv. Specify whether there was human-AI interaction in the handling of the input data, and what level of expertise was required of users100909115v. Specify the output of the AI intervention100949415vi. Explain how the AI intervention’s outputs contributed to decision-making or other elements of clinical practice1009494119i. Describe results of any analysis of performance errors and how errors were identified, where applicable. If no such analysis was planned or done, explain why not8851560.06725i. State whether and how the AI intervention and/or its code can be accessed, including any restrictions to access or re-use2545420.45Fisher exact test was used to compare item-specific concordance between trials published before and after the release of CONSORT-AI. Median values of item-specific concordance are shown.AI, artificial intelligence; CONSORT, Consolidated Standards of Reporting Trials.Concordance based on self-reported guideline For self-reported adherence to reporting guidelines, out of the 49 trials published post-CONSORT-AI (after 2020), 13 (27%) trials reported adherence to CONSORT-AI, 12 (25%) trials reported adherence to CONSORT, 2 (4%) trials reported adherence to other CONSORT guidelines (eg, EQUATOR network guidelines) and 22 (45%) trials did not explicitly report adherence to any RCT reporting guideline ( online supplemental table 3).SP910.1136/bmjonc-2025-000733.supp9Supplementary data Regardless of reporting guideline adherence, median concordance to all CONSORT and CONSORT-AI items across all trials, trials adherent to CONSORT-AI, trials adherent to CONSORT, trials adherent to other CONSORT guidelines and trials that did not report adherence to reporting guidelines was comparable across all trials (online supplemental table 3). As expected, trials that reported adherence to CONSORT-AI reporting guidelines had a higher median concordance to CONSORT-AI than trials that reported adherence to CONSORT, other EQUATOR guidelines or did not report adherence to any reporting guidelines.Concordance based on journal-recommended guideline For publishing journal-mandated adherence to reporting guidelines, out of the 49 trials published post-CONSORT-AI (after 2020), 8 (16%) trials were published in journals that mandated adherence to CONSORT-AI, 38 (78%) trials were published in journals that mandated adherence to CONSORT and 3 (6%) trials were published in journals that did not explicitly mandate adherence to any RCT reporting guideline. Median concordance with both CONSORT and CONSORT-AI items was highest among trials published in journals that recommended adherence to CONSORT-AI where applicable ( online supplemental table 4). Trials published in journals that recommended reporting according to CONSORT-AI guidelines had similar median concordance with CONSORT-AI items compared with trials published in journals that recommended reporting according to CONSORT guidelines and trials published in journals with no reporting guideline recommendation. Notably, trials published in journals that recommended CONSORT-AI adherence had the highest median concordance with CONSORT items compared with trials published in journals that recommended CONSORT adherence and trials published in journals with no reporting guideline recommendation.SP1010.1136/bmjonc-2025-000733.supp10Supplementary data Risk of bias and association with CONSORT and CONSORT-AI concordance Out of the 57 trials included in this study, the majority of trials exhibited low risk of bias (n=45, 79%) or moderate risk of bias (n=9, 16%), with a minority at serious risk of bias (n=3, 5%) (online supplemental figure 2). The most common domain with at least some concerns was bias arising from the randomisation process (n=16, 28%). Six articles scored high risk of bias in any domain and three articles were scored high overall risk of bias. Risk of bias assessment exhibited moderate inter-rater concordance based on the kappa score of 0.43.When we stratified RCTs by their overall risk of bias based on the RoB 2 tool, we observed that the overall risk of bias was associated with overall CONSORT reporting guideline concordance (figure 2 (online supplemental figure 3A)). As expected, trials at serious overall risk of bias were less concordant to CONSORT compared with trials at moderate or low overall risk of bias. Similar trends were observed when we conducted a subgroup analysis based on AI-specific CONSORT-AI items (figure 3B) and non-AI-specific CONSORT items (figure 2C (Figure 3C)), where greater concerns for overall risk of bias were associated with decreased reporting guideline adherence (online supplemental table 5). Comparing concordance with AI-specific CONSORT-AI items and non-AI-specific CONSORT items, we observed that trials within each overall risk of bias category had higher median concordance with AI-specific CONSORT-AI items than non-AI-specific CONSORT items.SP410.1136/bmjonc-2025-000733.supp4Supplementary data SP1110.1136/bmjonc-2025-000733.supp11Supplementary data Figure 2Association of study concordance with (A) all CONSORT (Consolidated Standards of Reporting Trials) items, (B) artificial intelligence (AI)-specific CONSORT-AI 2020 items and (C) non-AI-specific CONSORT 2010 items with low (n=45), medium (n=9) and serious (n=3) overall risk of bias (*p<0.05, *p<0.01, ***p<0.001).Discussion Key findings of RCT reporting adherence to CONSORT and CONSORT-AI This systematic review assessed 57 RCTs investigating AI interventions in oncology, offering an overview of the current state of reporting practices based on AI-specific CONSORT-AI guidelines as well as a comprehensive summary of the landscape of AI RCTs in oncology. Our findings reveal both promising trends and significant gaps in trial reporting practices, with implications for the development, evaluation and adoption of AI technologies in oncology.The review identified an overall low rate of self-reported adherence to CONSORT-AI (27%) and CONSORT (25%) guidelines among the included 49 trials published after CONSORT-AI (after 2020), indicating that despite consensus recommendations to promote standardised RCT reporting and improvements to the quality of RCT reporting,70 there remains a need to establish widespread familiarity and adoption of AI-specific RCT reporting standards. Notably, the overall median concordance of all RCTs with CONSORT and CONSORT-AI items was moderate at 82%, with similar rates of concordance regardless of self-reported use of CONSORT-AI, CONSORT and no guidelines. This may suggest that definitions of CONSORT items require further assessment to ensure that authors clearly understand the scope and depth of information required to report according to the expectations of the guideline. In comparison, we observed higher overall concordance in trials published in journals that mandated adherence to CONSORT-AI (88%) compared with CONSORT (82%) and no guideline mandate (80%). Our findings suggest that journals play a pivotal role in promoting high-quality reporting by enforcing adherence to reporting guidelines such as CONSORT.71 72 Compared to a 2022 review of AI RCT concordance with CONSORT reporting guidelines, we similarly observed that journal-mandated adherence to CONSORT rather than self-reported adherence to CONSORT was associated with higher CONSORT concordance, suggesting that editorial and funding mandates are potential driving factors that contribute to complete trial reporting.73 Our findings are consistent with those of Kwong et al, who applied APPRAISE-AI to evaluate studies of AI outcome prediction tools in non-muscle invasive bladder cancer contexts, reporting a median overall score of 37% and poor reporting of methodology necessary for reproducibility.74 Similarly, both Dhiman et al and Collins et al also demonstrated pervasive methodological shortcomings in reporting of machine learning-based prognostic models in oncology, such as inadequate sample size justification, study registration, protocol availability and data/code sharing.75 76 The observed decline in overall median concordance with CONSORT reporting guidelines among trials published in 2024 (71%) compared with 2021 (84%) is concerning, particularly given the lower median concordance to both non-AI CONSORT items. This trend suggests a potential lapse in adherence to CONSORT reporting guidelines as the initial awareness and resulting momentum following the release of CONSORT-AI in 2020 diminishes. Our finding of decreasing compliance mirrors patterns seen after publication of the original CONSORT statement, where Kane et al observed an early surge in adherence that plateaued post-publication, that may in part be due to sustained journal enforcement.70 Taken together, these findings underscore the critical need for ongoing efforts to promote and reinforce the adoption of reporting guidelines, ensuring that trials maintain rigorous transparency and reproducibility standards, especially in rapidly evolving fields such as AI-driven research.Under-reported features One of the most striking findings of this review is that among all 57 trials included in this review, we observed higher median concordance with AI-specific CONSORT-AI items (93%) compared with non-AI-specific CONSORT items (81%). This discrepancy may reflect the novelty and specificity of AI-focused reporting requirements, which align more closely with the interests and expertise of AI researchers developing these technologies. Among the 49 trials published post-CONSORT-AI, we observed high concordance with items describing the purpose and functionality of AI interventions, such as input (Item 5ii, 96%) and output (Item 5v, 100%) characteristics and intended use (Item 1ab-ii, 100%), suggesting that researchers are prioritising reporting these foundational aspects necessary for verifiability and reproducibility of AI applications in oncology. These items are important for contextualising the role of AI intervention in clinical workflows and ensuring that stakeholders, including clinicians and patients, understand the tool’s capabilities and limitations. 77 In contrast, non-AI-specific CONSORT items related to trial methodology and results, such as participant enrolment (Item 10, 57%) and reporting of harms (Item 19, 53%), exhibited relatively lower concordance. These deficits indicate persistent challenges in aligning AI-focused trials with broader standards of clinical research rigour.78 The high concordance with CONSORT items related to descriptions of the study background (Items 1b, 2a, 2b), data collection (Items 4a, 4b), outcomes (Item 6a) and limitations (Items 20, 21, 22) reflects the importance of context about data sources and limitations of AI interventions. Given that the context-dependent performance of AI tools is directly associated with the characteristics of its training and testing data, the prevalent reporting of data sources is a meaningful step towards transparency of AI algorithm performance evaluation. Similarly, the high concordance with AI-specific items related to the purpose, input and output of AI interventions (Items 1ab-i, 1ab-ii, 2a-i, 5ii, 5iv, 5v, 5vi) promotes clear descriptions of how the intervention would process and output data. By transparently reporting the rationale and outputs of AI interventions, researchers can help clinicians and other stakeholders make value-informed decisions about their implementation in clinical workflows and their potential impact on patient outcomes.79 Several reporting items related to reproducibility and data handling exhibited poor concordance with CONSORT and CONSORT-AI guidelines. Among CONSORT items, defining reasons for trial conclusion and reporting of harms and unintended effects were reported less frequently, pointing towards the need to include negative findings and reasons for study failure in study reporting.80 Despite the surge of positive findings in AI research, reporting negative results can reduce redundant scientific efforts by providing a complete understanding of an intervention’s efficacy and limitations. Coupled with the poor reporting of absolute and relative effect sizes among included RCTs, there remain concerns about limitations towards assessing the clinical safety and efficacy of AI tools which may pose risks to patient care if deployed without careful oversight.Stakeholder roles and future directions Among CONSORT-AI items, poor concordance with inclusion and exclusion of input data and handling of low-quality input data may lead to positive results that fail to generalise to broad, heterogeneous populations of patients with cancer. Clarification of exclusion criteria of outliers and handling of sparse, incomplete data in real-world datasets can help optimise AI applications to be more robust, generalisable and capable of delivering consistent performance across diverse clinical settings. 81 Moreover, poor reporting of algorithm versions, as well as data and code access, makes it difficult for external teams to independently validate AI tools outside of their training context.82 To promote consistent reporting adherence across all CONSORT and CONSORT-AI guidelines, journal editors can encourage authors to submit the completed checklist of the relevant reporting guideline on study submission. Editorial processes should encourage adherence to study design-specific reporting guidelines, such as CONSORT-AI for clinical trials evaluating AI interventions,11 TRIPOD-AI for prediction model studies,83 PROBAST-AI for risk of bias assessment of prediction model studies84 85 and TRIPOD-LLM for large language model studies.86 We hypothesised that trials with higher overall risk of bias would also be poorly concordant with reporting guidelines if perceived risk of bias is based in part on reporting completeness. As expected, trials with higher overall risk of bias demonstrated lower concordance with both CONSORT and CONSORT-AI items, suggesting that perceived methodological rigour is closely linked to reporting quality, alike to previous observations in RCTs evaluating AI for medical imaging.87 Notably, bias arising from the randomisation process remained a common concern that affected 28% of AI trials in oncology. These results underscore the importance of addressing methodological weaknesses in the trial design and conduct. By improving randomisation practices among other key aspects of AI trial conduct, researchers can reduce the risk of bias and enhance the validity and clinical applicability of their findings. This is especially essential in AI-driven research, where methodological rigour directly impacts the ability to discern validated clinical benefits from spurious associations or algorithmic biases that should be controlled as confounding variables.Possible ways to promote standardised reporting practices beyond CONSORT guideline adherence can include the design of RCT protocols that are concordant with RCT protocol-specific reporting guidelines such as SPIRIT-AI,88 open-access sharing of trial protocols through pre-print archives and protocol-specific publication mediums for peer review89 and registration of trials through RCT registries such as ClinicalTrials.gov to ensure relevant information used for trial registration is available.90 Consideration of trial design to address risks of bias coupled with standardised trial reporting collectively aims to improve translation of AI technologies towards responsible deployment in healthcare and meaningful advances in patient outcomes.In this review, we did not extract clinical endpoints and therefore cannot determine whether higher CONSORT-AI concordance or lower risk of bias was associated with observed clinical outcomes, highlighting a future direction of research. A second limitation of this review is that we did not distinguish which CONSORT items were optional or not applicable for particular trials, potentially biasing our concordance estimates by treating all checklist items as uniformly required.Conclusion This systematic review showed that AI trials in oncology report the majority of CONSORT and CONSORT-AI items well, with some critical under-reported items related to methodological transparency necessary for reproducibility. Addressing inconsistent reporting deficiencies and promoting broader adoption of RCT reporting guidelines can ensure that AI interventions in oncology are evaluated based on robust evidence from well-controlled RCTs that form the basis of evidence-based medicine. Fostering consistent, transparent and comprehensive reporting in AI-driven oncology trials is essential to building trust in these promising technologies and enabling their safe integration into clinical practice with reliable impacts on patient outcomes.