Document resource
WHAT IS ALREADY KNOWN ON THIS TOPIC Bayesian networks (BNs) are increasingly recognised as powerful artificial intelligence (AI) tools for prognostication in oncology, offering key-advantages over traditional regression models.WHAT THIS STUDY ADDS This review identifies and evaluates studies that develop and validate BNs as prognostic models in oncology. It highlights BN’s unique capacities, such as modelling causal relationships, handling missing data, adapting to new evidence and predicting multiple outcomes in one simultaneously.HOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE OR POLICY The findings demonstrate the potential of BNs to enhance shared medicine-making in oncology practice. The study also shows the importance of robust development and validation methods for these AI-based prognostic tools.Introduction In the last decades, artificial intelligence (AI) has made tremendous steps into creating efficient and accurate models to improve risk stratification and guide decision-making in healthcare. Among these, Bayesian networks (BNs) have emerged as powerful tools for probabilistic modelling and decision-making.BNs are graphical models representing variables and their probabilistic dependencies using directed acyclic graphs and provide information on a qualitative level (network structure) and a quantitative level (conditional probability distributions (CPDs)). The networks consist of nodes representing variables, and arcs representing conditional dependencies between the nodes, allowing the networks to express causal or statistical relationships. To illustrate, in a BN modelling cancer progression, depicted in figure 1, a variable presenting tumour stage may directly influence outcome variables such as survival probability, here reflecting their causal relationships. The node ‘tumour stage’ is in turn influenced by its parent nodes (tumour grade and lymph node metastasis). In this way, both direct and indirect effects among variables can be captured within an interpretable model. Each node is associated with a CPD specifying the probabilities of the node’s outcome given the combinations of parent states (figure 1).Figure 1Examples of a Bayesian network in which 5-year overall survival is modelled in relation to tumour stage, tumour grade and lymph node metastasis. The Bayesian network is depicted with its marginal probability distributions (A and B), as well as with evidence provided to the network and its conditional probability distributions (C and D). LNM, lymph node metastasis.These networks are computationally efficient as BNs explicitly represent independencies as well, here reducing the number of correlations that need to be evaluated. They rely on the Bayesian approach to statistical inference, which combines prior knowledge with observed data to update beliefs, essentially a revised probability, about uncertain events (belief updating).1 Belief updating is achieved by using Bayes’ rule that calculates the probability of an event by combining its baseline probability with the likelihood of the observed data. This belief updating process not only enhances predictive power but also enables the exploration of ‘what-if’ scenarios and counterfactual reasoning. For example, BNs can be used to estimate how altering a specific treatment decision may affect survival probabilities, offering valuable insights for personalised decision-making.There is a large variety of techniques applied to construct the BN structures. Generally, these graphical structures can be learnt with data-based inference, by using expert knowledge or information from medical literature, or a combination of both approaches. Many techniques are available to learn BN structures from data, which can be classified into three approaches:Constraint-based methods use a given dataset to identify conditional independence relationships between variables (eg, Markov conditions). These independencies are then used to constrain the underlying BN structure.2Score-based methods explore different structures and assign a score to each structure that indicates how well it fits with the given dataset. Exhaustive searches are usually not feasible because of the large number of potential structures. Therefore, most score-based algorithms have employed heuristic search techniques. One example is the hill-climbing algorithm that starts with a network structure and subsequently attempts every possible arc addition, removal and reversal compared with the initial candidate network structure.3 Only changes increasing the score become the next candidate network structure, until a change no longer increases the score.Hybrid algorithms that apply both constraint-based and score-based techniques, such as Max-Min Hill-Climbing.4Besides these machine learning (ML) approaches, expert knowledge can be used to force known direct causal relationships to be part of a ML proposed BN structure. It is known that these hybrid approaches improve structure learning processes.5BNs enable the integration of a diversity of predictive or prognostic factors such as clinical, genetic and imaging data (eg, discrete imaging data such as radiomics features) to model disease progression, treatment response and survival outcomes. In contrast to traditional statistical models such as regression models, BNs are able to capture complex interactions between variables and provide interpretable insights in causal dependencies within a disease domain. In contrast to black-box AI models such as deep neural networks, BNs are inherently explainable due to their transparent structure, which may enhance end-user trust in the model. Compared with regression models, several advantages make them very suitable for prognostic and predictive medicine in oncology, including their ability to manage missing data and their ability to be updated when new evidence is available (dynamic nature). Under the right assumptions (eg, correctly specified network structure using expert knowledge and absence of unmeasured confounding), BNs can also be used to support causal inference and estimate treatment effects from non-randomised data. These capacities may help to overcome disadvantages of using regression-based models, which include: each outcome of interest has to be trained using a different model; they cannot model causal structures; they operate under rigid assumptions about the relationships between variables, for example, linearity, independence. Compared to pure ML approaches BNs do not explicitly necessitate the use of massively large datasets as they can incorporate expert knowledge in circumstances where data are limited. Despite these advantages, clinical implementation of these models is currently lacking.6Within this systematic review, we identify the available literature on BNs for prognostication purposes in the field of oncology. Second, we will discuss how BNs can be used to overcome challenges associated with traditional risk prediction models and summarise recommendations for constructing BNs for clinical use in oncology.Material and methods The systematic review was performed according to the Cochrane guidance and reported according to the guidelines for the Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA). 7 The review protocol was registered prospectively in the PROSPERO database (CRD420251140130).Search strategy and selection criteria A literature search was performed in MEDLINE, EMBASE (via Ovid) from inception to 1 January 2025. The following keywords and all known synonyms for these keywords were used: ‘oncology’, ‘neoplasms’, ‘Bayesian network’. The search strategy can be found in online supplemental appendix A. Studies in a language other than English were excluded. Eligible studies investigated BNs as prognostic or predictive tools within oncology. Studies were included when the disease domain was within oncology (solid malignancies or haematological malignancies). Outcome measures included survival (overall survival or disease-specific survival); recurrence (local, regional or distant recurrence) or disease extent and staging (local tumour spread, lymph node metastasis, distant metastasis). Systematic reviews and conference abstracts were excluded. Studies in which BNs were applied as a tool to investigate genetic pathways or protein–protein interactions without correlation to one of the clinical outcome measures were excluded. Studies in which the BN structure was not included or could not be reconstructed were excluded as well.SP110.1136/bmjonc-2025-001040.supp1Supplementary dataData extraction A data extraction form was designed including information on study design, diagnosis, cohort (public, institutional, population-based, trial), type of BN, method for structure and parameter learning, included parameters, validation results and comparison of results with other prognostic models. Data extraction was performed by a single reviewer. Although no formal duplicate extraction was conducted, care was taken to ensure internal consistency by systematically checking extracted data for completeness and plausibility.Assessment of study quality The Prediction model study Risk Of Bias Assessment Tool, evaluating diagnostic or prognostic prediction models, was used for assessment of methodological quality and risk of bias of the included studies. 8 Risk of bias, referring to systemic flaws or limitations in the design possibly distorting study results, was either assessed as ‘low concerns’, ‘high concerns’ or ‘unclear concerns’ in the following four domains: participants, predictors, outcome and analysis. Concerns regarding applicability, referring to the extent to which the model from the primary study matches our systematic review question, were scored for the first three domains (participants, predictors and outcome) as ‘low concerns’, ‘high concerns’ or ‘unclear concerns’. The risk of bias domain ‘participants’ was scored based on the presence of appropriate data sources (cohort, case-control or randomised controlled trials) and the absence of inappropriate exclusions. For the risk of bias domain ‘predictors’ it was reviewed whether the predictors were assessed without knowledge of the outcome data, and whether all predictors were assessed in a similar way. For the risk of bias domain ‘outcome’ it was reviewed whether the outcome was determined appropriately and in a prespecified and similar way for all patients and without knowledge of the predictor information. For the risk of bias domain ‘analysis’ the following items were scored: presence of a reasonable number of participants, handling of continuous variables and missing data, external validation, avoidance of univariable predictor selection, handling of complexities (censoring, sampling of controls) and overfitting.Results Studies A systematic review of the literature yielded 1480 results, of which 860 remained after removal of duplicates ( online supplemental appendix B). After screening of title and abstract, 92 studies were deemed relevant for full text screening. A total of 52 studies were included after full-text screening.9–61The most important study characteristics are shown in table 1 and have been summarised in figure 2. Of all included studies, seven were published before 2014 (between 2006 and 2013), with most studies being published in 2022 (n=11).9 14 15 22 29 30 32 36 39 48 57 Most studies were on lung cancer (n=12),21 24 33 34 38 40 45 46 51 55 56 61 breast cancer (n=9)16 18 19 23 25 36 40 47 53 or prostate cancer (n=5).28 43 50 52 57 For construction of the BN, 22 studies used (multi-)institutional cohorts11–13 15 16 22–25 27 30 32 33 38 44 47 48 54 57 58 60 61; 3 studies used trial cohorts20 37 39; 10 studies used population-based cohorts34 35 43 46 49–51 55 56 59; 8 studies used public cohorts (eg, The Cancer Genome Atlas datasets)9 14 18 19 21 26 40 53; 8 studies used no construction cohorts or inferred data from systematic reviewing of the literature10 17 28 29 41 42 45 52; 1 study simulated a patient cohort.36 Median size for the construction cohorts was 438 patients when analysing all studies and 364 patients when excluding studies using population-based or public data.Figure 2Summary of study characteristics, including year of publication, construction cohort sizes, development techniques, validation, tumour sites and primary outcomes. HCC, hepatocellular carcinoma; TNM, tumour node metastasis.Table 1Study characteristicsAuthorTumour siteConstruction cohortOutcomeBN structure learningBN parameters learningType dataValidationBhawe et al9Mesenchymal glioblastomaPublic datasets (525 males, 354 females)Overall survivalML. No expert knowledgeParametric learning was performedGenetic aberrations, miRNA changes, clinical parametersStructure validated with cross-validation and sensitivity analysisBradley et al10Pancreatic ductal adenocarcinomaPubMed data (n=31.214)Postresection overall survivalAgenaRisk software. No EKNormalised weighting parent nodesInflammatory, tumour, patients factors, response to neoadjuvant therapy.External (n=387) AUC 0.7 (preop) and 0.8 (postop), AUC 0.7 when two variables missingCai et al12Hepatocellular carcinoma299 patients after hepatectomyPostresection overall survivalTAN Bayes algorithm. No EK.TAN Bayes algorithmPatient, tumour factors, treatment factorsInternal: accuracy and Fussell-VeselyCai et al11Gallbladder carcinoma366 patients after resection; 244 used for constructionPostresection overall survivalTAN Bayes algorithm. No EKTAN Bayes algorithmPatient, tumour factors, treatment factorsInternal (n=122): accuracy and AUC (0.78) and Fussell-VeselyCarreras et al13Diffuse large B-cell lymphoma105 patients; 72 for constructionOverall survivalML, no EKNot describedGene expression dataInternal (n=33). Acc was 93%.Carreras 202214Mantle cell lymphoma123 patients, 70% used for constructionOverall survivalML, no EKNot describedGene expression dataInternal, 30% of patients. Acc 85%.Carreras 202215Several haematological diagnosesPublic datasets n=2732, 70% for constructionOverall survivalML, no EKNot describedGene expression data and IHC expressionInternal, 30% of patients. Acc 89%.Chun et al16Breast cancerInstitutional cohort, n=2931Disease free survivalEKProbability tables inferred from expert opinionPatient, tumour factors, treatment factorsIn silico prediction of trial resultsCypko et al17HNSCCNoneTNM stagingEKUsing NCCN guidelinesPatient, tumour factorsIn 66 patients with HNSCC, AUC 0.98Ganapathy et al18Breast cancerOnline dataset, n=286, 70% for constructionRecurrenceTAN Bayes algorithm. No EKExpectation-maximisation methodPatient, tumour factors, treatmentInternal (30%), misclassification rate (36.4%), sens 11.7%, spec 87.2%, PPV 30.6%, NPV 65.6%Gevaert et al19Breast cancerPublic dataset (ITTACA), n=78, 70% used for constructionRecurrenceML. No EKMAP parameterisation of the Dirichlet distributionPatient, tumour factors, gene expression dataInternal (30%) AUC (0.747–0.793)Gupta et al20Renal cell cancerCheckmate 25 trial cohort, n=803Overall survivalML. No EKBayesian estimationPatient, tumour factors, serum measures, immune cellsExternal, IMDC dataset (n=2152), AUC 0.68Hozumi and Shimizu21Non-small cell lung cancerPublic datasets, n=315, 70% used for constructionClinical benefit immunotherapyTAN Bayes algorithm. HC. No EKLearnt from the datasetPatient, tumour factors, mutational dataInternal validation with 30% of patients, AUC 0.823Jang et al22Multicancer (a.o. sarcoma, breast cancer, melanoma)Institutional cohort, n=87Local controlAugmented naive Bayes algorithm.Target optimisation tree algorithmTumour factors, genetic information, radiotherapy detailsInternal, clinical BN (AUC 0.86), genomic BN (AUC 0.81), clinicogenomic BN (AUC 0.84)Jang et al23Breast cancerInstitutional cohort, n=2931RecurrenceUsing EKNot performedPatient, tumour factors, radiotherapy detailsNo validation performedJayasurya et al24Non-small cell lung cancerInstitutional cohort, n=322Overall survivalML. No EKExpectation-maximisation methodPatient factors, tumour factorsExternal, three cohorts (n=115), AUC 0.70–0.77Jiang et al26Breast cancerMETABRIC dataset, n=981Overall survivalML. No EKExpectation-maximisation methodTumour factors, molecular factorsInternal, AUC 0.666–0.688Jiang et al25Breast cancerInstitutional cohort, n=6726Distant metastasisMLExpectation-maximisation methodPatient, tumour factors, molecular factors, treatment factorsExternal, probability of being metastasis-free, n=1989Jochems et al27Lung cancerInstitutional cohort, n=698SurvivalUsing EKExpectation-maximisationPatient, tumour factors, RT detailsExternal, n=196, AUC 0.65–0.66Kalet et al28Prostate cancerNoneRecurrenceUsing EKNot performedPatient, tumour factors, RT detailsNo validation performedLadyzynski et al29Lymphocytic leukaemiaNoneSurvivalEK, literatureExpert estimates and literaturePatient, tumour factors, treatment detailsNo validation performedLi et al30Rectal cancerInstitutional cohort, n=705, 70% for constructionSurvivalTAN Bayes algorithmLearnt from the datasetPatient, tumour factors, molecular factors, tumour markersInternal validation with 30% of patients, AUC 0.801Luo et al32Hepatocellular carcinomaInstitutional cohort, n=77Local controlUsing EKLearnt from the datasetPatient, tumour factors, biochemical featuresExternal validation in 59 patients, AUC 0.72Luo et al33Non-small cell lung cancerInstitutional, n=118, 68 for constructionLocal controlML and EKLearnt from the datasetPatient, tumour factors, molecular factorsInternal (n=50), AUC 0.79–0.83Makond et al34Lung cancerNationwide cancer database, n=438SurvivalEKMax a Posteriori estimationPatient, tumour factors, primary treatment, regionInternal, acc 89.9%, sens 91.1%, spec 88.8%Makond et al35Two contingent primary cancersNationwide cancer database, n=7845Survival (5 years)Expert knowledgeMaximisation likelihood methodPatient, tumour factors, treatmentInternal, acc 82.0%, sens 96.4%, spec 68.1%Maksimenko et al36Breast cancerHypothetical cohortSurvivalEK, literatureInferred from literaturePatient, tumour factors, mutational dataNot performedNissan et al37Colorectal cancerTwo clinical trials, n=278Sentinel node statusML, no EKMaximisation likelihood methodPatient, tumour factors, treatmentInternal (PPV 83%, NPV 97%)Oh et al38Non-small cell lung cancerInstitutional cohorts, n=74Local controlMLLearnt from the datasetPatient, tumour factors, radiotherapy details, biochemical featuresInternal (acc 80.2%, AUC 0.83)Osong et al39Rectal cancer14 prospective clinical trials, n=6754, 80% for constructionLocal control (2, 3, 5 years)ML and EKLearnt from the datasetTumour factors, treatmentInternal (20%), AUC 0.8Park et al40Breast cancerPublic datasetsDistant metastasisMLLearnt from the datasetTumour factors, gene expression dataInternal, AUC 0.76–0.79Phillips et al41HNSCCNoneLymph node metastasisEK, literatureInferred from literatureTumour factors, radiology resultsNot performed; utility testingPouymayou et al42HNSCCNoneLymph node metastasisUsing EKMaximisation likelihood estimation, sens/specTumour factorsNo validation performedRegnier-Coudert et al43Prostate cancerTwo population-based cohorts, n=3236. One institutional cohort, n=85TNM stagingUsing machine-learning (K2GA algorithm)Learnt from the datasetPatient factors, tumour factors, tumour markersInternal validation, AUC 0.63 (OC), 0.58 (EPE), 0.69 (SVI), 0.81 (LNI).Reijnen et al44Endometrial cancerMulti-institutional cohort, n=763Lymph node metastasis, survivalML and EKMaximisation likelihood methodTumour factors, molecular factors, tumour markers, imagingExternal (n=830), AUC 0.82 (LNM), AUC 0.82–0.84 (survival)Schallenberg et al45Non-small cell lung cancerInstitutional cohort, n=447SurvivalEKNot performedPatient, tumour factors, molecular factorsNo, network used for identification of confoundersSesen et al46Lung cancerPopulation-based, n=36 480Survival (1 year)MLMaximum likelihoodPatient, tumour factors treatmentInternal, AUC 0.79 (K2), 0.81 (MCMC)Shahmirazlou 202347Breast cancer (recurrent)Institutional, n=220, 70% for constructionSurvivalScore-based MLLearnt from the datasetPatient, tumour factors, molecular factorsInternal (30%), average loss function 0.32–0.40Sheidaei et al48Gastric cancerInstitutional, n=760, 70% for constructionSurvivalScore-based MLLearnt from the datasetPatient, tumour factors, treatmentInternal (30%), bias assessmentSieswerda et al51Non-small cell lung cancerPopulation-based, n=1 46 084, 80% for trainingTNM staging, survivalEK (TNM-staging)Expectation-maximisation methodTumour factors (TNM staging)Internal (20%), AUC 0.81 (survival)Sieswerda 202349Colon cancerPopulation-based, n=982Recurrence, survivalEK and MLNot performedPatient, tumour factors, adjuvant treatmentNo, network used for identification of confoundersSieswerda 202350Prostate cancerPopulation-based, n=4121SurvivalEK and MLExpectation-maximisation methodPatient, tumour factors and markers, treatmentNot performedSmith et al52Prostate cancerNoneNodal+distant metastasisEKNot performed (inferred from literature)Tumour factors and markers, radiotherapy detailsNot performedSu 202354Peritoneal mesotheliomaInstitutional, n=154, 70% for constructionSurvivalEKLearnt from the datasetPatient, tumour factors, surgery detailsInternal (30%), AUC 0.74Su 202353Breast cancerPublic, n=23.384, 70% for constructionSurvivalML, EK for selectionBayesian estimationPatient, tumour factors, treatment factorsInternal (30%), AUC (0.9). External, AUC (0.871) (n=8128)Wang et al56Lung cancerPopulation-based, n=4555, 80% for constructionBrain metastasisEK, literatureMaximisation likelihood methodPatient, tumour factors, treatmentInternal (20%), acc 82%, sens 83.3%, spec 80.9%Wang et al55Lung cancerPopulation-based, n=1075, 80% for trainingSurvivalEK, literatureLearnt from the datasetPatient, treatment factorsInternal (20%), R² 93.6% (st I), 86.8% (st II), 67.2% (st III), 52.9% (st IV)Wee et al57Prostate cancerInstitutional, n=1915EPE, SVINaive Bayes classifierLearnt from the datasetTumour factors and markersInternal, AUC 0.76 (EPE), 0.88 (SVI), 0.70 (R1)Xu et al58Hepatocellular carcinomaInstitutional, n=995Recurrence (early and late)EKGibbs sampler, Broyden-Fletcher-Goldfarb-ShannoTumour factors, biochemical features, surgical detailsInternal, acc 0.57Zhang et al59Endometrial cancerPopulation-based, n=618Survival (5 years)TAN Bayesian algorithmLearnt from the datasetPatient, tumour factors, treatmentExternal (n=104), AUC 0.849Zhao et al60Pseudomyxoma peritoneiInstitutional, n=453, 80% for constructionSurvivalNot described,Not describedPatient, tumour factors, surgical detailsInternal (20%), AUC 0.735Zhong et al61Lung cancerInstitutional, n=5240, 2137 for constructionSurvival (3-year)EK, literature score-based MLMaximisation likelihood methodPatient, clinical factors, biochemical featuresInternal (n=1433), AUC 0.896Acc, accuracy; a.o., amongst others; AUC, area under the curve; BN, Bayesian network; EK, expert knowledge; EPE, extraprostatic extension; HC, hill-climbing; HNSCC, head and neck squamous cell carcinoma; IHC, immunohistochemistry; IMDC, International Metastatic Renal Cell Carcinoma Database Consortium; ITTACA, Integrated Tumor Transcriptome Array and Clinical data Analysis; LNI, lymph node involvement; LNM, lymph node metastasis; MAP, maximum a posteriori; MCMC, Markov Chain Monte Carlo; METABRIC, Molecular Taxonomy of Breast Cancer International Consortium; ML, machine learning; NCCN, National Comprehensive Cancer Network; NPV, negative predictive value; OC, organ confined; PPV, positive predictive value; RT, radiotherapy; sens, sensitivity; st, stage; SVI, seminal vesicle involvement; TAN, tree augmented naïve; TNM, tumour node metastasis.Most studies analysed overall survival, whereas 12 studies (also) analysed a measure of disease spread (local tumour spread, lymph node metastasis, distant metastasis).17 25 37 40–44 51 52 56 57Assessment of study quality For the domain ‘participants’ almost half of all studies had an unclear risk of bias because of the lack of description of inclusion and/or exclusion criteria ( online supplemental appendix C). The risk of bias and concerns regarding applicability are shown in figure 3. Studies that did not use a construction cohort were classified as high-risk of bias and low applicability. For the domain ‘predictors’ no concerns on risk of bias were raised. For the domain ‘outcome’ almost half of all studies were found to be at unclear or high risk of bias due to the lack of description of appropriate and standardised assessment of the outcome measure. For the domain ‘analysis’ more than half of all studies were found to be at unclear or high risk of bias because of the lack of description of handling complexities such as censoring, sampling or control patients; missing data and overfitting; the absence of external validation; and the appropriate handling of continuous data as usually continuous data are categorised for implementation into BNs. Finally, 44% were considered having low concerns on risk of bias, 27% (14 studies) were considered having unclear concerns and 29% (15 studies) were considered having low concerns for bias.Figure 3Summary plots for the risk of bias (A) and applicability (B) for the different domains within the PROBAST tool. PROBAST, Prediction model study Risk Of Bias Assessment Tool.Bayesian network structure learning There is a large variety in techniques applied to construct the BN structures, either by learning with data-based inference, using expert knowledge or information from medical literature, or a combination of both approaches. It is known that hybrid approaches improve structure learning processes. 5 Only six studies explicitly stated they used a combination of constraint-based or score-based ML techniques with expert knowledge,33 39 44 49 50 61 whereas most studies (n=27) applied ML techniques without mentioning the use of expert knowledge.9–15 18–22 24–26 30 37 38 40 43 46–48 53 57 59 60 Of these 27 studies, 8 studies included a tree-augmented naïve Bayes (TAN) classifier. In this technique a Naive Bayes model, that assumes all features to be conditionally independent given the target variable, is adapted so that each feature can directly depend on the target variable as well as one other feature. In this way, the simplicity of Naive Bayes is maintained while improving accuracy by modelling limited conditional dependencies (online supplemental appendix D). It is important to note that a TAN only differs from a Naive Bayes model when some predictors are unknown; when all are known, it becomes equivalent to a Naive Bayes model.Validation Eight studies performed external validation in a cohort that was independent from the construction cohort(s). 10 20 24 25 27 32 44 59 11 studies performed no formal validation or performed in silico prediction of results only.16 23 28 29 36 41 42 45 49 50 52 Of these, two studies used the BN for the identification of confounders to include into Cox proportional hazard analysis.45 49 External validation cohort sizes varied between 59 and 2152 (median 292) patients. For the external validation of survival outcomes (overall survival, recurrence-free survival), area under the curve (AUC) varied between 0.65 and 0.85 (median AUC 0.80). One study performing external validation evaluated the prediction of disease spread (lymph node metastasis), with an AUC of 0.82.44 One study evaluated the impact of missing variables on outcome prediction with an AUC decreasing from 0.8 to 0.7 when more than six data points were missing.10 The AUC for studies performing internal (cross-)validation only varied between 0.62 and 0.98.For studies combining expert knowledge with ML algorithms to construct the BN structure, AUCs varied between 0.82 and 0.84 in external validation44 and 0.79 and 0.90 in internal validation.33 39 61 For studies using expert knowledge only, AUCs were found between 0.65 and 0.6627 in external validation and 0.72 and 0.81 in internal validation.32 51 54 For studies using ML only, AUCs varied between 0.70 and 0.87 in external validation,10 24 54 and 0.68 and 0.90 in internal validation.19 20 38 40 43 46 53Most studies (n=32) did not compare BN performance with other types of prognostic models (online supplemental appendix C). 13 studies compared BN with a Cox proportional hazards model.9 13–15 18 26 30 34 35 48 53 59 61 In most of these comparisons, the BNs gave a higher AUC than Cox regression analysis, although mostly no formal statistical testing was performed to compare the AUC values.Discussion Main findings With this systematic review we have identified published literature on the development and validation of BNs as prognostic tools within the field of medical oncology. More than half of all studies harboured unclear or high concerns on risk of bias and results should be interpreted with caution. Most studies used ML techniques to construct the BN, and did not implement expert knowledge within the process. Only a minority applied external validation to test the developed BN. It was shown that hybrid construction techniques yielded high validation performance metrics, underlining the importance of combining expert knowledge with ML techniques. Studies comparing BN validation with Cox regression validation showed better performance for BNs in most studies.Advantages of Bayesian networks Traditional regression-based models (such as Cox, logistic and linear regression) estimate an outcome of interest by learning the mathematical relationship between predictors (covariates) and the target variable ( table 2). These models typically use a baseline risk or intercept, combined with a linear combination of covariates weighted by regression coefficients to estimate the outcome. The estimation of regression coefficients is usually performed by either maximising the likelihood (eg, logistic and Cox regression) or minimising the sum of squared errors (eg, linear regression). Cox proportional hazards regression (CPHR) models the hazard (risk over time) rather than a direct outcome. A widely known example is the PREDICT tool for breast cancer, in which a multivariable CPHR is used to estimate survival outcomes and the benefit of adjuvant therapy for patients with breast cancer using a set of covariates (age, tumour size, nodal status, histological grade, oestrogen receptor status, HER2 status, detection mode, adjuvant treatment).62 This model, extensively validated in large patient cohorts, provides clinicians and patients with individualised risk assessment using a web-based tool (https://breast.predict.nhs.uk).63 64 These models are well-known and widely adopted within the medical clinical field, are easy to implement and interpret and are able to model a variety of outcome types (binary, continuous, time outcomes). Although they can model correlations based on well-established statistical inference techniques, they are unable to actually model causal correlations between predictors. Importantly, in order to get a sensible estimation of the target variable, all required covariates have to be known. With other words, regression models are unable to handle missing data. Especially for complex multivariable models this can hinder clinical use as they may lead to biased estimates, reduced statistical power and exclusion of potentially valuable patient data, when complete-case analysis is used.Table 2Comparison between Bayesian networks and regression modelsAdvantagesDisadvantagesRegression models Well-known and widely adopted within medical science and clinical fieldCannot model causal correlations Computationally inexpensiveEach outcome of interest requires a separate model Provide well-established statistical inference techniquesThey operate under rigid assumptions about the relationships between variables (eg, linearity, independence) Easy to implement and interpretCannot handle missing data without imputation Able to model binary, continuous and time outcomesBayesian networks Graphical nature makes them easily interpretableLess widely known and adopted within medical science and clinical field Can explicitly reflect causal relationships between variablesRequires expert knowledge to model network structure correctly Can be used for individualised risk estimationsComputationally expensive for large, complex networks Can handle missing dataStructure learning can be challenging without sufficient data Capable of modelling complex, non-linear correlations Multiple outcomes can be modelled within a single framework Enable counterfactual reasoning (‘what-if’ scenarios) Can be developed even with limited data (using expert knowledge)On the contrary, BNs are graphical representations of joint probability distributions (JPD) of a set of variables of interest. These probabilities are visualised with probability tables (nodes), and arcs (edges) representing causal relationships between the variables. The probability table includes either marginal probability distributions (MPD, if a variable has no parent nodes or evidence for the parent node is not provided to the network), or CPDs (if a variable has one or more parent nodes, with evidence provided). CPDs can be used to answer queries about, for example, treatment effects or outcomes by computing the likelihood of different outcomes for a node based on observed values of its parent nodes (figure 1). This is relevant in oncology as the golden standard (randomised controlled trials are often not available, time-consuming and may have limited applicability due to strict inclusion criteria and small sample sizes. Importantly, including treatment effects into such models requires careful reporting of the target treatment(s) and target populations, how these treatments were assigned in the development data, and how they were handled during model development.65 In addition, effective and robust causal inference techniques, developed to estimate treatment effects using observational data, should be applied for adequate development.66As described, the probability distributions are based on Bayes’ rule, which combines a prior probability with the likelihood of the observed data to calculate posterior probability. As a BN represents a full JPD of all included variables, it allows information to propagate from the observed variables to the variables of interest through explicit connections defined in the graph. These probabilities can be learnt from datasets, but can also be inferred from expert knowledge, when insufficient data is available. To illustrate, figure 1 represents a simple BN for the estimation of 5-year overall survival, based on tumour grade, the presence of lymph node metastasis and tumour stage. In panel B, MPDs are shown, for example, 10% probability of lymph node metastasis and 70% probability of 5-year overall survival. If evidence is provided to the BN (panel C), for example, tumour grade 3, CPDs are updated. Variables directedly connected are updated (eg, 40% probability of stage III, compared with a marginal probability of 10%), as well as variables connected indirectly (eg, 5-year overall survival; information is propagated through the intermediate node ‘stage’). Several important advantages of BNs are illustrated within this BN. First, a (simplified) causal relationship is modelled within this framework, for example, high tumour grade causes lymph node metastasis, which is intuitively depicted within a graphical representation. As the relationships between the variables are visualised in a transparent way (eg, tumour grade is correlated directly to advanced tumour stage, but also indirectly through lymph node metastasis), it allows clinicians and patients to actually understand how predictions are achieved. Also, the individual contributions of the different variables to the prediction of the outcome of interest are being made comprehensible. Importantly, for BNs used for prognostication, relationships do not necessarily have to be causal. Hybrid techniques, combining expert knowledge with (score-based) ML methods, improve the performance of BNs by guiding algorithm learning (eg, to excluding non-causal or unrealistic correlations); reducing overfitting (by filtering out irrelevant variables or correlations); narrowing the search space of ML algorithms, thus improving efficiency and focus; and enhancing the identification of relevant confounders. A second advantage is that, using Bayes’ rule, the outcome of interest (in this case 5-year overall survival) can be estimated even when not all input variables are known. To illustrate, when only tumour grade (grade 3) is provided to the network (panel C), the 5-year overall survival is updated (50%). We can provide new information to the network, for example, lymph node metastasis=yes, updating the probability of 5-year overall survival to 40% (panel D). This clearly demonstrates both the simplicity and power of BNs for outcome prediction. Of the identified studies, only few explicitly exploited the quality of BNs of dealing with missing variables. Bradley et al developed a BN on survival in pancreatic ductal adenocarcinoma and evaluated the impact of at least two missing variables on the model’s AUC.10 Reijnen et al developed a BN for patients with endometrial cancer and allowed predictions for the external validations when input variables were missing, defining a ‘minimally required input variable set’.44 Third, these networks can be used for counterfactual reasoning (‘what-if’ scenarios) using Bayes’ rule. In this example, given that the patient actually had a grade 3 tumour and died within 5 years, we can investigate what the probability of 5-year survival would be in case of a grade 1 tumour. It can thus serve as ‘generative networks’ by simulating data using the causal structure of the network. Importantly, this concept of ‘what-if’ reasoning is especially appealing to compare the effect of different treatment options which are unobserved, even when randomised controlled data is not available. Even though randomised controlled trials are considered the gold standard to estimate treatment effects, they are costly to perform and are time-consuming, which can be a significant problem in a rapidly evolving medical field. Observational data are usually more readily available, but pose a challenge because of biases (by indication) that may overestimate or underestimate treatment effect. By structuring a BN using effective causal inference techniques, we can model how treatment variables influence outcomes, by also incorporating other relevant variables that may act as confounders (eg, age, sex, performance status) or mediators. For these estimations to be as unbiased as possible, it is of great importance that all confounding variables are modelled as such, for example, are incorporated as parent nodes for both the treatment and the outcome (common cause criterion).67 Other criteria to correct for confounders include the pretreatment criterion (select all pretreatment variables as confounders), the backdoor path criterion (similar to the common cause criterion, but also able to correct for indirect confounders) and the disjunctive cause criterion (correcting for variables assumed to be either a cause of exposure or outcome).67 Several studies included treatment as an input variable within the BNs (table 1), however only a few used formal causal inference techniques for structure learning and evaluation of the treatment effects. To illustrate, Sieswerda et al integrated causal knowledge into a BN to evaluate the effect of active treatment in prostate cancer. By applying the backdoor path criterion, they identified and controlled for confounders, effectively mitigating selection biases and improving the causal inference of treatment effects in observational data.50Fourth, whereas regression models do not allow for the estimation of multiple outcomes, and each outcome has to be evaluated with its own model, BNs do allow the estimation of multiple outcomes in one model (eg, disease stage and 5-year overall survival). Of the included studies, only a few exploited this capacity. Several studies developed BNs to predict oncological outcomes (eg, survival and recurrence) at several time points.26 39 49 58 Reijnen et al developed a BN for patients with endometrial cancer predicting both disease spread patterns (lymph node metastasis) and survival.44 Smith et al developed a BN to predict both lymph node metastasis and distant metastasis in patients with prostate cancer, and Wee et al developed a BN to predict extracapsular extension and seminal vesicle invasion in prostate cancer.52 57 Fifth, BN parameters can be updated continuously when new evidence becomes available (eg, a new dataset or extension of existing datasets with new participants), without the entire model having to be re-trained. In this way, BN can be kept up-to-date over time, in the case of evolving guidelines or new parameters becoming available.68 To illustrate, Reijnen et al recently re-optimised the Preoperative risk stratification in endometrial cancer BN with two novel molecular biomarkers, to create ENDORISK-2.69 Preoperative risk stratification in endometrial cancer ENDORISK BN with two novel molecular biomarkers, to create ENDORISK-2. 69Limitations of Bayesian networks Despite the advantages, developing BNs can be challenging because of the requirement for expert knowledge to adequately model the network’s structure. Also, computing the most optimal network structure can be computationally expensive. Several algorithms, for example, heuristic score-based algorithms, are available to decrease the number of possible structures and in this way the required computational power. Especially in complex networks, large datasets are required for network structure learning.Limitations This systematic review has several limitations to be acknowledged. First, the overall concerns on bias were unclear or even high in a majority of the included studies. This was mainly attributable to the lack of external validation, unclear description of predictor and outcome assessments and the lack of correction for complexities within the datasets. Especially the lack of external validation makes it hard to draw final conclusions on the actual performance of the included BNs. Only a minority of included studies applied a hybrid technique to develop the network’s structure, even though these are known to significantly improve ML processes. Most studies used ML only for structure development, and did not specify what type of algorithm was actually used.Future directions and clinical recommendations Despite the advantages and potentials of BNs several important steps and challenges must be addressed within the technical, clinical and regulatory domains for these networks to be actually used.As shown within this review, combining ML techniques with expert knowledge to construct the network structures, yielded high validation performance metrics and should be the standard for development of BNs for oncological care. BNs should be validated in an external cohort to independently test its performance and generalisability, preferably in multiple cohorts. After rigorous internal and external validation for assessment of the test qualities, the clinical utility should also be tested in a prospective manner. With these prospective trials several questions should be answered: does the network actually improve clinical decision making; does it outperform existing guidelines or prediction tools, and in this way actually improve clinical outcome? Another important aspect to consider within the health technology assessment (HTA) would be the investigation of cost-effectiveness, patient-reported outcomes and organisational aspects, such as impact of healthcare delivery. Even though randomised controlled trials, possibly in a stepped wedge design, are considered the gold standard to answer these questions, they are time consuming and costly to perform. Alternative evaluation strategies have therefore been developed such as the RAPID-RT research programme, using iterative ‘rapid-learning’ cycles with real-world data to evaluate the effect of interventions within the field of radiotherapy.70Closely related to rigorous testing with HTA, these AI-based models need to be tested according to the medical device regulation or in-vitro device regulation, and need to be subjected to close post-marketing surveillance to guarantee sustainable and high-quality access.If rigorous prospective testing has shown additional benefit, they have the potential to improve individualised decision-making by implementation into decision support systems, clinical workflows including multidisciplinary tumour boards and shared decision making with the patient.Conclusion In this systematic review we summarised the published BNs developed as prognostic tools within a broad range of oncological diagnoses. Quality of existing literature is varying with considerable risk of bias and lack of external validation, and results should be interpreted with caution. External validations show very good performance metrics when adequate hybrid construction techniques have been applied. Several advantages, including the possibility to model causal relationships, apply counterfactual reasoning, work with missing variables, be updated when new evidence is available and predict multiple outcomes in one model, make these networks very attractive for clinical implementation to guide decision-making in oncology.