BetaEntity Annotation Prototype
← Back to diseases

Annotated full text

Qualified prediction system for allograft failure in real world settings: extended validation study

bmjmed · 2026-05-11 · canonical JSON source

36 visible annotations · policy: published · automated confidence ≥ 75.00%

Document resource

WHAT IS ALREADY KNOWN ON THIS TOPIC Accurate prediction of long term kidney allograft outcomes is critical for patient management and clinical trial designThe integrative Box (iBox) system, integrating functional, immunological, and histological parameters, has been internationally validated and qualified by the European Medicines Agency as a surrogate endpointThe performance of the iBox system in heterogeneous real world settings, including simplified versions adapted to local practices and data availability, and under race-free estimated glomerular filtration rate (eGFR) equations, has not been comprehensively assessedWHAT THIS STUDY ADDS The iBox showed strong discrimination, overall fit, and clinical utility in 12 683 kidney transplant recipients from 21 centres worldwide, consistently outperforming eGFR slope and anti-HLA donor specific antibodiesPredictive performance was maintained across flexible iBox versions, race-free eGFR equations, diverse clinical scenarios, and immunosuppressive strategies, with prediction periods extending up to 10 yearsAccounting for the competing risk of death did not improve performanceHOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE, OR POLICY These results support the iBox as a robust and generalisable prognostic tool for real world risk stratification in kidney transplantationThe extensive validation strengthens its use as a surrogate endpoint in clinical trials and supports integration into routine clinical careThe iBox may also facilitate shared decision making and patient engagement by providing individualised, quantitative estimates of graft prognosisIntroduction Need for improving kidney transplant care Maximising allograft longevity in kidney recipients is a key objective, 1 given the worldwide organ shortage2 and the increasing number of patients with end stage kidney disease.3 Accurate prediction of long term kidney allograft survival has therefore become crucial because it may improve risk stratification, patient monitoring, and clinical decision making. Accurate prediction of survival might also enhance the design, costs, and length of clinical trials. This strategy has been recommended by the American Society of Transplant Surgeons and the American Society of Transplantation.3Surrogate endpoint of allograft failure In response to these needs, several research teams have attempted to develop allograft failure prediction models. 4–9 In 2019, our group, together with a transplant consortium of more than 20 key opinion leaders from 17 institutions in four countries, published the integrative Box (iBox),10 a prediction model for long term risk of allograft failure which integrates functional, immunological, and histological factors evaluated at any time point after transplantation. Subsequent studies further showed the clinical relevance of the iBox as a surrogate endpoint in two randomised controlled trials, the TRANSFORM (Transplant Efficacy and Safety Outcomes With an Everolimus Based Regimen) and felzartamab trials.11 12 The iBox also outperformed clinicians in predicting allograft failure,13 performed better than machine learning based prediction models,14 performed well in paediatric patients,15 and maintained consistent performance when accounting for death as a competing event compared with standard survival analysis.16The robustness and reliability of the iBox have been increasingly recognised by regulatory authorities.17 18 19 After its endorsement by the American Society of Transplantation and qualification by the European Medicines Agency as an efficacy surrogate endpoint for clinical trials,20 the US Food and Drug Administration has recently accepted its qualification plan as a co-primary endpoint, a major step toward its implementation as a validated endpoint in clinical trials and routine healthcare.Need for further validation Extensive validation of a prediction model is crucial to ensure its successful deployment and large scale use. 21 This process is particularly important given the heterogeneity in practices among transplant centres worldwide. Centres vary considerably in resource availability, protocols, and routine practices. Some centres do not regularly perform biopsy or anti-HLA antibody assessment.22 Other centres differ in their methods of measuring proteinuria (ie, with a dipstick or by using the protein-to-creatinine ratio).23The clinical context also presents multiple challenges that require specific validation, including the performance of the model when using the new race-free estimated glomerular filtration rate (eGFR) equations (ie, without including race as a factor in the calculation),24 25 in patients with recurrence of the initial nephropathy or BK virus nephropathy, and across different immunosuppressive strategies, such as calcineurin inhibitors and mTOR (mechanistic target of rapamycin) inhibitors. The accuracy of predictions over extended follow-up periods26–28 also warrants investigation. Furthermore, the diversity of healthcare systems and geographical regions make it essential to validate the model across multiple continents.Study objective In this study, we included a large number of patients and transplant centres with various allocation and medical systems and defined a series of clinical scenarios reflecting the diversity of current clinical practice. We then investigated the prediction performance of the iBox in each clinical scenario to provide comprehensive evidence of its reliability and adaptability across various real world settings.Methods Study design and participants Derivation cohort The derivation cohort of the iBox ( NCT03474003) comprised 4000 consecutive patients, aged >18 years, who were prospectively enrolled at the time of kidney transplantation from a living or deceased donor at four institutions, between 1 January 2005 and 1 January 2014, as described previously.10 Patients’ follow-up was updated until 1 November 2024 to investigate the long term performance of the iBox. All patients provided written informed consent at the time of transplantation. Clinical data were collected from each centre and entered into the Paris Transplant Group database system (French Data Protection Authority registration No 363505), with a structured protocol to ensure harmonisation across study centres.Validation cohorts External validations included 8683 kidney transplant recipients, aged >18 years, who received kidneys from either living or deceased donors, recruited from 17 centres in 10 countries. A total of 5442 recipients were recruited in Europe, 2570 recipients in North America, and 671 recipients in South America, all undergoing transplantation between 2000 and 2022. These patients had at least one computable iBox score after transplant, defined as at least one adequate allograft biopsy (according to the Banff criteria) as part of the standard of care, with available data at the time of biopsy for eGFR, urine protein-to-creatinine ratio, and anti-HLA donor specific antibodies, or a concomitant evaluation of eGFR and urine protein-to-creatinine ratio only, or together with donor specific antibodies, when a biopsy was not available. Online supplemental methods 1.1 has details about the external validation cohorts and data collection process.SP110.1136/bmjmed-2025-002088.supp1Supplementary dataOutcome measures The outcome of interest was kidney allograft loss, defined as a patient's definitive return to dialysis or pre-emptive kidney re-transplantation. For patients who died with a functioning allograft, allograft survival was censored at the time of death as a functioning allograft.iBox multimodal prediction algorithm The iBox prediction model was designed to predict long term death censored graft failure and was developed in the derivation cohort. Firstly, we performed univariable Cox proportional hazards regression analyses to identify prognostic factors associated with long term kidney allograft failure. Thirty one candidate predictors were tested, including recipient demographics, donor characteristics, transplant related variables, allograft functional parameters, immunological parameters, and allograft histopathological lesions, scored according to the Banff classification. Variables showing associations with allograft failure in univariable analyses (P<0.10) were subsequently included in a multivariable Cox proportional hazards model. We applied backward stepwise elimination, and only independent predictors of allograft failure were retained in the final model.The iBox model integrates eight independent parameters: (1) time between the date of transplantation and the date of risk evaluation, with allograft function assessed by (2) eGFR and (3) urine protein-to-creatinine ratio (log transformed), (4) circulating anti-HLA donor specific antibodies with a qualitative validated binary mean fluorescence intensity cut-off value of 1400, and allograft pathology data with (5) transplant glomerulopathy (chronic glomerulopathy Banff score), (6) microvascular inflammation (sum of glomerulitis and peritubular capillaritis Banff scores), (7) interstitial fibrosis and tubular atrophy Banff score, and (8) interstitial inflammation and tubulitis Banff score, scored according to the Banff classification.29We assessed kidney allograft function by the Modification of Diet in Renal Disease Study equation (MDRD186),30 and proteinuria level with the protein-to-creatinine ratio. Circulating donor specific antibodies against HLA-A, HLA-B, HLA-Cw, HLA-DR, HLA-DQ, and HLA-DP were assessed with single antigen flow bead assays (One Lambda, Canoga Park, CA, USA) on a Luminex platform in the derivation cohort and according to local centre practice in the validation cohorts. Kidney allograft pathology data, including elementary lesion scores and diagnoses, were recorded according to the 2019 Banff classification,29 with the derivation cohort additionally updated according to the Banff 2022 classification.31 Online supplemental methods 1.2 describes the rationale for variable types and transformations.The iBox was computed at the time of risk evaluation after transplantation, which was conducted at the time of allograft biopsy, performed for clinical indication or as per protocol, according to the centres' practices. At the time of risk evaluation, recipients underwent concomitant evaluation of eGFR and proteinuria, allograft biopsy (Banff lesion scores and diagnoses), and circulating anti-HLA antibodies. For patients without a biopsy in the external validation cohorts, the time of risk evaluation was defined as the time of assessment of kidney function (concomitant evaluation of eGFR and proteinuria), performed at a minimum of one month after transplantation.Length of follow-up started with the risk evaluation of the patient, up to the date of loss of the kidney allograft, death with a functioning graft, loss to follow-up, or the end of the study. For patients who were followed up for more than seven years, follow-up was right censored at seven years in the main analysis. The performances of the iBox can be evaluated at any time point after the risk evaluation up to seven years. For clarity of presentation, we focused on three, five, and seven years.Validation of the iBox algorithm across diverse healthcare environments The iBox requires patient level data that are inherently heterogeneous and reflect the diversity of the kidney transplant recipient population globally. To ensure the generalisability of the iBox while maintaining its prognostic capabilities, we investigated its performance: in differently informed datasets and medico-economic systems; with various kidney function equations; in different clinical scenarios and therapeutic strategies; with histological diagnoses instead of Banff scoring lesions; and in extended follow-up periods after risk evaluation. Online supplemental methods 1.3 and 1.4 have details on the different versions of the iBox algorithm described in the sections below, including the variables and their functional forms, coefficients, development and validation datasets, and prediction periods.Performance of the iBox across variably informed datasets and distinct healthcare economic systems To accommodate the heterogeneity in clinical practices among countries and transplant centres worldwide, we developed flexible versions of the iBox algorithm. These versions are designed for use in centres that do not perform biopsies or measure anti-HLA donor specific antibodies (functional iBox), centres that do not perform biopsies (functional-immunological iBox), centres that do not measure anti-HLA donor specific antibodies (functional-histological iBox), and centres that monitor patients with a dipstick instead of measuring urine protein-to-creatinine ratio (iBox with dipstick; online supplemental methods 1.5 shows the conversion for proteinuria).Performance of the iBox with different kidney function equations The iBox was originally developed with GFR estimated with the MDRD 186 equation.30 We investigated its performance in the derivation cohort when applying different eGFR equations, including: MDRD175 equation,32 Chronic Kidney Disease Epidemiology Collaboration (CKD-EPI) 2009 equation,33 race-free CKD-EPI 2021 equation,24 European Kidney Function Consortium equation,34 and race-free kidney recipient specific equation.25 The eGFR coefficient of the iBox algorithm remained unchanged and was based on the MDRD186 equation.Performance of the iBox in different clinical scenarios and therapeutic contexts We investigated the performance of the iBox in the derivation cohort in patients with recurrence of the initial nephropathy, in patients with BK virus associated nephropathy, in patients treated with calcineurin or mTOR inhibitors, in patients who had a protocol biopsy, and in patients who had an indication biopsy.Performance of the iBox based on histological diagnoses and histological indexes instead of Banff scoring lesions We developed versions of the iBox algorithm where Banff lesion scores were replaced with either histological diagnoses, graded according to the 2022 Banff classification, 31 or histological indexes, as described in Haas et al34 and in Vaulet et al35 (online supplemental methods 1.6 and 1.7 show the calculation of these indexes).Performance of the iBox with extended follow-up Lastly, we investigated the predictive performance of the iBox when extending the prediction period up to 10 years after risk evaluation in the derivation cohort because the prediction period was restricted to a maximum of seven years in the original publication.Performance of the iBox versus eGFR slope and circulating anti-HLA donor specific antibodies We compared the predictive ability of the iBox with eGFR slope. In the derivation cohort, we selected patients with creatinine measurements at about three and 12 months after transplantation and calculated the eGFR slope between these time points with the MDRD 186 equation. A Cox model with the eGFR slope and the timing of risk evaluation as predictors was fitted to the data. The iBox was applied to the same patient subset with the eGFR measurement at 12 months after transplantation.We also compared the performance of the iBox with that of the mean fluorescence intensity of anti-HLA donor specific antibodies, with a Cox model where mean fluorescence intensity (qualitative binary cut-off of 1400) and timing of risk evaluation were predictors. Figure 1 presents a comprehensive synthesis of the multiple iBox validations described in this section.Figure 1Design of the integrative Box (iBox) system extended validations. AMR/MVI=antibody mediated rejection/microvascular inflammation; CKD-EPI=Chronic Kidney Disease-Epidemiology Collaboration equation; CNI=calcineurin inhibitor; DSA=donor specific antibodies; eGFR=estimated glomerular filtration rate; EKFC=European Kidney Function Consortium equation; KRS=kidney recipient specific equation; MDRD=Modification of Diet in Renal Disease Study equation; mTOR=mechanistic target of rapamycin inhibitor; TCMR/TI=T cell mediated rejection/tubulointerstitial inflammationCompeting risks analysis To account for the competing risk of death with a functioning graft in the prediction of long term graft failure, we modelled all flexible versions of the iBox algorithm with Fine-Gray subdistribution hazards models and assessed their predictive performance in the derivation and external validation cohorts ( online supplemental methods 1.8).Statistical analysis Descriptive statistics Continuous variables are described as mean (standard deviation (SD)) or median (interquartile range (IQR)). Means and proportions between groups were compared with the Student's t test, analysis of variance (Mann-Whitney test for mean fluorescence intensity), or the χ 2 test (or Fisher's exact test if appropriate). P values <0.05 were considered significant, and all tests were two tailed.Evaluation of model performance We assessed the accuracy of the iBox prediction model based on its discrimination, calibration, overall fit, and clinical utility in the derivation cohort and in the three external validation cohorts. Online supplemental methods 1.9 provides comprehensive details on the calculation of prediction performances. The chosen prediction periods were seven and 10 years after risk evaluation in the derivation cohort, five years after risk evaluation in the European and North American validation cohorts, and 2.5 years after risk evaluation in the South American cohort, to align with the median follow-up in each cohort and ensure sufficient statistical power. Statistical comparisons of Brier scores and concordance values between models were conducted following Blanche et al36 and Therneau and Atkinson,37 respectively.Missing data The original iBox algorithm development excluded 59 patients (1.5%) from the derivation cohort, as described in the original publication. 10 In the validation cohorts, complete cases were used for iBox calculation. Depending on the iBox algorithm version used, we excluded 94-2049 patients (1.7-37.7%) from the European cohort, 70-614 patients (2.7-23.9%) from the North American cohort, and six patients (0.9%) from the South American cohort. The iBox deployment strategy requires complete predictor data for the chosen algorithm version, with different versions available to accommodate differences in systematic data availability across clinical centres.Software All analyses were performed with Stata (StataCorp 2021, Stata Statistical Software: release 17, College Station, TX, USA) and R (The R Foundation for Statistical Computing, Vienna, Austria).Study protocol The study protocol is available at https://osf.io/65ukn/. The study is reported consistent with the TRIPOD+AI (transparent reporting of a multivariable prediction model for individual prognosis or diagnosis+artificial intelligence) guidelines38 (online supplemental methods 1.10).Patient and public involvement Two kidney transplant recipients, representing patient support and advocacy consulting association, were involved in discussions about the study objectives and the potential value of prediction tools, such as the iBox. These recipients provided feedback on the clinical relevance and acceptability of the prediction model, particularly about its potential to enhance patient engagement, activation, and shared decision making in transplant care. The results of this study will be disseminated to kidney transplant recipients through patient associations.Results Characteristics of the derivation and validation cohorts The study comprised 12 683 participants, with 4000 in the derivation cohort and 8683 in the validation cohorts. Table 1 shows the characteristics of the different cohorts. Median time from transplantation to risk evaluation was 0.98 years (IQR 0.27-1.07) in the derivation cohort and 0.97 years (0.30 to 1.07) in the validation cohorts. Graft loss and death with a functioning graft were found in 549 (13.7%) and 414 (10.5%) patients, respectively, in the derivation cohort, and in 991 (11.4%) and 468 (7.9%) patients, respectively, in the validation cohorts. Median follow-up time after risk evaluation was 5.78 years (IQR 3.51-7.00) in the derivation cohort and of 4.68 years (2.48-7.00) in the validation cohorts.Table 1Characteristics of the derivation and validation cohorts French derivation cohort(n=4000)European validation cohort(n=5442)North American validation cohort(n=2570)South American validation cohort(n=671)Total No of patientsCharacteristics (No (%), mean (SD), or median (IQR))Total No of patientsCharacteristics (No (%), mean (SD), or median (IQR))Total No of patientsCharacteristics (No (%), mean (SD), or median (IQR))Total No of patients Characteristics (No (%), mean (SD), or median (IQR))Recipient characteristics  Mean (SD) age (years)400050 (14)544251 (14)257050 (14)67144 (15) No (%) men40002450 (61.3)54423426 (63.0)25571516 (59.3)671400 (59.6)Cause of end stage renal disease (No (%)) Glomerulonephritis 40001086 (27.2)43481395 (32.1)2559719 (28.1)671144 (21.5) Diabetes4000438 (10.9)4348467 (10.7)2559580 (22.7)67172 (10.7) Vascular4000296 (7.4)4348328 (7.5)2559423 (16.5)671129 (19.2) Other40002180 (54.5)43482158 (49.6)2559837 (32.7)671326 (48.6)Donor characteristics Mean (SD) age (years)400052 (16)404850 (15)237541 (15)67145 (15) No (%) men40002151 (53.8)41912293 (54.7)23741202 (50.6)671356 (53.1) No (%) hypertension39031005 (25.7)2185577 (26.4)1565243 (15.5)NANA* No (%) diabetes mellitus3861231 (6.0%)202665 (3.2)127147 (3.7%)NANA*Donor type (No (%)) Deceased donor40003327 (83.2)48323993 (82.6)25701397 (54.4)671511 (76.2) Death from cerebrovascular disease†33271864 (56.0)26391331 (50.4)746209 (28.0)NANA* Expanded criteria donor39951409 (35.3)2450785 (32.0)2058180 (8.75)671169 (25.2)Transplant characteristics No (%) previous kidney transplant 4000605 (15.1)4560585 (12.8)2474411 (16.6)67145 (6.71) Mean (SD) cold ischaemia time in deceased donors (hours)397616.2 (9.0)421612.8 (7.5)214311.4 (11.2)66617.0 (10.2) Mean (SD) HLA-A/B/DR mismatch number40004.8 (1.4)52204.1 (1.5)24224.7 (1.7)6623.7 (1.4) No (%) delayed graft function38971046 (26.8)44011421 (32.3)1801255 (14.2)671311 (46.3) Median (IQR) time from transplantation to risk evaluation (years)40000.98 (0.3-1.1)54420.5 (0.3-1.0)25701.0 (0.5-1.1)6712.1 (1.0-3.5)Functional parameters at time of risk evaluation Mean (SD) eGFR (mL/min/1.73 m2)400050 (19)544050 (21)253753 (22)67138 (19) Median (IQR) proteinuria (g/g)40000.2 (0.1-0.4)53480.2 (0.1-0.4)25320.1 (0.1-0.3)6660.3 (0.1-0.8)Immunological parameters at time of risk evaluation Anti-HLA donor specific antibody mean fluorescence intensity (No (%))  <140040003659 (91.5)35613347 (94.0)23271889 (81.2)671616 (91.8)  ≥1400 4000341 (8.5)3561214 (6.0)2327438 (18.8)67155 (8.2)*Data not available for the South American cohort.†Number and per cent were calculated among deceased donors.eGFR, estimated glomerular filtration rate; IQR, interquartile range; NA, not available; SD, standard deviation.Predictive performance of the various iBox algorithms adapted for data availability The different iBox algorithms were: full iBox (eGFR, continuous proteinuria, anti-HLA donor specific antibody mean fluorescence intensity, and histological lesions); functional iBox (eGFR and continuous proteinuria); functional-immunological iBox (eGFR, continuous proteinuria, and anti-HLA donor specific antibody mean fluorescence intensity); functional-histological iBox (eGFR, continuous proteinuria, and histological lesions); full iBox with dipstick (eGFR, dipstick proteinuria, anti-HLA donor specific antibody mean fluorescence intensity, and histological lesions); and functional iBox with dipstick (eGFR and dipstick proteinuria). Tables 2 and 3 show the predictive performances of these models in the derivation and external validation cohorts. Figures 2 and 3 show the calibration plots and decision curves for the full iBox for the derivation and external validation cohorts, and online supplemental figures 1-8 show the calibration plots and decision curves for the other iBox algorithms. Online supplemental tables 1-4 present time dependent discrimination assessed every year after the risk evaluation for the derivation and validation cohorts. Online supplemental table 5 shows average calibration across all time points after risk evaluation for the derivation and validation cohorts.Table 2Predictive performance (discrimination, calibration, and overall fit) of the different integrative Box (iBox) algorithms in the derivation cohortiBox algorithm*No of patientsDiscriminationCalibrationOverall fitHarrell's C index (95% CI)Slope (95% CI)Intercept (95% CI)Observed/expected ratio (95% CI)Brier score (95% CI)Index of predictive accuracy (%) (95% CI)Royston's R2DFull iBox39410.81 (0.79 to 0.83)0.86 (0.77 to 0.94)−0.12 (−0.23 to −0.02)0.93 (0.87 to 1.00)0.10 (0.09 to 0.11)26 (23 to 30)0.52Functional iBox39410.79 (0.77 to 0.81)0.87 (0.78 to 0.97)−0.13 (−0.23 to −0.03)0.93 (0.87 to 1.00)0.11 (0.10 to 0.12)23 (19 to 27)0.48Functional-immunological iBox39410.80 (0.78 to 0.82)0.85 (0.76 to 0.94)−0.14 (−0.25 to −0.04)0.93 (0.87 to 0.99)0.11 (0.10 to 0.11)25 (21 to 29)0.50Functional-histological iBox39410.80 (0.78 to 0.82)0.86 (0.78 to 0.95)−0.12 (−0.22 to −0.01)0.94 (0.88 to 1.01)0.10 (0.10 to 0.11)25 (21 to 29)0.51Full iBox with biopsy diagnostics39980.81 (0.79 to 0.83)0.84 (0.76 to 0.93)−0.13 (−0.24 to −0.03)0.93 (0.87 to 1.00)0.10 (0.09 to 0.11)26 (22 to 30)0.52Full iBox with dipstick39410.81 (0.79 to 0.83)0.85 (0.77 to 0.94)−0.13 (−0.24 to −0.03)0.93 (0.87 to 1.00)0.10 (0.09 to 0.11)26 (23 to 30)0.52Functional iBox with dipstick39410.79 (0.77 to 0.81)0.86 (0.77 to 0.95)−0.14 (−0.24 to −0.04)0.93 (0.87 to 1.00)0.11 (0.10 to 0.12)23 (19 to 27)0.49Discrimination, calibration, and overall accuracy performance metrics for the different iBox algorithms were assessed seven years after the risk evaluation. *Full iBox (estimated glomerular filtration rate (eGFR), continuous proteinuria, anti-HLA donor specific antibody mean fluorescence intensity, and histological lesions); functional iBox (eGFR and continuous proteinuria);functional-immunological iBox (eGFR, continuous proteinuria, and anti-HLA donor specific antibody mean fluorescence intensity); functional-histological iBox (eGFR, continuous proteinuria, and histological lesions); full iBox with biopsy diagnostics (eGFR, continuous proteinuria, anti-HLA donor specific antibody mean fluorescence intensity, and Banff diagnoses); full iBox with dipstick (eGFR, dipstick proteinuria, anti-HLA donor specific antibody mean fluorescence intensity, and histological lesions); and functional iBox with dipstick (eGFR and dipstick proteinuria).CI, confidence interval.Table 3Predictive performance (discrimination, calibration, and overall fit) of the different integrative Box (iBox) algorithms in the external validation cohortsiBox algorithm*DiscriminationCalibrationOverall fitHarrell's C index (95% CI)Slope (95% CI)Intercept (95% CI)Observed/expected ratio (95% CI)Brier score (95% CI)Index of predictive accuracy (%) (95% CI)Royston's R2DEuropean cohort Full iBox (n=3393)0.82 (0.79 to 0.84)0.73 (0.63 to 0.84)−0.37 (−0.51 to −0.22)0.78 (0.71 to 0.86)0.09 (0.08 to 0.09)11 (5 to 17)0.45 Functional iBox (n=5348)0.80 (0.78 to 0.82)0.82 (0.72 to 0.92)−0.39 (−0.50 to −0.27)0.75 (0.70 to 0.82)0.08 (0.08 to 0.09)10 (6 to 15)0.43 Functional-immunological iBox (n=3467)0.80 (0.78 to 0.83)0.72 (0.61 to 0.82)−0.44 (−0.58 to −0.30)0.76 (0.69 to 0.84)0.09 (0.08 to 0.10)8 (2 to 15)0.41 Functional-histological iBox (n=3395)0.81 (0.79 to 0.84)0.78 (0.67 to 0.89)−0.39 (−0.53 to −0.25)0.77 (0.70 to 0.85)0.08 (0.08 to 0.09)12 (5 to 18)0.45 Full iBox with dipstick (n=3393)0.80 (0.78 to 0.83)0.71 (0.60 to 0.82)−0.28 (−0.42 to −0.14)0.85 (0.77 to 0.93)0.09 (0.08 to 0.10)9 (3 to 15)0.4 Functional iBox with dipstick (n=5348)0.79 (0.77 to 0.81)0.75 (0.66 to 0.85)−0.27 (−0.39 to −0.15)0.86 (0.79 to 0.93)0.08 (0.08 to 0.09)8 (4 to 12)0.37North American cohort Full iBox (n=1956)0.83 (0.81 to 0.86)0.79 (0.67 to 0.91)0 (−0.17 to 0.17)0.99 (0.89 to 1.10)0.08 (0.07 to 0.09)31 (25 to 37)0.55 Functional iBox (n=2500)0.83 (0.81 to 0.86)0.85 (0.72 to 0.99)0.1 (−0.07 to 0.27)1.08 (0.98 to 1.19)0.09 (0.08 to 0.10)27 (21 to 33)0.53 Functional-immunological iBox (n=2274)0.83 (0.81 to 0.86)0.83 (0.70 to 0.96)−0.03 (−0.20 to 0.15)0.98 (0.89 to 1.09)0.08 (0.07 to 0.09)30 (23 to 36)0.54 Functional-histological iBox (n=1963)0.83 (0.81 to 0.86)0.77 (0.65 to 0.89)0.05 (−0.12 to 0.22)1.03 (0.93 to 1.15)0.09 (0.07 to 0.10)29 (23 to 35)0.55 Full iBox with dipstick (n=1956)0.83 (0.81 to 0.86)0.82 (0.70 to 0.95)−0.01 (−0.18 to 0.16)0.97 (0.87 to 1.08)0.08 (0.07 to 0.09)31 (25 to 37)0.55 Functional iBox with dipstick (n=2500)0.83 (0.81 to 0.86)0.87 (0.73 to 1.01)0.07 (−0.10 to 0.24)1.05 (0.95 to 1.16)0.09 (0.08 to 0.10)26 (21 to 32)0.53South American cohort Full iBox (n=665)0.88 (0.85 to 0.91)1.52 (1.21 to 1.84)0.41 (0.22 to 0.60)1.20 (1.01 to 1.41)0.09 (0.07 to 0.11)37 (31 to 43)0.66 Functional iBox (n=665)0.86 (0.83 to 0.89)1.49 (1.17 to 1.82)0.17 (−0.01 to 0.36)1.01 (0.86 to 1.19)0.09 (0.08 to 0.11)36 (30 to 42)0.63 Functional-immunological iBox (n=665)0.87 (0.83 to 0.89)1.42 (1.12 to 1.71)0.25 (0.06 to 0.44)1.08 (0.92 to 1.28)0.09 (0.08 to 0.11)36 (30 to 42)0.63 Functional-histological iBox (n=665)0.88 (0.85 to 0.91)1.59 (1.24 to 1.94)0.39 (0.21 to 0.58)1.18 (1.00 to 1.39)0.09 (0.08 to 0.11)37 (31 to 42)0.67 Full iBox with dipstick (n=665)0.88 (0.85 to 0.91)1.47 (1.17 to 1.77)0.38 (0.19 to 0.56)1.17 (0.99 to 1.38)0.09 (0.08 to 0.11)36 (30 to 42)0.66 Functional iBox with dipstick (n=665)0.87 (0.83 to 0.89)1.37 (1.09 to 1.66)0.15 (−0.04 to 0.34)0.99 (0.84 to 1.17)0.09 (0.08 to 0.11)34 (28 to 41)0.65Discrimination, calibration, and overall accuracy performance metrics for the different iBox algorithms were assessed five years after the risk evaluation in the European and North American validation cohorts, and 2.5 years after the risk evaluation in the South American external validation cohort. *Full iBox (estimated glomerular filtration rate (eGFR), continuous proteinuria, anti-HLA donor specific antibody mean fluorescence intensity, and histological lesions); functional iBox (eGFR and continuous proteinuria);functional-immunological iBox (eGFR, continuous proteinuria, and anti-HLA donor specific antibody mean fluorescence intensity); functional-histological iBox (eGFR, continuous proteinuria, and histological lesions); full iBox with dipstick (eGFR, dipstick proteinuria, anti-HLA donor specific antibody mean fluorescence intensity, and histological lesions); and functional iBox with dipstick (eGFR and dipstick proteinuria).CI, confidence interval.Figure 2Calibration of the full integrative Box (iBox) in the derivation and validation cohorts. Calibration plots are presented for seven years after the risk evaluation in the derivation cohort, five years after the risk evaluation in the European and North American validation cohorts, and at 2.5 years after the risk evaluation in the South American validation cohort. In each group (eight groups for the derivation, European, and North American cohorts, and five groups for the South American cohort), the median of the predicted risks was plotted against the observed event probability estimated by (1−) the Kaplan-Meier estimator. Flexible calibration curves were derived from pseudo-observations plotted against predicted risks and smoothed with locally estimated scatterplot smoothing regression (span=0.75). The diagonal line at the origin represents the perfectly calibrated model. The histograms represent the distribution of the predicted risks for the individual models according to event statusFigure 3Decision curve analysis of the full integrative Box (iBox) in the derivation and validation cohorts. Decision curves are presented for seven years after the risk evaluation in the derivation cohort, five years after the risk evaluation in the European and North American validation cohorts, and 2.5 years after the risk evaluation in the South American validation cohort. Net benefit is plotted against threshold probabilities for predicted risk of allograft failure, ranging from 0 to 0.4, along with treat all and treat none strategiesIn the derivation cohort, the full iBox achieved a C index of 0.81 (95% confidence interval (CI) 0.79 to 0.83) (table 2). Although all of the iBox algorithms showed good discrimination (Harrell's c index range 0.79 to 0.81), discrimination was significantly improved with the original iBox algorithm compared with the functional iBox (P<0.001), functional-immunological iBox (P<0.001), functional-histological iBox (P=0.005), and functional iBox with dipstick (P<0.001). Online supplemental table 6 shows all pairwise concordance comparisons between the iBox algorithms. All iBox algorithms showed close and good agreement between the predicted and observed risks in the derivation cohort, as reflected in the calibration plot and the calibration metrics (table 2, figure 2, and online supplemental figure 2).For overall fit, we found a moderate but significant improvement for the Brier score of the full iBox model compared with the functional iBox (P<0.001), functional-immunological iBox (P=0.006), functional-histological iBox (P=0.004), and functional iBox with dipstick (P<0.001; online supplemental table 7 has all pairwise comparisons). Clinical utility, assessed with decision curves, showed that the different iBox algorithms had positive and similar net benefits across decision thresholds up to 40% (figure 3 and online supplemental figure 5). Specifically, at a threshold of 20%, net benefit was 0.07 for the full iBox model, functional iBox, functional-immunological iBox, functional-histological iBox, full iBox with dipstick, and functional iBox with dipstick, and 0.08 for the full iBox with biopsy diagnostics. In the external validation cohorts, all iBox algorithms showed good discrimination, with the C index ranging from 0.79 to 0.88 (0.82, 95% CI 0.79 to 0.84 and 0.80, 0.78 to 0.82 in the European cohort, 0.83, 0.81 to 0.86 and 0.83, 0.81 to 0.86 in the North American cohort, and 0.88, 0.85 to 0.91 and 0.86, 0.83 to 0.89 in the South American cohort for the iBox and functional iBox, respectively) (table 3).In terms of calibration, agreement between the predicted and observed risks for the iBox algorithms was adequate in the North American validation cohort five years after the risk evaluation (figure 2 and online supplemental figure 3). In the European validation cohort, all models tended to overestimate the risks, as reflected in the calibration slopes of <1 (0.73, 95% CI 0.63 to 0.84 and 0.82, 0.72 to 0.92 for the full iBox and functional iBox, respectively), and the negative calibration intercepts (−0.37, 95% CI −0.51 to −0.22 and −0.39, −0.50 to −0.27 for the iBox and functional iBox, respectively; table 3, figure 2, and online supplemental figure 2). In the South American validation cohort, the predicted risks were underestimated, as reflected by positive calibration intercepts (0.41, 95% CI 0.22 to 0.60 and 0.17, −0.01 to 0.36 for the full iBox and functional iBox, respectively) and calibration slopes >1 (1.52, 95% CI 1.21 to 1.84 and 1.49, 1.17 to 1.82 for the iBox and functional iBox, respectively; table 3, figure 2, and online supplemental figure 4).For clinical utility, all iBox algorithm versions showed comparable decision curves with positive net benefit across threshold probabilities from 0 to 0.4 within each validation cohort (net benefit at a 20% threshold of 0.03, 0.06-0.08, and 0.10-0.11 across algorithms in the European, North American, and South American validation cohorts, respectively; figure 3 and online supplemental figures 5-8).Predictive performance of the various iBox algorithms when accounting for competing risks To account for the competing risk of death with a functioning graft, the different iBox algorithm versions were modelled with Fine-Gray subdistribution hazards models. Online supplemental tables 8-10 show the predictive performance of these models. Online supplemental figures 9-18 show the calibration plots and decision curves, and online supplemental tables 11 and 12 show the net benefit at the 20% threshold. The Fine-Gray models gave similar or slightly worse discrimination, regardless of the cohort or iBox algorithm version (C index of 0.80, 0.79, 0.82, and 0.88 for the Fine-Gray models compared with 0.81, 0.80, 0.82, and 0.88, for the full iBox in the derivation, European, North American, and South American cohorts, respectively; tables 2 and 3, and online supplemental tables 8-10).Calibration improved in the derivation cohort and European external validation cohort when we used Fine-Grey models for all algorithm versions, with calibration slopes of 0.90 (95% CI 0.81 to 0.99) and 0.74 (0.62 to 0.86) in the derivation and European cohorts, respectively, compared with 0.86 (0.77 to 0.94) and 0.70 (0.59 to 0.82) for the full iBox (online supplemental tables 9 and 10). In the South American validation cohort, however, miscalibration increased for all algorithm versions when we used Fine-Gray models (calibration slope of 1.64, 95% CI 1.31 to 1.98 compared with 1.52, 1.21 to 1.84 for the full iBox; table 3 and online supplemental table 9). In the North American validation cohort, we saw no systematic improvement in calibration with either modelling approach across algorithm versions (online supplemental tables 9 and 10).Overall fit varied across metrics, with similar Brier scores for both iBox and Fine-Gray models, consistently improved values for the index of predictive accuracy for the iBox algorithms, and slightly improved Royston's R2D values for the Fine-Grey models, across all cohorts and algorithm versions (tables 2 and 3 and online supplemental tables 8-10). In terms of clinical utility, net benefit was slightly worse across the range of threshold probabilities when we used Fine-Gray models, regardless of the cohort or algorithm version (net benefit at the 20% threshold of 0.06, 0.03, 0.07, and 0.10 for the Fine-Gray models compared with 0.07, 0.04, 0.08, and 0.11, for the full iBox in the derivation, European, North American, and South American cohorts, respectively; online supplemental tables 11 and 12).Predictive performance of the iBox across multiple validated eGFR equations Mean eGFR was 48.18 mL/min/1.73 m² (SD 19.36) when estimated with the MDRD 186 equation, 45.28 (18.21) with the MDRD175 equation, 48.32 (20.6) with the CKD-EPI 2009 equation, 51.11 (21.3) with the CKD-EPI 2021 equation, 48.48 (SD 19.51) with the European Kidney Function Consortium equation, and 49.98 (15.78) with the race-free kidney recipient specific equation (online supplemental figure 19). All equations had similar discrimination and overall fit, with a C index of 0.81 (CI 0.79 to 0.83) and a Brier score of 0.10 (CI 0.09 to 0.11) (table 4). Calibration and clinical utility showed comparable results across equations (online supplemental figures 20 and 21).Table 4Predictive performance of the integrative Box (iBox) according to different estimated glomerular filtration rate (eGFR) equationseGFR equationNo of patientsDiscriminationCalibrationOverall fitHarrell's C index (95% CI)Slope (95% CI)Intercept (95% CI)Observed/expected ratio (95% CI)Brier score (95% CI)Index of predictive accuracy (%) (95% CI)Royston's R2DMDRD186*39410.81 (0.79 to 0.83)0.86 (0.77 to 0.94)−0.12 (−0.23 to −0.02)0.93 (0.87 to 1.00)0.10 (0.09 to 0.11)26 (23 to 30)0.52MDRD17539410.81 (0.79 to 0.83)0.88 (0.79 to 0.97)−0.21 (−0.31 to −0.11)0.87 (0.81 to 0.93)0.10 (0.09 to 0.11)27 (23 to 30)0.52CKD-EPI 200939410.81 (0.79 to 0.83)0.84 (0.75 to 0.92)−0.15 (−0.26 to −0.05)0.93(0.87 to 0.99)0.10 (0.09 to 0.11)26 (23 to 30)0.52CKD-EPI 202139410.81 (0.79 to 0.83)0.82 (0.74 to 0.90)−0.06 (−0.17 to 0.04)1.00 (0.93 to 1.07)0.10 (0.09 to 0.11)26 (22 to 30)0.52EKFC39410.81 (0.79 to 0.83)0.85 (0.76 to 0.93)−0.12 (−0.22 to −0.02)0.94 (0.88 to 1.01)0.10 (0.09 to 0.11)26 (23 to 30)0.51KRS39410.81 (0.79 to 0.83)0.89 (0.80 to 0.98)0.02 (−0.08 to 0.12)1.01 (0.95 to 1.09)0.10 (0.09 to 0.11)26 (22 to 29)0.52Discrimination, calibration, and overall accuracy performance metrics for the different iBox algorithms were assessed seven years after the risk evaluation in the derivation cohort. *Because the MDRD186 equation was used in the original iBox development, its performance metrics serve as the reference values.CKD-EPI, Chronic Kidney Disease-Epidemiology Collaboration; EKFC, European Kidney Function Consortium; KRS, kidney recipient specific; MDRD, Modification of Diet in Renal Disease Study.iBox performance across diverse subpopulations and clinical scenarios In the derivation cohort, 127 of 3941 (3.2%) patients had recurrence of the initial nephropathy, 175 patients (4.4%) had BK virus associated nephropathy, 3658 patients (92.8%) were treated with calcineurin inhibitors, 239 patients (6.1%) were treated with mTOR inhibitors, 1160 (29.4%) patients had a protocol biopsy, and 2781 patients (70.6%) had an indication biopsy. The iBox prediction ability was accurate in these scenarios, with a C index of 0.79 (95% CI 0.71 to 0.85), 0.74 (0.66 to 0.80), 0.81 (0.79 to 0.83), 0.87 (0.79 to 0.93), 0.80 (0.77 to 0.82), and 0.81 (0.76 to 0.85) for recurrence, BK virus associated nephropathy, calcineurin inhibitor subpopulation, mTOR subpopulation, indication biopsies, and protocol biopsies, respectively ( online supplemental table 13). Online supplemental figures 22 and 23 show the calibration plot and decision curve for the calcineurin inhibitor subpopulation. The small number of patients in the scenarios of recurrence of initial nephropathy, BK virus associated nephropathy, and mTOR inhibitor treatment did not allow for an assessment of calibration and clinical utility.iBox performance with histological diagnoses and histological indexes rather than individual Banff lesion scores When we used Banff 2022 diagnoses instead of Banff lesion scores in the iBox, the following independent determinants were identified, in addition to eGFR, proteinuria, and anti-HLA donor specific antibody mean fluorescence intensity: antibody mediated rejection/microvascular inflammation (hazard ratio 1.45 95% CI 1.19 to 1.77, P<0.001), mixed rejection (hazard ratio 1.94, 1.34 to 2.81, P<0.001), and recurrence of the initial nephropathy (hazard ratio 1.69, 1.21 to 2.37, P=0.002). The model maintained good discrimination and calibration in the derivation cohort with a C index of 0.81 (CI 0.79 to 0.83) ( table 2 and online supplemental figures 1 and 5).Replacing Banff lesion scores with dichotomised indexes (activity, chronicity, or both combined) did not modify the performance of the iBox in the derivation cohort, with the C index remaining stable at 0.82 (95% CI 0.80 to 0.84) for all versions except chronicity index (0.81, 0.79 to 0.83) (online supplemental table 14 and online supplemental figure 24). We saw similar results with continuous histological indexes (activity, chronicity, antibody mediated rejection/microvascular inflammation, and T cell mediated rejection/tubulointerstitial inflammation; online supplemental table 15 and online supplemental figure 25).iBox versus eGFR slope and versus circulating anti-HLA donor specific antibodies We compared the performances of the iBox computed at one year after transplantation with those of eGFR slope, comprising two creatinine measurements between three months and one year after transplantation. In the derivation cohort, 524 patients met this criterion. Median eGFR slope was −2.39 (IQR −11.00 to 5.42). Discrimination of the iBox in this subpopulation was 0.86 (95% CI 0.82 to 0.90), whereas discrimination of the Cox model with the eGFR slope and the timing of risk evaluation was 0.68 (0.61 to 0.74, P<0.001; online supplemental table 16).The iBox was also compared with circulating anti-HLA donor specific antibodies assessed at the time of risk evaluation for predicting graft failure. The C index of the Cox model with the anti-HLA donor specific antibody mean fluorescence intensity and the timing of risk evaluation was 0.57 (95% CI 0.54 to 0.59, P<0.001).Extension of iBox predictions up to 10 years after risk evaluation We extended follow-up of patients to a maximum of 10 years after the risk evaluation in the derivation cohort. In this setting, median follow-up time was 9.36 years (IQR 5.22-10.00) and we found 587 (19%) graft losses. The iBox achieved a C index of 0.79 (95% CI 0.77 to 0.81) and maintained good calibration ( online supplemental table 17 and online supplemental figure 26).Discussion Principal findings In this study, we found that the iBox system for prediction of kidney allograft failure maintained good discrimination and overall fit, and demonstrated clinical utility, when used in abbreviated forms adapted for centres without routine histological or immunological assessment. Calibration was adequate in some but not all of the external validation cohorts, with slight overestimation or underestimation of predicted risks. The iBox consistently performed well with different eGFR equations, and in various clinical scenarios, including recurrence of the initial nephropathy and in different immunosuppressive strategies. Furthermore, the iBox maintained good performances when dipstick proteinuria was used instead of continuous measurements, which further affirms its versatility and value as a surrogate endpoint in various transplantation settings.Different versions of the iBox to adapt to centres’ practices The development of multiple iBox algorithms addressed the heterogeneity in practices in transplant centres worldwide. In centres where routine biopsies are rarely or not performed, we validated two abbreviated scores that showed good prediction performances: the functional iBox, based on eGFR, proteinuria, and timing of assessment, and the functional-immunological iBox, combining functional parameters with anti-HLA donor specific antibodies. For centres where immunological parameters are not routinely assessed, the functional iBox can be used, or the functional-histological iBox that we specifically developed for this scenario, which incorporates biopsy findings. Both algorithms maintained good prediction performances.Furthermore, recognising that some centres may not grade histological lesions according to the Banff classification or do not have these data available, we developed the diagnostic iBox with Banff diagnoses instead of histological lesions scoring. This algorithm showed good prediction performances while accommodating different histological assessment practices.Dipstick proteinuria versus continuous proteinuria Adaptation of the iBox to use dipstick proteinuria instead of continuous proteinuria measurements represents a major advancement in its practical applicability. Although previous studies 23 suggested that urine protein-to-creatinine ratio testing might have advantages over dipstick testing alone in predicting kidney failure, our findings showed that both methods gave comparable performances in the iBox model. This validation enables centres that use only dipstick measurements to implement the iBox effectively.iBox versus eGFR slope and versus donor specific antibodies eGFR slope has been used as a surrogate endpoint for randomised controlled trials, particularly in chronic kidney disease affecting native kidneys. 39–41 Although eGFR slope is useful in this population, the transplant setting offers the possibility of a more comprehensive risk assessment because regular follow-up in transplant centres enables the integration of multiple parameters. This integration, as captured in the iBox, provides a more complete risk prediction tool than eGFR slope alone. Our findings showed the superior predictive ability of the iBox over eGFR slope (discrimination of 0.86 v 0.68, P<0.001). Similarly, we showed the superior predictive ability of the iBox over circulating anti-HLA donor specific antibodies (discrimination of 0.81 v 0.57, P<0.001). We also showed that the use of histological indexes, capturing activity and chronicity, did not modify the performances of the model compared with the use of Banff lesion scores.Competing risks We also investigated the performance of the different iBox algorithms when accounting for the competing risk of death with a functioning graft. In the derivation and validation cohorts, the iBox with and without competing risk modelling showed comparable predictive performance across discrimination, overall fit, and clinical utility metrics. In terms of calibration, however, the results varied across cohorts, with no modelling approach being systematically better than any other. In our study, the rates of death with a functioning graft were lower than the rates of graft loss in all cohorts. These findings support the robustness of the current iBox approach.Difficulty of validating a prediction model Validating a prediction model is a challenging process. 42 21 43 Studies attempting to validate a prediction model often lack comprehensive assessment of performance, with many limiting their evaluation to one discrimination metric,44 without assessment in external cohorts45 or diverse clinical scenarios. In our study, we used several key performance metrics (discrimination, calibration, overall fit, and clinical utility) and conducted external validation across multiple cohorts from many centres spanning different continents.Our study complements a longstanding effort by our group to thoroughly investigate the performance of the iBox in various healthcare systems and practices. The iBox is now one of the most validated prediction systems in nephrology: the iBox can predict the findings of randomised controlled trials,11 12 outperform clinicians in predicting allograft failure,13 perform similarly or better than machine learning models,14 outperform eGFR slope, perform well in a large series of subpopulations, and can now adapt to specific centre practices.iBox as a tool to improve patient activation The iBox prediction system may also be useful as a tool for promoting patient engagement and shared decision making. By providing a probabilistic and personalised estimation of allograft longevity, the iBox can help patients better understand their transplant trajectory, anticipate future scenarios, and take an active role in planning personal, medical, or professional decisions. This information could help with, for example, decisions about parenthood, lifestyle adaptations, treatment changes, or even long term financial and geographical planning. In this sense, the iBox may contribute to improving patient activation, which is a determinant of better health outcomes. 45Sharing the iBox prediction The decision for a patient to access their iBox prediction must remain theirs, following an informed discussion with their physician. We advocate that iBox prediction should never be communicated in isolation, but always in a clinical context and interpreted jointly by the physician and patient. Prediction is not absolute truth, and can be affected by transient factors, such as infection, dehydration, or acute events.Limitations of this study We acknowledge several limitations of our study. Firstly, although the iBox was developed on a prospective derivation cohort, performances were assessed in retrospective validation cohorts. Prospective validation remains desirable. Secondly, the model incorporating dipstick proteinuria was tested in three European centres. Future studies could aim to evaluate the performance of the iBox in centres from other continents. Thirdly, non-invasive biomarkers associated with allograft failure are increasingly published. The iBox could be used in future research investigating the predictive power of these biomarkers when adjusted for the routine clinical parameters in the iBox.Conclusions The iBox model showed robust discrimination, overall fit, and clinical utility across diverse real world settings and transplant systems worldwide. The successful validation of the iBox in multiple abbreviated algorithms adapted to specific clinical scenarios supports its broad implementation in both clinical practice and trials. This comprehensive validation establishes the iBox as a versatile and reliable tool for predicting the long term outcomes of kidney allografts.