BetaEntity Annotation Prototype
← Back to diseases

Annotated full text

Interpretable machine learning model for predicting recurrence in patients with diabetic foot ulcers

bmjdrc · 2025-11-12 · canonical JSON source

39 visible annotations · policy: published · automated confidence ≥ 75.00%

Document resource

WHAT IS ALREADY KNOWN ON THIS TOPIC Diabetic foot ulcer (DFU) is a serious diabetes mellitus complication with a chronic course and high recurrence, complicating patient management. Accurate DFU recurrence prediction is crucial for improving patient care through timely interventions. Before this study, there were few machine learning (ML) models for predicting 3-year DFU recurrence risk, as existing methods often lacked integrated feature selection and thorough algorithm comparisons, highlighting the need for this research.WHAT THIS STUDY ADDS This study offers new insights by including 494 with DFU, splitting them into training and test sets, and using four feature selection methods to create a robust predictor set. It assessed seven ML algorithms for predicting 3-year DFU recurrence, optimized them via cross-validation and identified the best model. This model was calibrated with Platt scaling and enhanced with SHapley Additive exPlanations analysis, resulting in an accurate and interpretable DFU recurrence prediction tool.HOW THIS STUDY MIGHT AFFECT RESEARCH, PRACTICE OR POLICY This study offers a framework for using multiple feature selection methods and comparing ML algorithms for predicting DFU recurrence, laying the groundwork for future research. Clinically, the interpretable ML model aids in assessing the 3-year DFU recurrence risk, facilitating personalized treatment and potentially reducing recurrence rates and healthcare costs. Policy-wise, the study’s findings could encourage policymakers to support the integration of ML predictive tools in DFU management, including funding for training and updating clinical guidelines.Introduction Diabetic foot ulcer (DFU) and peripheral vascular disease are serious and common complications of diabetes mellitus (DM) that greatly affect patients by increasing pain, limiting mobility and reducing overall quality of life. 1–3 Global epidemiological studies estimate the prevalence of DFU at 6.3%, with recurrence rates presenting a critical challenge.4 Following successful initial healing, DFUs exhibit a tendency for recurrence, with approximately 40% of cases recurring within 1 year and 60% within 3 years posthealing.5 This recurring cycle of ulceration not only worsens patient morbidity but also raises the risk of serious complications, such as infection and amputation. Additionally, DFU is a major cause of non-traumatic lower limb amputations, contributing to nearly 85% of these cases. Following an initial amputation, the likelihood of reamputation within 5 years reaches 50%, with associated mortality rates nearing 70%.6 The medical, economic and social burden of DFU is substantial, placing significant strain on healthcare systems worldwide. In developed countries, DFU-related healthcare expenditures comprise nearly 20% of public health resources annually, with potentially even higher proportions in developing nations.7–9 For instance, a comparative cost analysis across multiple countries found that the annual treatment expenses per DFU patient are approximately US$3000 in the United States, Sweden and the Netherlands.10–12 In contrast, costs in China vary significantly, ranging from $1700 for mild cases to $21 000 for severe cases.13 These challenges underscore the urgent need for effective predictive strategies to mitigate recurrence risk and improve patient outcomes and quality of life.Accurately predicting DFU recurrence risk and identifying patient-specific risk factors are critical for facilitating early intervention in high-risk individuals.14 Proactive risk stratification can help reduce recurrence rates and decrease the incidence of amputations, ultimately improving patient outcomes. Recent studies have emphasized the growing role of machine learning (ML) models in clinical applications, demonstrating their potential in enhancing predictive accuracy and clinical decision-making support.15 16 Unlike traditional scoring systems such as the Diabetic Ulcer Severity Score and the Meggitt-Wagner Classification, ML models offer greater scalability, accommodating a broader range of predictive variables while capturing both linear and non-linear relationships.17 Notably, we have previously conducted several studies focusing on DFU, including the application of ML to predict in-hospital amputation outcomes in patients with DFU, which has been endorsed by the International Working Group on the Diabetic Foot (IWGDF) guidelines.3 18 19 As an advanced artificial intelligence technology, ML has been widely applied across various medical fields due to its ability to process complex datasets, identify intricate patterns and improve predictive accuracy. This advanced capability improves the precision of risk assessment and decision-making in clinical practice.20 21 This study aimed to develop ML-based predictive models for assessing the 3-year recurrence risk in patients with DFU. By incorporating multidimensional clinical features, we evaluated model performance and clinical applicability to investigate the effectiveness of feature utilization. Our goal is to develop data-driven decision support tools that empower clinicians to enhance personalized patient management and improve long-term outcomes.Materials and methods Study design and participants This retrospective cohort study was conducted across three hospitals in Southwest China from 2016 to 2022. Each hospital created its own database and handled its own data entry, verification and storage. Quarterly, hospitals submitted standardized data to Chongqing University’s Affiliated Hospital, maintaining their own data management authority. In collaboration with clinical specialists in endocrinology and wound care, we established standardized inclusion and exclusion criteria to ensure data integrity and consistency within the cohort. Exclusion criteria included patients under 18 years of age, those who had undergone bilateral leg amputation prior to hospital admission, those with unhealed ulcers at discharge, clinical data with more than 50% missing values, lack of complete follow-up information or those with concurrent malignant tumors or critical conditions. An overview of the study workflow is depicted in figure 1.Figure 1The overall flowchart of the study (A). The algorithm chart of the study (B). MRNs, Medical Record Numbers; BP, blood pressure; DFU, diabetic foot ulcers; GBDT, gradient boosting decision tree; HbA1c, glycated hemoglobin; HR, heart rate; LASSO, least absolute shrinkage and selection operator; LR, logistic regression; MRMR, minimum redundancy maximum relevance; RFE, recursive feature elimination; RR, respiratory rate; SVM, support vector machine.Patient data were extracted from hospital electronic health records, including demographic information, neighborhood characteristics, clinical and laboratory parameters. The Wagner ulcer classification, Ischemia and foot Infection classification were used to assess ulcer severity and vascular status, respectively. Clinical outcomes at the time of hospitalization were also recorded and extracted. After applying rigorous inclusion and exclusion criteria, a total of 494 patients were included in this study. For variables with <50% missing values, we excluded those with a missing rate of 30% or more to prevent bias. For variables with <30% missing data, we used k-nearest neighbor imputation with k=4, based on Euclidean distance. This involved estimating missing values using the average of the four most similar samples. The choice of k=4 was optimized by testing k values from 1 to 10, using the area under the curve (AUC) of a logistic regression (LR) model to ensure reliable imputation. To ensure model training consistency by eliminating variable dimension differences, we applied Z-Score standardization. This method transforms data using the formula  xnorm=x−x-σ , where x  is the original value, x- is the mean and σ is the SD. The resulting standardized data have a mean of 0 and an SD of 1, enhancing data comparability across centers.22 23 Patients were stratified into two groups based on DFU recurrence status. The recurrent group included patients who experienced ulcer recurrence following initial treatment and recovery during the study period. The non-recurrent group comprised patients who remained ulcer-free after treatment and recovery throughout the study period. Baseline characteristics of patients in the recurrence group and non-recurrence group are shown in online supplemental table 2.SP310.1136/bmjdrc-2025-005242.supp3Supplementary data This study is a retrospective observational analysis using anonymized data from routine clinical records. All identifying information was removed to ensure participant privacy, eliminating any risk of identity disclosure. As there was no direct participant contact or intervention, the study posed no impact on their well-being.Statistical analysis We performed statistical analyses using IBM SPSS Statistics V.27, (IBM, Armonk, New York). Continuous variables with a normal distribution were analyzed using a t-test or analysis of variance and reported as mean±SD. Non-normally distributed variables were assessed using the Wilcoxon rank-sum test and presented as median (IQR). Categorical variables were analyzed with the χ 2 test and reported as counts and percentages. P<0.05 was considered statistically significant.Model development To identify the most relevant predictive features, we applied four feature selection techniques— least absolute shrinkage and selection operator, minimum redundancy maximum relevance, Fisher and recursive feature elimination—on the training dataset. 24–27 The final feature set consisted of the 24 variables shared across all four methods. The preprocessed data were divided into training and testing sets using train-test-split, followed by fivefold cross-validation. The dataset was randomly split into five equal subsets, with each subset serving as the test set once, while the others formed the training set. This process was repeated five times to ensure each subset was used as the test set once. The training set included 395 patients (257 recurrent and 138 non-recurrent), while the test set contained 99 patients (61 recurrent, 38 non-recurrent).This study employed seven ML models to develop predictive models for assessing the 3-year recurrence risk of DFU: LR, support vector machine (SVM), random forest (RF), gradient boosting decision trees (GBDT), AdaBoost, extreme gradient boosting (XGBoost) and light gradient boosting machine (LightGBM).28–32 Hyperparameter tuning was performed on all algorithms using grid search to ensure that the parameter combinations achieved optimal generalization performance on the validation set.33 Model performance was evaluated using multiple metrics, including area under the receiver operating characteristic curve (AUROC), accuracy, recall, specificity, negative predictive value (NPV) and positive predictive value (PPV). The model with the highest overall predictive performance was selected for further refinement. To improve calibration, Platt scaling was applied to the optimal model, with its calibration effectiveness assessed using the Brier score. Additionally, the SHapley Additive exPlanations (SHAP) method was used to interpret the contributions of individual features to the final model predictions.All model development, validation and interpretation were conducted using standard Python libraries (Python V.3.8.2). https://www.python.org/downloads/release/python-3820/. Source code, processed datasets and trained model weights are available in online supplementary3.SP410.1136/bmjdrc-2025-005242.supp4Supplementary data Results Characteristics of participants A total of 494 DFU patients were recruited from three hospitals. The characteristics of patients in the training set and test set are summarized in table 1. Within the training cohort, the median age was 70.25 years (IQR 69.09–71.41), and 271 (69.1%) participants were male, and recurrence was observed in 243 (63.0%) patients. For external validation, models were tested on a separate dataset of 98 DFU patients from three independent hospitals. In the validation cohort, the median age was 69.73 years (IQR 67.57–71.90), with 69 (70.4%) male participants, and recurrence was observed in 67 (68.4%) patients. The characteristics across training set and validation set are statistically comparable in terms of their feature distributions(online supplemental file 2).SP210.1136/bmjdrc-2025-005242.supp2Supplementary data Table 1Baseline characteristics of patients with DFU in the training and test setsVariablesTraining set(n=392)Validation set(n=98)Demographic data  Age, years70.25 (69.09, 71.41)69.73 (67.57, 71.90)  Sex, %  Male271 (69.1)69 (70.4)  Female121 (30.9)29 (29.6)  Body mass index, kg/m2 22.62 (22.29, 23.96)22.88 (22.24, 23.53)  Hospital duration, days27.00 (24.53, 29.49)29.40 (18.94, 39.86)  SI, %  <800 units215 (54.8)47 (47.9)  >800 units177 (45.2)51 (46.9)Medical history  Infection, n (%)152 (38.8)29 (29.6)  Hypertension, n (%)199 (50.8)43 (43.9)  Coronary heart disease, n (%)265 (67.6)75 (76.5)  Cerebral infarction, n (%)128 (32.7)35 (35.7)  Diabetic nephropathy, n (%)207 (52.8)49 (50)  Diabetic neuropathy, n (%)343 (87.5)91 (92.9)  Ischemia, n (%)248 (63.3)61 (62.2)  Diabetic retinopathy196 (50)50 (51)  Tumor, n (%)20 (5.1)8 (8.2)  Recurrence, n (%)247 (63.0)67 (68.4)Wagner  0–2, n (%)146 (37.2)36 (36.7)  3–5, n (%)246 (62.8)62 (63.3)SI = smoking years ×number of cigarettes smoked per day; <800 units, the number of years a person has smoked and the number of cigarettes smoked per day is under 800; >800 units, the product of the number of years a person has smoked and the number of cigarettes smoked per day exceeds 800.DFU, diabetic foot ulcer; SI, smoking index.Model performance The ROC curves and the corresponding AUC values for the seven evaluated models are depicted in figure 2A. Additional performance metrics—including accuracy, sensitivity, specificity, NPV and PPV—are detailed in table 2. The results reveal that the XGBoost model attained the highest AUC of 0.924 (95% CI 0.867 to 0.967). In contrast, the RF model exhibited greater stability, with an AUC of 0.918 (95% CI 0.867 to 0.960), and achieved the highest sensitivity and NPV. The XGBoost model surpassed all other models in terms of accuracy, specificity and PPV, while its sensitivity and NPV were comparable to those of the RF model. Although the RF model excelled in sensitivity, its lower specificity limited its capacity to accurately identify non-recurrent cases. Moreover, the calibration curve for the XGBoost model closely approximated the ideal 45° line, with a low Brier score of 0.096, indicating excellent probability calibration (figure 2B).Figure 2Discrimination and calibration performance of the models. (A) Receiver operating characteristic curves for the LR, RF, SVM, LGBM, XGBoost, AdaB and GBDT models. (B) Calibration curve for the XGBoost model. (C) Precision-recall (PR) curves comparing the performance of different machine learning models. AdaB, AdaBoost; AP, average precision; GBDT, gradient boosting decision tree; LGBM, light gradient boosting machine; LR, logistic regression; RF, random forest; SVM, support vector machine; XGBoost, extreme gradient boosting.Table 2The results of model trainingModelAccuracySensitivitySpecificityPPVNPVLR0.7170.9120.2900.7380.600RF0.8590.9260.7100.8750.815SVM0.6770.8240.3550.7370.478LightGBM0.8480.9120.7100.8730.786XGBoost0.8690.8980.8060.9100.781AdaBoost0.8180.8820.6770.8570.724GBDT0.8480.8970.7420.8840.767AdaBoost, Adaptive Boosting; GBDT, gradient boosting decision tree; LightGBM, light gradient boosting machine; LR, logistic regression; NPV, negative predictive value; PPV, positive predictive value; RF, random forest; SVM, support vector machine; XGBoost, extreme gradient boosting.To evaluate the XGBoost model’s robustness and generalizability, we conducted fivefold cross-validation on the entire dataset. The AUROC values for the folds were 0.944 (95% CI 0.905 to 0.976), 0.906 (95% CI 0.84 to 0.965), 0.946 (95% CI 0.898 to 0.973) and 0.904 (95% CI 0.828 to 0.967), with a mean of 0.923 (SD=0.018). The narrow 95% CIs indicated stable generalization. XGBoost outperformed LR and AdaBoost, demonstrating superior discriminative power. These results, supported by online supplemental figure 1, confirm the model’s stability and high discriminative ability across data partitions.SP110.1136/bmjdrc-2025-005242.supp1Supplementary data The precision-recall curves and average Precision (AP) values are presented in figure 2C. Notably, XGBoost achieved the highest AP value of 0.942, with RF closely following at an AP of 0.937. Both LightGBM and GBDT also demonstrated high AP values of 0.916, highlighting the effectiveness of ensemble methods in managing imbalanced datasets. In contrast, LR and SVM yielded comparatively lower AP values of 0.835 and 0.845, respectively. These results collectively indicate that XGBoost offers superior discriminative capability and calibration, thereby justifying its selection as the optimal predictive model in this study.Interpretable prediction model Analysis of the overall risk factors This study employs SHAP to interpret the calibrated XGBoost model, visualizing feature importance through the beeswarm plot presented in figure 3, which intuitively illustrates the impact of 24 features on the model’s predictions. In the plot, the relative importance of features decreases from top to bottom. The color of the points indicates the feature value, with red points representing larger values and blue points representing smaller values. The horizontal position of the points represents the relative importance of the features to the predicted values; the closer a point is to the center line, the smaller the absolute SHAP value of that feature, indicating a lesser impact on the prediction. Conversely, a point further from the center line indicates a greater impact. The top 10 factors influencing recurrence risk included smoking index, age, hemoglobin levels, ischemia, glycated hemoglobin, body mass index (BMI), length of hospital stay, low-density lipoprotein (LDL) cholesterol, creatinine and white blood cell count. Statistical analysis confirmed significant differences between the recurrence and non-recurrence groups for these 10 features, as shown in table 3. These findings support the validity and relevance of the feature interpretations derived from the SHAP method.Figure 3Feature importance ranking from SHAP analysis. Each dot represents the impact of a feature on the prediction for an individual patient. Red dots indicate higher feature values, while blue dots indicate lower feature values. Dots positioned to the left of the x-axis represent features associated with a decreased recurrence prediction, whereas dots positioned to the right indicate features associated with an increased recurrence prediction. BMI, body mass index; Hb, hemoglobin; HbA1c, glycated hemoglobin; LDL-C, low-density lipoprotein cholesterol; HS, Hospital day; SHAP, SHapley Additive exPlanations; SI, Smoking Index = smoking years × number of cigarettes smoked per day; WBC, white blood cell count.Table 3Baseline characteristics of DFU patients from the training setCharacteristicRecurrence (n=247)Unrecurrence (n=145)P valueSI800.0 (800.0,1000.0)0.0 (0.0,600.0)<0.001Age, years73.0 (72.0,81.0)68.0 (67.0,75.0)<0.001Hb, mmol/L116.0 (114.0,133.0)125.0 (122.0,135.0)0.001Ischemia, 0/185/17262/760.020HbA1c, %8.3 (7.9,11.8)8.0 (7.8,9.6)0.017BMI, kg/m2 23.2 (22.9,25.7)22.2 (21.4,24.6)0.004LDL-C, mmol/L2.4 (2.3,2.7)2.2 (2.0,2.6)0.018Hospital duration, days16.0 (15.0,27.0)26.0 (25.0,48.0)<0.001Scr, umol/L83.8 (77.0,124.0)69.5 (65.1,97.0)0.002WBC, G/L8.8 (8.1,13.0)7.1 (6.8,9.1)0.003SI = smoking years ×number of cigarettes smoked per dayBMI, body mass index; DFU, diabetic foot ulcer; Hb, hemoglobin; HbA1c, glycated hemoglobin; LDL-C, low-density lipoprotein cholesterol; Scr, serum creatinine; SI, smoking index; WBC, white blood cell count.Analysis of personalized risk factors Building on the identified risk factors, we developed a personalized risk prediction tool based on SHAP values. This tool offers a scale from 0 to 1, reflecting the contribution of each feature to the predicted recurrence risk. We demonstrated the application of this personalized tool with examples from both recurrent and non-recurrent patients in the test cohort. For a recurrent patient in their 60s with a smoking index of 1000, a hospitalization duration of 39 days, an LDL cholesterol level of 1.2 mmol/L, BMI of 26.07 kg/m², blood sugar level of 23.1 mmol/L, creatinine level of 100 µmol/L, triglycerides of 0.63 mmol/L and ischemia, the baseline risk prediction was E[f(x)]=0.717. After accounting for all factors, the calculated risk of recurrence was 0.987. Specific contributions to this risk were made by age (+0.08), smoking index (+0.41), hospital stay (−0.38), LDL cholesterol (−0.34), BMI (+0.31), blood sugar (+0.18), creatinine (+0.08), triglycerides (+0.08) and ischemia (−0.08), as illustrated in figure 4A. This analysis aligns with clinical observations, confirming the model’s accuracy in predicting recurrence. A similar analysis can be performed to predict the likelihood of non-recurrence, as shown in figure 4B.Figure 4Examples of personalized risk factor analysis. (A) An example of personalized risk factor analysis for a patient in the test set (clinical outcome was recurrence). (B) An example of personalized risk factor analysis for a patient in the test set, where the actual clinical outcome was non-recurrence. BMI, body mass index; LDL-C, low-density lipoprotein cholesterol; BG, Blood Glucose; HS, Hospital day; TG, Triglyceride; Scr, serum creatinine; SI, Smoking Index = smoking years × number of cigarettes smoked per day; WBC, white blood cell count.Discussion Although researches on the recurrent risk factors of DFU have become increasingly profound, there is currently no systematic model to integrate these risk factors and accurately predict recurrent probability. 34 35 In this study, we developed and validated the first multidimensional prediction model for DFU recurrence, which integrates metabolic, hematological, inflammatory and clinical management factors using multicenter cohort data. The model includes 10 easily accessible predictors: smoking pack-years, HbA1c, ischemia, LDL-C, hemoglobin, age, BMI, serum creatinine, white blood cell count and hospitalization duration. This model provides a comprehensive tool that allows healthcare providers to assess and stratify the recurrence risk of DFU on an individual basis, enabling more personalized and effective management in routine clinical practice.As previous studies have shown, chronic hyperglycemia accelerates advanced glycation end product (AGEs) accumulation, while tobacco smoke further promotes the generation of AGEs through the activation of oxidative stress, such as Nicotinamide Adenine Dinucleotide Phosphate (NADPH) oxidase, which synergistically impairs endothelial function and collagen remodeling.36 This ‘dual-engine’ driving mechanism of metabolic-oxidative stress is particularly prominent in diabetic microangiopathy, potentially explaining the double risk of ulcer recurrence observed in patients with concurrent smoking and hyperglycemia. This also supports our previous ML study, which demonstrated that patients with hyperglycemic crises tend to have worse outcomes.16 37 The predictors in our DFU recurrence model—anemia, ischemia, leukocytosis and elevated LDL-C—collectively delineate a hypoxia-inflammation-metabolic vicious cycle. Anemia combined with tissue ischemia disrupts oxygen supply–demand equilibrium,38 39 whereas leukocytosis indicates persistent microenvironmental inflammation that may deplete local oxygen reserves. Concurrently, elevated LDL-C exacerbates lower limb ischemia through foam cell formation and atherosclerotic plaque destabilization, establishing a metabolic-ischemic vicious cycle.40 This is consistent with the previous studies.41 42 Obesity is a well-known risk factor of plantar foot ulcer recurrence, as previous studies have shown.11 43 44 Our findings validate these results, demonstrating the substantial influence of weight on foot ulcer recurrence. Therefore, patients should be closely monitored in the clinic until their weight issues are addressed.While prior DFU recurrence prediction models consistently overlooked hospitalization duration, our study found that short-term hospitalization (<7 days) may leave an increased risk of relapse due to residual biological risk of incomplete treatment (eg, debridement or inadequate course of antibiotics), whereas appropriate prolonged hospitalization reduces the risk of recurrence. This finding is consistent with the experience of the clinicians. Patients who require long-term hospitalization are often those with severe conditions and a higher risk of death. Appropriate duration of treatment increases the likelihood of being cured, hence reduces the risk of readmission. Thus, clinical practice demands dynamic stratification: enhancing transitional care for short-stay patients while initiating early rehabilitation in prolonged inpatients to disrupt the frailty-recurrence cycle. Thus, clinical practice demands dynamic stratification: enhancing transitional care for short-stay patients while initiating early rehabilitation in prolonged inpatients to disrupt the frailty-recurrence cycle.Given the superior performance of ML techniques in managing high-dimensional and non-linear data, an increasing number of researchers have recently applied these approaches to the study of DFU. Both ML and deep learning algorithms have demonstrated remarkable predictive capabilities by utilizing clinical features. In this study, the XGBoost model exhibited exceptional predictive performance, achieving an AUROC of 0.924, with a 95% CI of 0.867 to 0.967. While traditional statistical models such as LR and Cox regression are prevalent in clinical prediction due to their interpretability, they encounter limitations when dealing with high-dimensional non-linear relationships. The LR model developed by Aan de Stegge et al (AUROC=0.690)45 and the Cox regression model by Wang and Pan (AUROC=0.796)46 demonstrated significantly lower performance compared with our model.In studies specific to DFU, the conditional inference tree model developed by Stefanopoulos et al (AUROC=0.880)47 and the Naive Bayes model by Wang et al (AUROC=0.864)48demonstrated inferior performance compared with our model. This discrepancy may be attributed to variations in prediction endpoints: the aforementioned studies concentrated on ‘short-term healing’, whereas our research focused on the more complex endpoint of ‘recurrence’. The incorporation of 3-year longitudinal data significantly enhanced the accuracy of long-term risk prediction. Although the multimodel ensemble by Zhang et al (AUROC=0.937)49 exhibited marginally superior performance compared with our model, it did not report calibration metrics. In contrast, our model’s Brier score substantiates the reliability of its probability predictions, thereby rendering it more suitable for clinical risk quantification (online supplemental table 4).SP710.1136/bmjdrc-2025-005242.supp7Supplementary data In conclusion, our XGBoost model demonstrates markedly superior discriminative capability compared with traditional regression models and the majority of DFU-specific models. Its performance is on par with high-performance ML models, while offering calibration and interpretability that are more suited to clinical practice. These findings advocate for the use of ensemble learning models in predicting DFU recurrence, especially within patient populations characterized by complex comorbidities.By integrating a range of clinical factors, including established disease scoring systems, into an ML framework, our study significantly enhances the predictive accuracy of the model, setting it apart from previous iterations. Furthermore, we have developed an online prediction tool designed to provide users with easy access to patient data and reliable prognostic assessments, potentially aiding general practitioners (GPs) in the early evaluation of DFU recurrence risk. Presently, the tool is deployed exclusively in Chongqing, with its initial development grounded in domestic research data. It is currently in the beta testing phase, during which local GPs and specialists are actively contributing to its ongoing training and refinement. Simultaneously, efforts are being made to broaden the data sources and optimize the tool, with the ultimate aim of extending its application to support diabetic foot patients on a global scale.Our study has several limitations. First, the model was developed retrospectively using data from multiple hospitals in Southwest China, which may introduce potential biases. Second, the small sample size of the external validation set could impact the robustness of the model’s evaluation. Additionally, the variability in clinical data across different hospitals presents challenges in developing a universally applicable predictive model. To address these issues, future studies should include larger, multicenter validation cohorts to enhance the model’s generalizability. Third, although the model is designed to handle missing data, certain important features, such as the time between symptom onset and hospital admission, were frequently missing, which limits further in-depth analysis.Conclusions In conclusion, the results of this study indicate that the XGBoost model is an effective and promising tool for predicting the risk of DFU recurrence. The proposed online prediction tool can assist non-specialist physicians or community healthcare providers in the early diagnosis and assessment of DFU recurrence. Given the complexity of the pathogenesis of diabetic wounds, 50 further multicenter, prospective studies are needed to validate these findings and explore additional factors influencing recurrence, thereby enhancing the model’s predictive accuracy. Overall, this study offers a novel approach to DFU management, with the potential to significantly improve clinical outcomes for patients with diabetes.SP510.1136/bmjdrc-2025-005242.supp5Supplementary data SP610.1136/bmjdrc-2025-005242.supp6Supplementary data