Document resource
Introduction Systemic sclerosis (SSc) is a rare disease with substantial heterogeneity in long-term outcomes. Prognostic models using simple clinical predictors such as the modified Rodnan Skin Score (mRSS) and organ involvement factors have been proposed, but their external validity remains uncertain. Reliable external validation is essential before such models can be considered for clinical decision support.Material and Methods We conducted a preliminary external validation of two previously published deep learning models predicting 5-year mortality, one based on model 1 (mRSS plus World Health Organization functional class (WHO-FC) II) and model 2 (mRSS plus WHO-FC III). An independent cohort of 344 patients (62 deaths, 18% mortality) was used. Data were preprocessed using the original MinMaxScaler parameters, and model predictions were generated without re-training. Discrimination was assessed by Area under the Receiver Operating Characteristic Curve (AUROC) and Area Under the Precision-Recall Curve (AUPRC), calibration by Brier score, calibration-in-the-large, and calibration slope (slope). Threshold-based performance (sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and F1) was evaluated at cut-off 0.5 and by Youden’s J statistic. Recalibration was performed to improve probability estimates. Bootstrap resampling (B=1000) was used to calculate 95% confidence intervals.Results For the model 1, AUROC was 0.80 (95% CI: 0.75 – 0.86), AUPRC 0.55, and Brier score 0.23; recalibration reduced the Brier score to 0.15 while discrimination was unchanged. For the model 2, AUROC was 0.76 (95% CI: 0.79 – 0.84), AUPRC 0.55, and Brier score 0.25, improving to 0.15 after recalibration. At the 0.5 threshold, both models showed high specificity (>85%) but limited sensitivity (26 –61%) and PPV (0.50 – 0.60). Youden’s index suggested alternative thresholds with improved sensitivity 0.65 and greater at the expense of specificity.Conclusions In this preliminary external validation, both models showed moderate discrimination and acceptable NPV but suboptimal calibration and limited PPV, restricting their readiness for clinical deployment. Recalibration improved probability estimates but did not materially after discrimination. These findings highlight the importance of external validation and recalibration before translation to clinical decision support systems. Further refinement and validation in larger and more diverse external cohorts are required.