Document resource
The growing demand for accessible, high-quality and privacy-preserving health data has led to increased interest in synthetic health data as a promising solution to overcome data scarcity and legal barriers.1 Synthetic data refer to information that has been created artificially to mimic real-world observations. This is particularly relevant in the context of rare diseases, where real-world data are often fragmented, siloed or insufficient for robust artificial intelligence (AI) development and clinical research.2 This paper summarises the outcomes of a multidisciplinary Sandpit workshop involving experts with lived experiences in rare diseases, as well as experts in clinical medicine, data science, cybersecurity and medical informatics. The goal was to define a shared vision and roadmap for a synthetic health data repository (SHARE).