Document resource
Background Prospective and historic (back to 1998) germline testing data have been collated across all English diagnostic testing laboratories by the National Disease Registration Service (NDRS), forming a dataset of over 85,000 HBOC patients tested for BRCA1/BRCA2, offering great potential to inform UK variant classification. We have developed a methodology for leveraging enriched nationally-collected datasets within case-control analysis via adjusted integration with unselected datasets using novel likelihood ratio tools.Methods We estimated enrichment of pathogenic variants in clinically-ascertained laboratory data from both the NDRS (England) and Ambry (USA) using truncating variant prevalence, and paired data with equivalent population controls. Using our published PS4-LRCalc tool, we calculated a combined log likelihood ratio (LLR) across five datasets (three unselected, and two enriched).Results Data were combined for 10,820 missense variants from 325,255 female breast cancer patients and 671,350 controls of Western European ancestry for five breast cancer susceptibility genes (BRCA1, BRCA2, PALB2, ATM, CHEK2). A combined LLR was produced for 5,360 missense variants; 934 variants received evidence towards pathogenicity (LLR≥1), and 3,791 received evidence towards benignity (LLR≤1).Conclusion This novel variant-level methodology leverages value from nationally collected laboratory variant data and empowers flexible incorporation of case-control data (PS4) into variant classification.