BetaEntity Annotation Prototype
← Back to institutions

Annotated full text

Foundation models in computational pathology: methods, applications and clinical implications

bmjonc · 2026-05-08 · canonical JSON source

30 visible annotations · policy: published · automated confidence ≥ 75.00%

Document resource

Introduction Pathology foundation models (PFMs) built on artificial intelligence (AI) frameworks are catalysing a paradigm shift in diagnostic pathology, moving the field beyond narrowly optimised task-specific AI toward general-purpose computational representations of tissue morphology, that is, PFMs. Since the nineteenth-century, following the principles of cellular pathology established by Rudolf Virchow, the definitive diagnosis of cancer has relied on the microscopic examination of histological glass slides. 1 This diagnostic paradigm remained largely unchanged for more than a century, until the early 2000s marked a pivotal transition with the commercial introduction of whole-slide imaging (WSI) scanners, enabling histopathology to enter a fully digital domain.2 3 Over the past decade, improvements in scanner resolution and image quality, combined with enhanced data transmission, expanded storage capacity and declining costs, have facilitated the widespread adoption of digital pathology, transforming histology into a data-rich resource suitable for large-scale computational analysis.4In parallel with this digitisation, AI has emerged as a powerful tool for automated histopathological analysis, with early successes driven primarily by deep learning architectures such as convolutional neural networks (CNNs).5 These models, designed for specific tasks, have shown robust performance in narrowly defined applications, such as tumour detection, grading and cell quantification. However, their clinical scalability has been constrained by a rigid reliance on supervised learning and expert annotations, leading to substantial brittleness under real-world conditions. Variability in scanner hardware, staining protocols and laboratory workflows can result in pronounced performance degradation, while the labor-intensive requirement for dense, pathologist-generated labels creates a persistent annotation bottleneck particularly for rare cancers and underrepresented tumour subtypes.6Traditional AI approaches in histopathology have largely focused on optimizing specialised models for individual diagnostic tasks.7 Over the past decade, numerous task-specific AI platforms have been developed for clinical applications spanning diagnosis, prognosis and treatment response prediction.8 Despite the rapid expansion of pathology AI, only three products have received US Food and Drug Administration (FDA) authorisation for clinical use, via either 510(k) clearance or De Novo pathways. Two of these products, Paige AI (authorised via the De Novo pathway) and Ibex (cleared through 510(k)), focus on prostate cancer diagnosis and support detection by highlighting regions of malignancy. The ArteraAI prostate platform is unique in receiving FDA De Novo authorisation for prognostic risk stratification in prostate cancer using prostate biopsy specimens and complementary clinical data,9 supporting clinical decision-making for treatment planning. In contrast, the vast majority of pathology AI platforms remain unauthorised for clinical use, including PathAI (breast, gastrointestinal and liver-specific models for research use only),10 11 Proscia/Concentriq AI (tumour-specific models for workflow and decision support),12 NovinoAI (prostate, gastrointestinal, urine cytology models as well as workflow improvement models),13 14 Aiforia (breast, skin, gastrointestinal and kidney models with Conformite Europeenne (CE) marking),15 16 Deep Bio (prostate, breast and gastric cancer detection models with CE marking)17 and Ibex (breast and gastrointestinal models with CE marking), among others.PFMs address these systemic limitations by shifting the learning objective from task-level supervision to large-scale self-supervised and weakly supervised representation learning. By pre-training on millions of unannotated or minimally labelled WSIs spanning diverse tissues, institutions and technical conditions, these models learn transferable, task-agnostic embeddings that encode general histomorphological patterns. This strategy substantially reduces reliance on extensive manual annotation, improves robustness to technical heterogeneity and enables efficient adaptation to multiple downstream tasks using comparatively limited labelled data. Rather than being optimised for a single diagnostic objective, PFMs function as reusable computational infrastructure capable of supporting cancer detection, subtyping, grading, biomarker prediction and prognostication within a unified framework.18While PFMs are increasingly being explored across multiple domains in oncology, including radiology, genomics and multimodal clinical prediction, this review is specifically focused on their development and application in computational pathology (CPath), with particular emphasis on histopathology image analysis using WSI. We detail the design and scope of the review, describe the literature search strategy and outline inclusion and exclusion criteria, study selection, data extraction and synthesis approach. In the results and evidence synthesis, we present an overview of PFMs, including their technical foundations, evolution, major trends, model types and capabilities. We further explore the transition from foundation models to agentic AI systems and discuss their potential integration into clinical workflows. Regulatory and governance considerations are addressed to contextualise the clinical translation of these models. Finally, we critically examine their limitations and translational challenges, providing a foundation for future research and practical guidance for the responsible adoption of foundation models in digital pathology.Methods Review design and scope This study was conducted as a narrative review synthesising recent developments in PFMs applied to digital pathology for cancer diagnosis to investigate: (1) conceptual advances, (2) model design choices and (3) translational implications. We focused on general-purpose, pretrained models with an emphasis on vision, vision-language, multimodal and emerging agentic frameworks.Literature search strategy A structured literature search was performed in PubMed, Europe PMC, ClinicalKey and the Cochrane Database of Systematic Reviews, the latter primarily for contextual and background references. The search covered publications from January 2023 to December 2025, corresponding to the period during which PFMs have rapidly emerged and matured.The search strategy was designed around three core conceptual domains: digital pathology, foundation models and oncological applications. Within each domain, relevant keywords and controlled vocabulary terms were combined to capture variations in terminology. Concepts related to digital pathology included WSI and histopathology; foundation models concepts encompassed large pre-trained vision-only models, vision–language models and multimodal architectures; and clinical application terms focused on cancer diagnosis, classification, prognosis and molecular or biomarker prediction.To ensure comprehensive coverage of emerging multimodal approaches, additional terms were included to identify studies integrating pathology with other data modalities, such as radiology–pathology fusion and image–text modelling. Reference lists of included articles were manually screened to identify additional relevant studies not captured by the database searches.As model names and terminology evolved across publications, search terms were updated to include newly identified model names and architectural descriptors to ensure comprehensive coverage. Reference lists of key foundation model publications (including Universal Network for Imaging (UNI), Clinical Histopathology Imaging Evaluation Foundation model (CHIEF), CONtrastive learning from Captions for Histopathology (CONCH), Virchow, Transformer-based pathology Image and Text Alignment Network (TITAN) and related multimodal models) were manually screened to identify additional relevant studies. This search strategy was intended to support a narrative synthesis of representative and influential foundation models, rather than an exhaustive systematic review.Inclusion criteria Studies were included in this review if they met all of the following criteria: (1) publication between January 2023 and December 2025; (2) original research articles published in peer-reviewed journals or conference proceedings; and (3) primary focus on digital or CPath, or on multimodal AI models incorporating pathology data. In addition, eligible studies were required to report at least one of the predefined outcomes of interest relevant to foundation or multimodal model development.Exclusion criteria Studies were excluded if they met any of the following conditions: (1) use of conventional supervised, task-specific models without large-scale pretraining or foundation model characteristics; (2) exclusive focus on radiology-based applications without integration of, or relevance to, pathology data; (3) investigation of non-oncological applications unrelated to cancer; (4) presentation of purely technical or methodological demonstrations lacking evaluation on real-world pathology datasets; (5) publication as review articles, editorials, commentaries or opinion pieces; or (6) publication in languages other than English.Study selection and data extraction Titles and abstracts were independently screened for relevance by multiple authors. Screening and full-text assessment were conducted by domain experts in pathology, AI and oncology. Studies meeting the predefined inclusion criteria underwent full-text review.For each included study, data were systematically extracted using a standardised framework, including: (1) model architecture and modality (vision-only, vision–language or multimodal); (2) training data characteristics, including data type, scale and supervision strategy; (3) primary diagnostic, prognostic or predictive applications; (4) reported performance characteristics, including evaluation datasets and performance metrics; and (5) reported limitations and translational considerations, including interpretability or explainability approaches and stated translational intent (research-only, workflow support, or diagnostic use).Synthesis approach Models were grouped according to their conceptual approach, scope of applicability, technical design, clinical relevance and reported limitations. Tables 1–3 provide structured comparative summaries of representative PFMs and multimodal models identified through the literature search. While the main text focuses on selected examples, the tables additionally include models not discussed in depth in order to provide broader contextual coverage of the rapidly evolving research landscape. Models are grouped by modality and ordered chronologically to illustrate trends in model scale, architectural complexity and downstream task scope. An expanded summary of each model, including architectural details, training scale and evaluation benchmarks, is presented in online supplemental table 1.SP110.1136/bmjonc-2026-001102.supp1Supplementary dataTable 2Vision–language foundation modelsModel nameYear releasedLicense typeArchitectureParametersMultiresolution supportDataset scaleEvaluated downstream tasksPLIP522023-03CC BY-NC-SA 4.0 - research use onlyViT-B/16 (CLIP-based)86 millionNo (native 20×)208 414 image-text pairs (OpenPath)Zero-shot subtyping and cross-modal image-text search.MI-Zero532023-04CC BY-NC-SA 4.0 - research use onlyViT-B/16 (CLIP backbone) + aggregator86 millionNo (native 20×)Built on CONCH (1.17M pairs) pre-trained encoderZero-shot slide-level subtyping and CUP prediction.CONCH542023-05CC BY-NC-SA 4.0 - research use onlyViT-B/16 (CLIP-based)86 millionNo (native 20×/0.5 MPP)1.17 million image-caption pairsZero-shot classification, image-text retrieval and slide-level subtyping.QuiltNet552023-06Apache 2.0 - commercial use permittedViT-B/1686 millionNo (native 20×)1 million image-text pairs (Quilt-1M)Zero-shot classification and pathology-specific image-text retrieval.PathChat562024-01CC BY-NC-SA 4.0 - research use onlyViT-L/14+Llama-2-13B (LLM)13.3 billionNo (native 20×)1.17 million image-caption pairs (CONCH-based)VQA, diagnostic reasoning and board examination simulation.PathCLIP572024-01Apache 2.0 - commercial use permittedViT-B/1686 millionNo (native 20×)208 414 image-text pairs (OpenPath)Zero-shot classification and cross-modal image-text retrieval.Quilt-LLaVA582023-12Apache 2.0 - commercial use permittedViT-L/14+LLM (Vicuna)~7.3 billionNo (native 20×)1 million image-text pairs (Quilt-1M)Visual instruction tuning, pathology VQA and diagnostic chat.PRISM592024-05Proprietary (Paige) - commercial via agreementSwin-B/ViT+generative decoder558 millionYes (multiresolution support)587 196 WSIs (proprietary)Generative slide modelling and slide-level feature representation.PathGen-1.6M602024-06CC BY-NC 4.0 - research use onlyViT-L/14 (vision–language)300 millionNo (native 20×)1.6 million image-text pairsSynthetic image generation and zero-shot patch-level classification.EXAONE Path 2.0612024-07Proprietary (LG AI) - research use onlyViT-L/14300 millionYes (5×, 10×, 20×, 40×)100 000 Patches (TCGA-based 1.0 version)Patch classification, image retrieval and subtyping benchmarks.KEEP622024-12CC BY-NC-SA 4.0 - research use onlyKnowledge-enhanced ViT-B/1686 millionNo (native 20×)1.2 million image-text pairs (Quilt-1M+OpenPath)Knowledge-enhanced diagnosis, zero-shot classification and retrieval.TITAN632024-11CC BY-NC-SA 4.0 - research use onlyViT+long-range slide transformer86 millionNo (native 20×/0.5 MPP)340 000 WSIs (Mass-340K dataset)Whole-slide cancer subtyping and survival/prognosis prediction.MUSK642025-01Research Only - research use onlyViT-L/14+LLM fusion300 millionYes (multiscale aggregation)50 million patches+1 million image-text pairsPrecision oncology tasks, molecular alteration prediction and subtyping.MR-PLIP652025-04CC BY-NC-SA 4.0 - research use onlyMultiresolution vision transformer~300 millionYes (Multiresolution: 5×, 10×, 20×, 40×)~208 000 image-text pairs (multiresolution OpenPath)Multiresolution image-text retrieval and zero-shot diagnostic classification.CLIP, Contrastive Language–Image Pretraining; CONCH, CONtrastive learning from Captions for Histopathology; CUP, cancer of unknown primary; LLM, Large Language Model; MPP, Microns Per Pixel; TCGA, The Cancer Genome Atlas; TITAN, Transformer-based pathology Image and Text Alignment Network; ViT, vision transformer; VQA, visual question answering; WSI, whole-slide imaging.Table 1Vision-only foundation modelsModel nameYear releasedLicense typeArchitectureParametersMultiresolution supportDataset scaleEvaluated downstream tasksCTransPath662022-03Apache 2.0 - commercial use permittedHybrid CNN-Swin Transformer28 millionNo (native 20×/0.5 MPP)32 220 WSIs (TCGA+PAIP)Patch-level classification and slide-level cancer subtyping.HIPT672022-03MIT - commercial use permittedHierarchical Vision Transformer (ViT-S/16)<10 millionYes (hierarchical: 5×, 10×, 20×)11 000 WSIs (TCGA)Slide-level subtyping, survival prediction and CUP origin.UNI (CPath)272023-08CC BY-NC-SA 4.0 - research use onlyViT-L/16300 millionNo (native 20×/0.5 MPP)100 426 WSIs (Mass-100K dataset)30+ tasks including subtyping, Gleason grading and biomarker (MSI) prediction.Phikon682023-09Apache 2.0 - commercial use permittedViT-B/16 (iBOT-based)86 millionNo (native 20×)6100 WSIs (TCGA subsample)Subtyping, survival and standard histology patch benchmarks.Virchow692023-09Virchow License - non-commercial research onlyViT-H/14632 millionNo (native 20×/0.5 MPP)1.5 million WSIs (proprietary MSKCC)MSI prediction, cancer subtyping and prostate Gleason grading.RudolfV702024-01Apache 2.0 - commercial use permittedViT-L/14300 millionNo (native 20×)130 000 WSIs (proprietary+TCGA)Search and retrieval, pan-cancer patch classification and subtyping.Kaiko-B82024-04Apache 2.0 - commercial use permittedViT-B/886 millionNo (native 20×/0.5 MPP)29 000 WSIs (TCGA+internal)Patch classification (CRC, lymph node) and WSI-level subtyping.Prov-GigaPath712024-05GigaPath License - non-commercial research onlyLong-context ViT (dilated attention)1.1 billionNo (native 20×/0.5 MPP)171 189 WSIs (providence dataset)Pan-cancer subtyping, survival and genetic mutation prediction (from H&E).GigaPath712024-05GigaPath License - non-commercial research onlyViT-L/14 (dilated attention)1.1 billionNo (native 20×/0.5 MPP)171 189 WSIs (1.3 billion patches)Subtyping, genetic mutation prediction and slide-level survival analysis.Hibou-B722024-06Apache 2.0 - commercial use permittedViT-B/1486 millionYes (mixed 0.25–2.0 MPP)1.14 million WSIs (936k H&E, 202k non-H&E)Pan-cancer subtyping, survival analysis and metastasis detection.Hibou-L722024-06Apache 2.0 - commercial use permittedViT-L/14300 millionYes (mixed 0.25–2.0 MPP)1.14 million WSIs (936k H&E, 202k non-H&E)Pan-cancer subtyping, survival analysis and metastasis detection.H-Optimus-02024-07Apache 2.0 - commercial use permittedViT-g/14 (giant)1.1 billionNo (native 20×/0.5 MPP)500 000 WSIs (proprietary)Subtyping, PANDA Gleason grading and CRC patch classification.Virchow2732024-08Virchow License - non-commercial research onlyViT-H/14632 millionYes (mixed 5×, 10×, 20×, 40×)3.1 million WSIs (MSKCC+international)Clinical biomarker prediction, subtyping and mixed-magnification analysis.Virchow2G732024-08Virchow License - non-commercial research onlyViT-g/14 (giant)1.85 billionYes (mixed 5×, 10×, 20×, 40×)3.1 million WSIs (MSKCC+international)SOTA clinical biomarker prediction and pan-cancer subtyping.CHIEF742024-09CC BY-NC-SA 4.0 - research use onlyViT-B/16 with hierarchical aggregation86 millionYes (5×, 10×, 20×, 40×)60 530 WSIs (spanning 19 anatomical sites)Cancer subtyping, survival prediction and HRD (genomic) status prediction.Phikon-v2752024-09Apache 2.0 - commercial use permittedViT-L/14 (DINOv2-based)300 millionNo (native 20×)58 400 WSIs (PANCAN-XL collection)PANDA Gleason grading, metastasis detection and cancer subtyping.H-Optimus-12024-10Apache 2.0 - commercial use permittedViT-g/14 (giant)1.1 billionNo (native 20×/0.5 MPP)1 million+ WSIs (from 800 000+ patients)Broad subtyping, metastasis detection and prostate grading.PLUTO762024-5Proprietary (PathAI) - commercial via agreementViT-B (vision transformer)86 millionYes (5×, 10×, 20×, 40×)160 000 WSIs (proprietary)Clinical diagnostic tasks, subtyping and Paige proprietary benchmarks.PLUTO-4772025-11Proprietary (PathAI) - commercial via agreementViT-L (scaled ViT)1.1 billionYes (5×, 10×, 20×, 40×)160 000+WSIs (scaled proprietary)Scale-up clinical diagnostics and pan-cancer subtyping.Atlas782025-01Proprietary (Aignostics/Mayo) - research onlyViT-H/14 (vision transformer)632 millionNo (Native 20×/0.5 MPP)61 300 WSIs (1.3 billion patches)Pan-cancer subtyping, metastasis detection and clinical cohort validation.UNI-2h2025-01CC BY-NC-SA 4.0 - research use onlyViT-h/14-reg8 (hierarchical)632 millionYes (hierarchical 5×, 10×, 20×)350 000 WSIs (H&E and IHC slides)Hierarchical subtyping, survival analysis and complex diagnostic tasks.PathOrchestra792025-03Apache 2.0 - commercial use permittedViT (DINOv2 backbones)1.1 billionNo (Native 20x) - (evaluated on multi-mag tasks)300 000 WSIs (multicentric)100+ clinical tasks including subtyping, grading and mutation prediction.Midnight 12k802025-04Apache 2.0 - commercial use permittedViT-g/141.1 billionNo (native 20×)12 000 WSIs (TCGA)Standard subtyping and patch classification benchmarks.Midnight 92k802025-04Apache 2.0 - commercial use permittedViT-g/141.1 billionNo (native 20×)92 000 WSIs (TCGA+NKI-80k)Pan-cancer subtyping, survival prediction and biological feature extraction.Midnight 92k/392802025-04Apache 2.0 - commercial use permittedViT-g/141.1 billionNo (native 20×, high-res tile support)92 000 WSIs (high-res training on 512px tiles)Nuclei segmentation, cell-level classification and high-res feature mapping.OpenMidnight2025-05Apache 2.0 - commercial use permittedViT-g/14 (replication)1.1 billionNo (native 20×)12 000 or 92 000 WSIs (midnight replication)Cancer subtyping and standard patch-level classification (replication).CNN, convolutional neural network; CRC, colorectal cancer; CUP, cancer of unknown primary; HRD, homologous recombination deficiency; iBOT, image BERT pretraining with online tokenizer; IHC, Immunohistochemistry; MPP, microns per pixel; MSI, microsatellite instability; MSKCC, Memorial Sloan Kettering Cancer Center; NKI, Netherlands Cancer Institute; PAIP, Pathology Artificial Intelligence Platform; PANCAN-XL, Pan-Cancer Extra Large collection; PANDA, Prostate cANcer graDe Assessment; SOTA, state of the art; TCGA, The Cancer Genome Atlas; UNI, Universal Network for Imaging; ViTs, vision transformers; WSI, whole-slide imaging.Table 3Multimodal foundation modelsModel nameYear releasedMultimodal typeLicense typeArchitectureParametersMultiresolution supportDataset scaleEvaluated downstream tasksmSTAR352024-07Vision, language, genomicsProprietary - research use onlyViT-L/14+multimodal transformer300 millionNo (native 20×)10,000 WSIs (TCGA+multimodal reports)Survival analysis, clinical outcome prediction and multimodal genomic correlation.spEMO812025-01Vision, spatial multi-omicsMIT - commercial use permittedMultimodal transformerMultimodalYes (multiscale spatial features)620 000 patch-level image-gene pairsSpatial transcriptomics alignment and multi-omics niche identification.Threads822025-01Vision, molecular/genomicsProprietary (ArteraAI) - research use onlyHierarchical vision transformer11.3 millionYes (multiresolution hierarchical)47 171 WSIs (MGH+BWH+TCGA)HRD prediction, MSI prediction and patient survival analysis.PRISM2832025-06Vision, language, clinical dialogueProprietary (Paige) - commercial via agreementMultimodal transformer (vision+text)4.6 billionYes (multiresolution support)2.3 million WSIs+clinical dialogue dataMultimodal medical dialogue and automated clinical report generation.EXAONE Path 2.5362025-12Vision, language, spatial omicsCC BY-NC-SA 4.0 - research use onlyViT-g/14 (Giant) with cross-modal bridge1.1 billionYes (mixed 5×, 10×, 20×, 40×)23 099 WSIs (multi-omics aligned)Gene expression prediction and pan-cancer subtypingBWH, Brigham and Women’s Hospital; HRD, Homologous Recombination Deficiency; MGH, Massachusetts General Hospital; MIT, Massachusetts Institute of Technology; mSTAR, Multimodal Self-TAught PRetraining; TCGA, The Cancer Genome Atlas; ViT, vision transformer; WSI, whole-slide imaging.Results and evidence synthesis Pathology foundation models In this context, our discussion centres on PFMs designed for computational pathology workflows, where representation learning is typically performed on digitised histopathology images and aggregated for downstream slide-level or case-level tasks. In general, PFMs in CPath employ a two-stage training paradigm ( figure 1). The first stage involves task-agnostic pre-training of transformer-based architectures on large, heterogeneous datasets, typically comprising thousands to millions of WSIs across diverse tissue types and staining protocols, and may also incorporate complementary data modalities such as pathology reports and molecular profiles. This pre-training typically relies on self-supervised learning (SSL),19 which enables the model to extract rich, generalisable features without extensive manual annotation. In the second stage, the pre-trained embeddings are fine-tuned or adapted for specific downstream applications, such as cancer detection, tumour subtyping, biomarker assessment and prognostic prediction.Figure 1Two-stage training paradigm for pathology foundation models: large-scale self-supervised pretraining on unlabelled data to learn general tissue representations (Step 1), followed by task-specific fine-tuning with minimal labelled data for clinical downstream applications (Step 2).To understand how these capabilities are achieved, it is essential to first review the technical principles that underpin the development of PFMs.Overview of technical aspects for PFMs This section outlines the essential technical concepts required to support a shared understanding of foundation models, providing sufficient background for readers from diverse disciplines to follow the material presented in the sections that follow.Self-supervised learning in digital pathology SSL underpins modern foundation model development across domains, enabling large-scale representation learning without reliance on exhaustive expert annotations. In digital pathology, SSL has been particularly impactful due to the vast number of WSIs and the scarcity of detailed expert labels. In SSL, supervisory signals are derived from the intrinsic structure of the data through pretext tasks, allowing models to exploit unlabelled image repositories effectively. This approach is well suited to pathology, where manual annotation is costly, time-consuming, error-prone and often infeasible at scale, while histological slides exhibit rich morphological patterns across multiple spatial resolutions. By learning invariant and transferable representations, SSL-based models provide robust initialisations that can be adapted to a variety of downstream diagnostic, prognostic and predictive tasks. 20 21Contrastive learning in digital pathology Contrastive learning 22 is a widely used SSL strategy that has played a key role in foundation model development for digital pathology. The core idea is to learn representations by bringing semantically similar samples closer in the embedding space while pushing dissimilar samples apart. In practice, this is achieved by constructing positive pairs, such as alternative augmentations or spatially related regions from the same tissue slide, alongside negative pairs sampled from unrelated images. By optimizing these relationships, models learn invariant features that capture essential morphological patterns. Several PFMs in CPath have leveraged contrastive pre-training to achieve improved generalisation across diverse datasets and institutions, highlighting its practical utility in real-world clinical scenarios.20 21 23Masked image modelling in digital pathology Masked image modelling (MIM) 24 is a SSL strategy in which a subset of image patches or tokens is masked, and the model is trained to reconstruct the missing content from the remaining visible context. This objective encourages the learning of contextual representations that capture both local morphology and broader spatial dependencies. In CPath, MIM has been primarily implemented using transformer-based architectures, which are well-suited to modelling long-range relationships across gigapixel WSIs. In practice, MIM provides complementary benefits to contrastive learning approaches, with the two paradigms often yielding distinct but partially overlapping feature representations.25Representation learning at whole-slide scale WSIs pose unique challenges for representation learning due to their gigapixel resolution and complex tissue architecture. To address this, PFMs typically divide WSIs into smaller patches or tiles, from which feature embeddings are extracted using self-supervised strategies such as contrastive learning or MIM. These patch-level embeddings are subsequently aggregated through methods including attention-based pooling, graph representations or hierarchical encoders to produce slide-level representations that capture both local morphological details and global tissue context. This multiscale strategy allows foundation models to generalise across diverse histological patterns, supporting a wide range of downstream tasks. 26 27Vision transformers and large-context modelling Vision transformers (ViTs) have emerged as a general-purpose architecture in computer vision and are increasingly adopted in CPath due to their capacity to model long-range spatial dependencies in high-resolution images. Unlike CNNs, which rely on localised receptive fields and hierarchical feature aggregation, ViTs tokenise images into fixed-size patches and employ self-attention mechanisms to capture interactions across both local and global tissue contexts. This global receptive field is particularly advantageous for WSIs, where diagnostically and prognostically relevant features may span multiple spatial scales and anatomically distant regions. When coupled with SSL paradigms, ViTs enable PFMs to learn high-dimensional, transferable representations that integrate fine-grained cellular morphology with broader tissue architecture, supporting robust performance across diverse downstream tasks. 23 25Multimodal alignment in digital pathology Multimodal alignment refers to computational frameworks that jointly model heterogeneous biomedical data modalities—including histopathology WSI, pathology reports, radiological imaging, genomic and other omics profiles and structured clinical data—to learn shared or coordinated representation spaces. 28 29 By aligning embeddings across modalities, these approaches enable the integration of complementary, non-redundant information, allowing models to capture associations between tissue morphology, molecular alterations and clinical context that are not recoverable from unimodal data alone. Multimodal alignment is commonly achieved through contrastive learning objectives that maximise agreement between paired samples (e.g, image–report or image–genomic profiles), as well as through cross-modal attention and fusion architectures that model intermodality dependencies at feature and token levels. These strategies support a range of downstream tasks, including cross-modal retrieval, classification and survival modelling, while enabling clinically relevant applications such as prediction of molecular subtypes or mutational status directly from histomorphology and grounding image-based predictions in interpretable textual or clinical features.Types of pathology foundation models PFMs can be organised according to the types of input data they process and the range of tasks they support. Broadly, they fall into three categories: vision-only models, which learn exclusively from histopathology images; vision–language models, which integrate images with textual pathology reports; and specialised multimodal models, which combine histology with additional data types such as genomics or radiology. This classification provides a practical framework for structuring the literature, comparing architectural designs, pre-training strategies, data requirements and downstream applications in CPath and precision oncology.Vision-only pathology foundation models Vision-only PFMs represent the first frontier of large-scale representation learning in the field, operating exclusively on histopathology imagery such as WSIs or high-resolution image tiles ( table 1). To circumvent the annotation bottleneck inherent in manual expert labeling, these models use SSL to internalise the complex morphological language of human tissue from vast, unannotated slide repositories. Methodologically, the field has evolved from early contrastive learning20 frameworks like SimCLR and MoCo, which focus on feature invariance to more self-distillation methods such as DINOv230 and MIM.24 These contemporary approaches encourage models to reconstruct missing tissue segments or match global-to-local crops, fostering a deep understanding of spatial heterogeneity. Architecturally, there has been a significant transition from traditional CNNs toward ViTs.UNI/UNI2 UNI is a general-purpose PFM introduced by Chen et al and published in Nature Medicine in 2024.27 Developed by the Mahmood Lab, UNI was the first vision-based foundation model designed to operate broadly across CPath tasks. It was pretrained using SSL on the ‘Mass-100K’ dataset, comprising over 100 million image patches extracted from 100 426 diagnostic H&E-stained WSIs (>77 TB) spanning 20 major tissue and organ types, with data drawn from large public cohorts such as The Cancer Genome Atlas (TCGA), Prostate Cancer Grade Assessment(PANDA) and Cancer Metastases in Lymph Nodes (CAMELYON). The model employs Vision Transformer-Large architecture (ViT-L/16) as a backbone and was trained using the DINOv2 framework with self-distillation.Across 34 downstream clinical tasks—including cancer classification, disease subtyping, tissue and organ classification and transplant assessment—UNI outperformed prior state-of-the-art models such as CTransPath and REMEDIS. It enabled new capabilities in CPath, including resolution-agnostic tissue analysis, slide-level prediction beyond patch aggregation and broad cancer subtyping with up to 108 OncoTree classes, while also showing strong performance on rare and diagnostically challenging cancer types.27 In January 2025, the Mahmood Lab released UNI2 (UNI2-h), a scaled successor based on a Vision Transformer-Giant architecture (ViT-H/14). UNI2 was pre-trained on more than 200 million image tiles from over 350 000 diverse H&E and immunohistochemistry slides, resulting in consistent performance improvements over the original UNI, including higher accuracy on the TCGA Uniform Tumour classification task (0.675 vs 0.595) and strong external validation, with a reported macro area under the curve (AUC) of 0.780 for complex tasks such as endometrial cancer molecular subtyping.Virchow/ Virchow2 Virchow is among the largest image-only foundation models developed for CPath, created through a collaboration between Paige and Microsoft Research. 31 The model is pretrained using SSL on approximately 1.5 million H&E stained WSIs derived from diverse tissue and specimen types, primarily sourced from Memorial Sloan Kettering Cancer Centre. Virchow employs a tile-based representation learning strategy in which small image patches extracted from WSIs are processed by a deep feature extractor to generate high-dimensional embeddings that encode histomorphological information. These embeddings can be aggregated to support downstream classifiers, enabling applications across cell-level and tissue-level tasks, including classification, segmentation and morphology-driven analysis. In benchmarking experiments, Virchow demonstrated strong performance in pan-cancer detection and achieved high discriminative accuracy in several rare cancer types, with reported area under the receiver operating characteristic curve (AUROC) values approaching 0.95.Virchow2 extends this framework by adopting a larger ViT-H/14 vision transformer architecture and incorporating substantially expanded pre-training data.32 The updated model was trained on an additional ~3.1 million WSIs, encompassing multiple scanning magnifications (5×, 10×, 20× and 40×), thereby enhancing multiscale representation learning. In large-scale comparative evaluations, Virchow2 achieved the highest overall performance across datasets such as TCGA, Clinical Proteomic Tumour Analysis Consortium (CPTAC) and multiple external cohorts when assessed against 31 PFMs over 41 downstream tasks, including morphological classification, biomarker prediction and prognostic modelling. Earlier benchmarking studies reported slightly lower relative rankings,33 and subsequent analyses have identified residual performance degradation in certain settings despite the scale of pre-training data.Vision–Language foundation models in pathology Vision-language models (VLMs) bridge the fundamental gap between visual morphology and clinical semantics by aligning histopathology images with natural language descriptions. This group of models are typically pre-trained on massive datasets of image-caption pairs curated from medical literature and diagnostic report archives ( table 2). Technically, these models employ a dual-encoder architecture consisting of a vision encoder (eg, ViT) and a text encoder (eg, Bidirectional Encoder Representations from Transformers (BERT) or Robustly Optimised BERT Pretraining Approach (RoBERTa)). Modality-specific encoders map images (eg, WSI tiles) and text (eg, pathology reports) into a shared embedding space via projection heads. The model is trained using a contrastive loss (eg, InfoNCE), which maximises the scaled cosine similarity between matched image–text pairs while minimizing similarity across mismatched pairs within a batch, thereby aligning semantically related representations across modalities. During training, text is processed through subword tokenisation and masked language modelling to capture complex medical syntax. This cross-modal alignment allows the model to associate specific visual patterns, such as glandular crowding, with their corresponding medical terminology. The integration of language enables transformative capabilities, including zero-shot classification, where a model can identify novel disease states via text prompts without task-specific fine-tuning. Furthermore, VLMs facilitate cross-modal retrieval and visual question answering, providing an interactive interface where a pathologist can query the model regarding specific histological features through natural language.CONCH CONCH 34 is a multimodal vision–language foundation model trained on approximately 1.17 million histopathology image–text pairs using contrastive, task-agnostic pretraining. The training corpus integrates paired data from publicly available sources, including the PubMed Central Open Access subset, TCGA, the CPTAC and manually curated biomedical image–caption datasets. Architecturally, CONCH follows a dual-encoder paradigm, consisting of a visual encoder and a text encoder trained to project images and corresponding captions into a shared embedding space, enabling cross-modal retrieval and representation learning. Unlike generative multimodal models, CONCH does not rely on a fusion decoder but instead uses contrastive objectives to align modalities.By leveraging natural-language supervision, CONCH captures semantically rich representations that extend beyond purely visual features, facilitating improved performance across a range of downstream tasks, including classification, retrieval and weakly supervised prediction. In a recent large-scale benchmarking study comparing multiple PFMs, CONCH demonstrated among the highest overall performance across diverse evaluation tasks, highlighting the utility of multimodal pre-training in CPath.33TITAN TITAN 29 is a multimodal, whole-slide foundation model pretrained on 335 645 WSIs spanning 20 organs and tissue types, accompanied by 182 862 matched pathology reports and 423 122 synthetic captions generated using multimodal generative AI tools for pathology. The model learns slide-level representations through a combination of self-supervised visual learning and vision–language alignment, enabling the extraction of robust embeddings from ultra-large WSIs without relying on computationally intensive downstream multiple instance learning frameworks. By integrating visual features with textual information, TITAN can generate general-purpose slide representations and produce automated pathology reports, facilitating AI-assisted diagnostic documentation and seamless integration into clinical workflows. A notable strength of TITAN is its adaptability to resource-limited clinical settings and rare disease scenarios, where annotated data are scarce, allowing for accurate slide-level inference and report generation even in challenging or low-data contexts.Specialised and multimodal pathology foundation models The current frontier of the field lies in specialised multimodal foundation models, which move beyond simple image–text pairs to integrate a comprehensive spectrum of clinical and molecular data ( table 3). These models aim to replicate the holistic diagnostic process by synthesising histopathology with genomics, transcriptomics and longitudinal electronic health record (EHR) data. By employing fusion techniques, such as cross-attention mechanisms or multimodal embedding alignment, these frameworks can jointly process heterogeneous data types ranging from polygenic risk scores to detailed patient histories. Models like Multimodal Self-TAught PRetraining (mSTAR) and EXAONE Path 2.5 exemplify this approach, combining whole-slide histology with molecular and clinical data to generate integrated patient-level representations capable of supporting complex tasks such as disease subtyping, biomarker prediction, treatment response forecasting and survival estimation. These specialised models function as integrative diagnostic engines, addressing the primary limitation of earlier iterations by contextualising tissue morphology within the broader clinical and biological profile of the patient. Consequently, they move computational pathology closer to the realisation of personalised medicine, where diagnostic and therapeutic decisions are informed by the totality of a patient’s molecular, histological and clinical data.mSTAR mSTAR is a multimodal PFM designed to integrate complementary clinical information beyond histology alone. 35 Rather than relying solely on vision-only pre-training, mSTAR incorporates three distinct data modalities: whole-slide H&E images, expert-written pathology reports and gene expression profiles within a unified self-supervised framework. Its pre-training dataset comprises 26 169 slide-level multimodal pairs across 32 cancer types, representing more than 116 million pathological patch images, curated from TCGA and other sources. The architecture combines a slide-level contrastive learning stage, which aligns representations across modalities, with a subsequent patch-level ‘self-taught’ training stage, where multimodal context learnt at the slide level is propagated into the patch feature extractor, enabling comprehensive whole-slide and multimodal representations. Across a spectrum of 97 oncological tasks covering pathological diagnosis, molecular prediction, report generation, survival prediction, multimodal fusion and zero-shot classification, mSTAR outperforms state-of-the-art vision-only models, demonstrating that integrating multimodal clinical data can substantially enhance foundation model performance without requiring much larger vision-only datasets.EXAONE Path 2.5 EXAONE Path 2.5 represents an advanced multimodal PFM that explicitly integrates histologic and multi-omics data to capture a richer biological representation of cancer than image-only approaches. 36 Unlike traditional vision-only architectures, EXAONE Path 2.5 jointly models WSIs alongside genomic, epigenetic and transcriptomic datasets, generating unified representations that reflect tumour biology across morphological and molecular layers. Its architecture incorporates three key innovations: a multimodal sigmoid loss for language-image pretraining (SigLIP) that enables all-pairwise contrastive alignment across heterogeneous data types, a fragment-aware rotary positional encoding module that preserves spatial structure and tissue topology within WSIs and domain-specialised internal foundation encoders for both WSI and RNA sequencing data that produce biologically grounded embeddings for robust multimodal alignment. Trained on a multimodal cohort of 23 099 patients with paired imaging and omics measurements, EXAONE Path 2.5 demonstrates high data and parameter efficiency, achieving competitive or superior performance on both internal clinical benchmarks and the public Patho-Bench suite covering 80 tasks when compared with leading unimodal and multimodal pathology models. These results highlight the value of biologically informed multimodal design in linking genotype to phenotype for next-generation precision oncology.Advantages and capabilities of foundation models Advantages of foundation models in pathology The advantages of PFMs are: (1) Scalability and Transferability: Because the large diverse datasets with whole slide labels are used to train models, it can be adapted to many different pathology tasks—reducing the need for expensive, labour-intensive manual annotation for every new task. (2) Versatility: From basic tissue classification to complex cancer subtyping, PFMs provide a unified ‘backbone’. This can standardise workflows and accelerate development of CPath tools. (3) Efficiency in Low-Data or Rare Disease Settings: Few-shot/low-label performance means it can generalise to rare diseases, uncommon tissue types or resource-limited settings where labelled data are scarce. (4) Bridging Research and Clinical Use: Slide-level capability and resolution-agnostic behaviour bring the technology closer to real-world pathology applications, where WSIs and varying resolutions are the norm.Multimodal image intelligence: pathology and radiology AI Recent work in AI increasingly frames cancer diagnosis as a multimodal problem, where radiology AI and pathology AI provide complementary information rather than competing solutions. Radiology AI provides macroscopic, in vivo tumour characteristics such as spatial extent, heterogeneity and temporal evolution, while pathology AI offers microscopic features related to cellular morphology, tissue architecture and tumour microenvironment from digitised histology slides. The reviewed benchmarking and foundation-model studies emphasise that integrating these modalities can improve cancer detection, subtyping, prognostic stratification and biomarker prediction by linking whole-organ imaging phenotypes with cellular-level tissue patterns. 37Within this framework, PFMs such as Virchow and UNI serve as key enablers of multimodal approaches by learning robust, generalisable tissue representations that can be integrated downstream with radiologic, clinical and molecular data. Multimodal deep learning models have been developed that jointly analyse radiologic imaging (eg, CT or MRI) and digitised histopathology to improve tumour grading, survival prediction and treatment response assessment, consistently outperforming unimodal approaches. Additional work has focused on radiology pathology registration, aligning histologic ground truth with in vivo imaging to enhance tumour localisation and imaging interpretation. Similarly, general vision foundation models developed in radiology serve as complementary encoders of macroscopic imaging phenotypes. Although the reviewed studies do not implement end-to-end radiology and pathology fusion systems, they articulate a clear translational direction toward interoperable AI ecosystems in oncology, where modality-specific foundation models act as shared building blocks for integrated cancer characterisation. Collectively, the literature suggests that combining macroscopic radiologic features with microscopic tissue representations enables more comprehensive characterisation of tumour biology, while remaining positioned primarily as research and decision-support tools rather than standalone diagnostic systems.From foundation models to agentic AI systems The pathology foundation models and agentic AI The PFMs and agentic systems are complementary in that PFMs provide a generalisable representation layer, while agentic frameworks enable task execution and system-level interaction ( figure 2). In this architecture, the foundation model—whether vision-only or multimodal—functions as a feature encoder and semantic interpreter, transforming high-dimensional histologic imagery and associated clinical text into structured representations. These representations capture morphological patterns, contextual relationships and cross-modal associations learnt through large-scale pre-training.Figure 2Complementary relation between pathology foundation models and agentic AI. AI, artificial intelligence.However, PFMs are inherently passive: they do not initiate actions or maintain goals. Agentic systems operationalise these models by embedding them within a decision-making loop that includes state tracking, task planning and tool use. In this context, the agent queries the PFM to extract clinically relevant features (eg, tumour morphology, biomarker expression or spatial patterns) and integrates these outputs with external systems such as laboratory information systems (LIS), digital slide repositories or reporting platforms. Furthermore, agentic AI takes this integrated information to actions—such as case triage, ancillary test recommendation or report generation—thereby bridging the gap between representation learning and executable clinical workflows.From passive prediction to active autonomy A critical distinction in this evolution is the transition from reactive, single-step predictions to proactive, multistep autonomy. Traditional AI models operate on a linear ‘input-to-output’ basis—for example, receiving a tissue patch and returning a probability score for malignancy. In contrast, agentic pathology AI is goal-oriented rather than input-driven. When tasked with a broad objective, such as ‘confirm the primary site of this metastatic carcinoma’, an agentic system does not simply provide a label. Instead, it formulates a multistage plan: it may first use a vision–language model to identify the most representative regions of interest, autonomously trigger a virtual staining module to assess specific protein expressions and finally cross-reference these findings against a pan-cancer database. This capacity for self-correction and iterative refinement allows the system to resolve diagnostic ambiguities that would typically stall a conventional foundation model.Clinical use cases for agentic systems in pathology The deployment of agentic systems introduces transformative efficiencies across the pathology workflow, extending far beyond simple image analysis. In the prediagnostic phase, agents can act as autonomous triage engines, identifying critical cases—such as transplant rejection or necrotizing fasciitis—immediately on slide digitisation and escalating them to the top of a pathologist’s worklist. 13 14 During the diagnostic phase, agentic AI functions as a collaborative assistant capable of automated quality control; it can detect technical artefacts like tissue folding or poor staining and autonomously order a re-scan or re-stain before a human ever views the case. Furthermore, in precision oncology, these systems can synthesise complex multimodal data to assist in clinical trial matching. By autonomously ‘reading’ both the histological slide and the patient’s genomic profile, the agent can identify candidates for targeted therapies and draft a comprehensive integrative report, significantly reducing the cognitive load on clinical teams and accelerating the time to treatment.Regulatory and governance frameworks Several international organisations, including the WHO and the Organisation for Economic Co-operation and Development (OECD), have developed guidelines and recommendations for AI governance. WHO provides non-binding but globally influential guidance and recommends six core principles: protect human autonomy; promote human well-being and safety; ensure transparency and explainability; foster responsibility and accountability; ensure inclusiveness and equity; and promote responsive and sustainable AI (WHO, Ethics & Governance of Artificial Intelligence for Health, 2021). Similarly, OECD developed a global reference framework, representing the first international standards for trustworthy AI, which have been adopted by the European Union (EU), USA and 71 other jurisdictions. The five principles recommended by OECD align closely with WHO’s and include inclusive growth, sustainable development and well-being; human-centred values and fairness; transparency and explainability; robustness, security and safety; and accountability (OECD, AI Principles, 2024). This framework embodies the vision of a ‘Good AI Society’, in which the development, governance and use of AI benefit humanity, respect human rights and minimise harm.38In the EU, AI products are regulated under the Medical Device Regulation (MDR) and In Vitro Diagnostic Regulation (IVDR), as well as the EU AI Act (2024).39 Most radiology and pathology AI products for diagnosis, prognosis and prediction are classified as Class IIa–III under MDR or Class C/D under IVDR. Compliance requires demonstration of clinical evidence and performance evaluation, implementation of a quality management system and post-market surveillance and vigilance, along with CE marking via a notified body.40 Under the EU AI Act, medical AI systems used for diagnosis, prognosis or treatment decisions are considered high-risk, necessitating comprehensive documentation and risk management, data governance and bias mitigation, representative datasets, technical documentation and traceability, human oversight and post-market monitoring including incident reporting.41 Additionally, the EU AI Act permits in-house AI or hospital/laboratory-developed AI (not for commercial distribution), provided performance and safety requirements are met.In the USA, medical AI systems are broadly categorised as commercially available products or laboratory-developed tests (LDTs), both regulated by federal/state agencies and overseen by non-federal organisations. To be marketed for clinical use, AI tools require the FDA authorisation. The FDA has published a framework for Software as a Medical Device (SaMD) to support the development of innovative, safe and effective medical devices (FDA, Global Approach to Software as a Medical Device, 2022). AI systems without FDA authorisation may only be marketed as research-use only (RUO), and laboratories using RUO products must follow LDT regulations. Laboratories are responsible for establishing clinical performance metrics, including sensitivity, specificity and accuracy, in accordance with the Clinical Laboratory Improvement Amendments of 1988 (CLIA88) (federal regulation agent) and standards from the College of American Pathologists (CAP) (non-federal organisation). CAP requirements include rigorous pre-implementation validation, ongoing monitoring and integration into existing quality systems. Even FDA-authorised AI tools must be validated in the local laboratory context to ensure performance, clinical relevance and mitigation of data bias, with pathologists overseeing digital workflows in their role as CLIA directors. Additional guidance is provided by professional associations, including the American Society for Clinical Pathology and the Association for Pathology Informatics. Readers can refer to several recent reviews for further information.40 42 43Discussion Limitations of pathology foundation models Despite rapid advances in computational pathology, PFMs exhibit significant limitations related to scalability, generalisability, data bias, interpretability, weak clinical endpoint linkage, workflow integration and interoperability, clinical validation benchmarks, etc. Therefore, task-specific supervised models may still play important roles in clinical settings.PFMs impose substantial computational demands, including pre-training on gigapixel WSIs requires large-scale Graphics Processing Unit (GPU) infrastructure, and inference remains resource-intensive due to patch-based and multiscale processing, therefore, heavily relying on the availability of resources. Generalisability is further limited by dataset heterogeneity. Training data aggregated across institutions introduce confounding non-biological signals, and models trained on resources such as TCGA may encode site-specific artefacts rather than disease biology. Demographic imbalances can also produce performance disparities across patient groups,44 while sensitivity to batch effects and institutional variation persists.45 46Most PFMs rely on self-supervised or weakly supervised learning47, which may capture spurious correlations (‘deceptive learning’) rather than clinically meaningful morphology.48 Finally, limited interpretability of learnt representations, despite biologically relevant embeddings49 50, and variability across models,51 remains a barrier to clinical trust and regulatory adoption.Integration of foundation models into clinical workflows remains a significant challenge. Current PFMs are not yet validated for clinical use, and further studies are needed to evaluate their clinical performance metrics, such as sensitivity, specificity and accuracy, in addition to common metrics used in model development (AUROC, F1, concordance, etc).45 Quality control of PFMs in clinical environment needs to be implemented to identify temporal domain shift and performance drift over time,Furthermore, the existing isolated LIS and image management system (IMS) hinder efficient integration and limit the practical application of the broad capabilities offered by foundation models. To address this gap, our team developed FlexLIS, an integrated pathology information system that provides a potential solution for streamlined deployment and utilisation of foundation models in routine clinical practice. FlexLIS is an integrated platform for LIS, whole-slide IMS and AI to work together, where LIS organises clinical information, IMS manages the images and AI communicates between LIS and IMS. In FlexLIS, AI models and foundation models can retrieve clinical information from LIS which is interfaced with EHR and analyse images which are stored in IMS, therefore, allowing language and vision models to work together in the clinical settings.13 14Regulatory and translational challenges Current regulations, guidelines and standards were primarily developed for task-specific AI systems. Given that foundation models are general-purpose and have broader applications, tailored guidance and updated standards are needed to govern their uses. In response, WHO recently released new guidance ( Ethics and Governance of Artificial Intelligence for Health: Guidance on Large Multi-Modal Models, 2024), which provides over 40 recommendations for governments, technology companies and healthcare providers to ensure the responsible development and deployment of large multimodal models, with the goal of promoting and protecting population health.For clinical implementation of foundation models, it is far more complicated and challenging due to lack of clear regulatory frameworks. Therefore, it is important to follow current laboratory standards8 and newly developed recommendations including: (1) establishing scope and intended use of AI models (screen, assist diagnosis, assist workflow, admin, etc) which determine the classification (SaMD vs non-device admin tools) and scope of validation and regulatory compliance; (2) developing policies covering accountability (pathologist roles), transparency (limitation and scope of use), safety and security (HIPAA, cybersecurity and data protection) and downtime procedures; (3) completing validation which includes technical performance (sensitivity, specificity, and related metrics), clinical relevance (meet the criteria of intended uses) and identification of AI related artefacts (hallucinations, omissions and wrong answers); and (4) establishing quality control and quality assurance procedures, including monitoring input drift, output quality, error captures/incidence report, update controls (model/version changes and re-validation protocol) and total product lifecycle as defined by the FDA (FDA, Predetermined Change Control Plans for Machine Learning-Enabled Medical Devices: Guiding Principles. 2025).2Conclusion We expect that many new PFMs and platforms will be developed to address the limitations discussed above. Given that current foundation models are not ready for clinical use, future study will focus on the improvement of foundation models and integration into clinical workflows such as assisting cancer screening and diagnosis (especially rare cancer types), risk stratification (prognosis) and predicting treatment response. Finally, with the maturity of agentic AI, the agentic pathology systems will evolve as central hubs in participating in patient care with unprecedented accuracy and efficiency, a quality that is critically important in a resource-limited world.