BetaEntity Annotation Prototype
← Back to treatments

Annotated full text

Beyond auto-segmentation: the case for planning and dosimetry AI in head and neck radiation oncology

bmjonc · 2026-06-29 · canonical JSON source

39 visible annotations · policy: published · automated confidence ≥ 75.00%

Document resource

Introduction Artificial intelligence (AI) has been rapidly integrated across multiple stages of the radiation oncology workflow driven by advances in deep learning and increasing amounts of large, annotated imaging data sets. In clinical practice, AI-enabled tools support workflow automation across radiotherapy, including automated segmentation of target volumes and organs at risk (OARs), treatment planning optimisation, and quality assurance processes. This reflects a broader transition towards data-driven radiotherapy, in which complex imaging and clinical data are leveraged to guide contouring, plan generation and evaluation. 1 To date, much of the progress in the space has focused on auto-segmentation of OARs and more recently target volumes.2–8 This emphasis is practical as contouring is lengthy and subject to interobserver variability. Various studies have shown that auto-segmentation can improve efficiency and standardisation, justifying its integration into commercial treatment planning systems with clinician oversight.9Head and neck radiotherapy represents a unique and complex clinical domain. Treatment planning requires steep dose gradients across many critical structures within a confined space. Small variations in planning decisions can have downstream effects on treatment-related toxicity. In addition to significant advancements in auto-segmentation, further improvements in workflow may impact clinical outcomes. In head and neck radiotherapy, interplanner variability persists even when contours are standardised. Additionally, treatment plan quality remains highly dependent on planner expertise and optimisation strategy. These limitations suggest that although auto-segmentation is a critical prerequisite, advances in AI for treatment planning and dosimetry optimisation may offer improvements in both clinical outcomes and system-level performance.Progress of auto-segmentation: from bottleneck to enabling infrastructure The rapid progress of auto-segmentation represents a tangible success of AI in radiation oncology. Over the past decade, advances in deep learning architectures, publicly annotated imaging data sets and multi-institutional validation efforts have transformed automated delineation from a research concept to a clinical tool. Commercial deep-learning segmentation tools as well as home-grown systems have been incorporated into departmental workflows after multidisciplinary evaluation, producing contours that often require only minor adjustment before use in planning. 10 The level of clinical confidence is further reflected in their use within prospective, multicentre clinical trials, where automated contouring has been leveraged not merely as a convenience tool but as a mechanism for standardised practice.In the TROG V.18.01 NINJA trial, an automated contour quality assurance tool for prostate clinical target volume was used to reduce interobserver variability and manual review burden.11 In a multivendor benchmarking study, Doolan et al demonstrated that five commercial auto-segmentation systems achieved good geometric agreement with expert contours (median volumetric Dice similarity coefficients ranging from 0.82 to 0.88) while providing substantial time savings, particularly in head and neck cases where contouring time was reduced by approximately 74–93 min per patient.12 The HaN-Seg challenge demonstrated robust multi-institutional performance of head and neck OAR segmentation across CT and MRI, supporting generalisability across data sets and imaging platforms.4 Complementing this, recent real-world implementation studies such as one conducted by Young et al have shown that AI-assisted contouring can achieve favourable dosimetric acceptability and clinician uptake in routine workflows.13 Together, these studies extend the role of auto-segmentation beyond workflow acceleration towards clinically scalable deployment.Consistent with this trial-level validation, widespread real-world adoption has also been reported. For example, survey data from German-speaking countries indicate that a majority of radiation oncologists and medical physicists report using AI-based auto-contouring in routine clinical practice.14 Approximately two-thirds of respondents reported current use and 90% reported the tools as helpful.14 In parallel, the scope of automated contouring has expanded substantially within head and neck radiotherapy, evolving from limited OAR delineation to contouring of complex anatomical structures, nodal levels and target volumes.15Several factors explain the speed and scale of this progress in auto-segmentation. Contouring is a universal prerequisite of head and neck radiotherapy workflows, as reflected in international consensus guidelines that standardised target and OAR delineation across disease subsites and treatment techniques.16 From a methodological standpoint, segmentation is well suited for early AI development because segmentation is a bounded task with clearly defined inputs and outputs. This structure enables objective performance evaluation using standardised geometric metrics, such as Dice similarity and surface distance, and has facilitated benchmarking, comparison across studies, and rapid iteration and validation supported by large data sets.17 18 Clinically, the contouring process is both time-intensive and repetitive, creating an immediate and easily quantifiable opportunity for efficiency gains. Multiple clinical implementation studies have demonstrated substantial reductions in contouring time, allowing institutions to directly measure workflow and productivity benefits.19 20 Importantly, auto-segmentation outputs remain subject to clinician review prior to treatment delivery, thus allowing for lower regulatory and safety barriers to adoption.21 Together, these factors explain why auto-segmentation has emerged as the most mature and widely adopted application of AI in radiation oncology, while also highlighting why its success does not directly translate to more complex planning and dosimetric decision-making tasks.Radiomics and AI-enabled prognostication in head and neck cancer Beyond geometric automation, parallel advances in AI have focused on extracting prognostic information from imaging data through radiomics. In head and neck cancer, radiomic analyses have been explored as predictors of treatment response, locoregional control and radiation-induced toxicity, with applications like the prediction of mucositis in patients receiving volumetric modulated radiotherapy. 22 These approaches aim to capture imaging features that reflect tumour biology or normal tissue susceptibility.23 24 These efforts extend AI from workflow optimisation towards individualised risk stratification. More recently, AI-based prognostication in head and neck oncology has expanded beyond handcrafted radiomic features to include deep learning approaches applied to digital pathology and multimodal data to predict outcomes such as recurrence and survival. For example, Tian et al developed a multimodal deep learning model integrating CT, whole-slide histopathology and clinical features from 1087 patients with head and neck squamous cell carcinoma to predict prognosis and postoperative radiotherapy response.25 Emerging studies have demonstrated that integrating imaging, histopathology, genomic and clinical features can improve survival prediction in oral squamous cell carcinoma.26 Although these approaches are not yet used clinically on the same scale as prostate cancer prognostic assays such as the guideline-endorsed ArteraAI Prostate Test (Artera, Los Altos, CA, USA), they highlight the growing role of AI in predicting biological behaviour and radiosensitivity.27Despite this promise, radiomics-driven prognostication in head and neck radiotherapy remains limited in clinical impact. Many models are retrospective, sensitive to imaging acquisition parameters and difficult to generalise across institutions. For example, radiomics features extracted from pretreatment imaging have been shown to predict the risk of severe radiation-induced xerostomia following intensity-modulated radiotherapy (IMRT).28 These findings remain largely correlative and have not been prospectively translated into treatment decision algorithms. Broader reviews have similarly highlighted methodological limitations related to variability in imaging protocols, limited external validation and reproducibility challenges in head and neck radiomic studies.29 As a result, radiomic predictions are often generated in isolation, identifying elevated risk without directly informing how treatment planning should be modified.These limitations underscore a critical gap between prognostic insight and actionable intervention. Recent work illustrates how this gap can be closed in practice. Van den Bosch et al developed comprehensive normal tissue complication probability (NTCP) profiles that map dose across 14 OARs to 22 specific toxicity endpoints, quantifying each patient’s individual risk landscape prior to planning.30 Van der Laan et al then showed these NTCP profiles can be directly used as optimisation objectives, and their quality-of-life-weighted volumetric modulated arc therapy planning approach reduced predicted dysphagia risk by 6%–7% by redistributing dose away from high-risk structures, without compromising target coverage.31 This sequence, from patient-specific risk quantification to dose redistribution guided by that risk, demonstrates the mechanism through which prognostic information can modify treatment planning rather than remain an isolated prediction. Prognostic signals derived from imaging, pathology or clinical features are most impactful when coupled to systems capable of operationalising them within the treatment planning process. In this context, planning and dosimetry AI provide a natural interface for integrating prognostic information, enabling risk-aware optimisation strategies that adapt dose distributions to patient-specific anatomy and biological vulnerability.32 Such integration offers a pathway for moving AI-based prognostication beyond risk prediction towards clinically meaningful personalisation of head and neck radiotherapy.Clinical leverage of planning and dosimetry AI in head and neck radiotherapy In head and neck radiotherapy, clinical outcomes and long-term quality of life are driven by geometric accuracy and by nuanced dosimetric trade-offs across many critical structures. Even when target volumes and OARs are perfectly delineated, substantial interplanner variability often persists. 33 This variability is due to differences in optimisation priorities, constraint weighting, and compromise strategies that meaningfully affect dose–volume distributions and downstream toxicity. Head and neck cancer is particularly sensitive to subtle dosimetric differences. Small changes in mean dose to the parotid glands, pharyngeal constrictor muscles, laryngeal structures, or oral cavity can affect long-term functional outcomes.34 Planning and dosimetry AI can address this challenge by systematically navigating complex and multiobjective spaces to develop consistent and high-quality plans that balance tumour control with organ preservation.35In addition, planning AI may also facilitate real-time evaluations of the dosimetric ‘costs’ associated with specific planning decisions, such as more generous clinical target volumes or larger planning target expansions. Specifically, planning AI can model how incremental changes in target volume definition translate to measurable increases in dose to adjacent OARs, enabling clinicians to weigh the oncologic benefit of broader coverage against the expected toxicity cost before committing to a final plan.36 Commercial tools such as Pinnacle Auto-Planning (Philips Healthcare, Fitchburg, WI, USA), Varian RapidPlan and Ethos Intelligent Optimisation Engine (IOE) (Varian Medical Systems, Palo Alto, CA, USA) and RayStation multicriteria optimisation (RaySearch Laboratories, Stockholm, Sweden) have demonstrated the feasibility of automated or semiautomated planning in complex head and neck cases. In a multi-institutional comparison of five automated treatment planning solutions, Krayenbuehl et al showed that multiple commercial auto-planning systems achieved clinically acceptable target coverage and OAR sparing while substantially reducing effective planning time, supporting the maturity of these tools for clinical deployment.37 Beyond knowledge-based and rule-based automation, newer generative AI approaches such as GPT-RadPlan suggest a further evolution towards reasoning-enabled planning systems capable of interfacing with commercial treatment planning environments to automate optimisation decisions.38In adaptive radiotherapy settings for head and neck cancer, AI-guided reoptimisation has been shown to consistently reduce dose to multiple OARs, including the parotid glands, submandibular glands, oral cavity and larynx, while maintaining comparable target coverage.39 Similarly, studies employing machine-learning-driven optimisation engines within commercial adaptive platforms such as Varian Ethos demonstrate that clinically acceptable plans balance target coverage and OAR sparing without iterative manual replanning.40 Critically, not all patients with head and neck cancers benefit equally from adaptive workflows, making patient selection essential for scalable implementation. Mastella et al systematically reviewed how AI can automate the CT-based adaptive radiotherapy chain—from synthetic CT generation and OAR segmentation to predictive models that identify candidates most likely to benefit from adaptation, with reported accuracy exceeding 80%.41 Aristophanous et al further supported this with clinical data from a conventional C-arm linear accelerator-based offline adaptive radiotherapy programme, where their custom-developed Automated Watchdog for Adaptive Radiotherapy Environment system monitored volumetric changes using cone-beam CT and deformable image registration.42 The system identified patient-specific anatomical triggers for replanning, including a parotid volume reduction of 7% or greater and a nodal gross tumour volume decrease of 29% or greater, both of which were associated with clinically meaningful dosimetric improvements.42 These findings underscore that planning AI may support not only adaptive reoptimisation itself but also prospective identification of patients most likely to benefit from adaptive radiotherapy.Beyond improvements in plan quality and consistency, planning AI can substantially shorten planning timelines, enabling more streamlined and scalable radiotherapy workflows and facilitating earlier treatment initiation with potential clinical benefit for patients with head and neck cancer. This implementation underscores the potential of planning AI to support efficient and reproducible adaptive workflows. Collectively, these studies suggest that planning and dosimetry AI can improve consistency, reduce planning time and achieve clinically acceptable OAR sparing in selected cohorts. However, most published evaluations emphasise dosimetric endpoints, workflow efficiency, or blinded plan preference rather than prospective reductions in toxicity, improvements in patient-reported quality of life, or survival outcomes.39 40 43–45 Future prospective multicentre studies and real-world implementation analyses will be necessary to determine whether these technical gains translate into clinically meaningful patient benefit.Technical roadmap: pattern learning to reasoning-based planning AI In head and neck radiotherapy, where outcomes are driven by planning trade-offs rather than contouring accuracy alone, recent technical advances in planning AI outline a clear pathway beyond auto-segmentation. As illustrated in figure 1 (Evolution of Radiotherapy Planning), the development of planning AI has progressed from static, database-driven prediction models towards more adaptive and reasoning-oriented systems.46–49 Early efforts in improving outcomes in radiotherapy focused on knowledge-based planning, in which statistical models trained on databases of high-quality previously treated patient plans are used to estimate achievable dose–volume relationships and guide optimisation.50Figure 1Evolution of radiotherapy planning. AI, artificial intelligence; DVHs, dose–volume histograms; HN, head-and-neck; IMRT, intensity-modulated radiotherapy; KBP, knowledge-based planning; LLMs, large language models; NICE, National Institute for Health and Care Excellence. Created in BioRender. Choi, S. (https://BioRender.com/tj34of4) is licensed under CC BY 4.0.A key limitation of knowledge-based planning is that model performance is inherently dependent on the quality, consistency, and representativeness of the training plan library. Poorly curated or heterogeneous data sets may propagate suboptimal planning strategies rather than encode expert-level performance. This dependency has been demonstrated directly in head and neck radiotherapy, where iterative retraining of RapidPlan models using improved model-generated outputs produced stronger regression performance and reduced mean dose to parallel OARs such as the parotid glands, suggesting that model quality can improve through deliberate maintenance strategies rather than remain fixed at deployment.51 Importantly, this limitation may therefore be manageable rather than prohibitive, provided institutions invest in high-quality plan libraries, periodic retraining and ongoing model refinement.52 These considerations help explain why later approaches sought methods capable of moving beyond static database-derived priors.These approaches were subsequently extended by neural network-based methods capable of learning more complex spatial dose patterns and optimisation outcomes directly from data.53 While such models address the limitations of handcrafted features and improve consistency in dose prediction, they primarily operate as static predictors rather than decision-making agents.36However, deep learning–based dose prediction introduces its own limitations, particularly related to generalisability across institutions and patient populations. A recent systematic review and meta-analysis found substantial heterogeneity across studies (I² >99%), with performance varying by radiotherapy technique, network architecture, and disease subsite, highlighting the instability of results across settings.54 Similarly, external validation studies have demonstrated meaningful performance degradation when models are transferred across institutions, although transfer learning using small external data sets has shown partial restoration of model performance, suggesting a practical pathway for cross-institutional adaptation rather than a fundamental barrier to deployment.55 More recent architectures incorporating physics-informed priors have also demonstrated improved generalisability across tumour sites, including head and neck, without institution-specific retraining.56 These findings suggest that while generalisability remains a recognised limitation, it is increasingly being addressed through technical and model adaptation strategies.These limitations have also been recognised at the field level. The ESTRO-AAPM joint guideline noted that widespread clinical adoption of AI planning models remains constrained by concerns regarding availability, applicability, quality, generalisability, interpretability, and safety, with much of the validation literature remaining retrospective and single-institutional.43 At the same time, emerging multicentre evidence suggests this landscape may be evolving. In a three-institution evaluation, over 80% of AI-generated plans met clinical criteria and 60% were preferred over manual plans in blinded review, supporting the view that AI planning may already provide a high-quality foundation for clinician refinement, even if not yet a replacement for expert oversight.45An underexplored but clinically important concept in treatment planning is the degree of precision needed for contours. Prior work suggests that in some clinical contexts, similar dose distributions can be achieved across a range of target volume definitions, highlighting the importance of robust planning strategies. Ploquin et al compared intensity-modulated and three-dimensional conformal plans for oropharyngeal cancer under simulated setup errors and found that even the more geometrically complex IMRT plans maintained acceptable OAR doses despite positional perturbations, demonstrating that dosimetric stability is driven more by the underlying optimisation strategy than by millimetre-level contouring precision.57 Understanding and exploiting these tolerance limits represents an important opportunity for planning AI, shifting emphasis away from contour precision towards dose robustness, trade-off navigation and sensitivity analysis.Established multicriteria optimisation and Pareto navigation frameworks provide one example of how such uncertainty-aware planning can be operationalised in practice.46 47 For example, human-like intelligent automatic treatment planning systems incorporate a Virtual Treatment Planner, an AI agent designed to mimic expert dosimetrists by iteratively evaluating plan quality.58 Similarly, GPT-RadPlan evaluates dose distributions and dose–volume histograms, generates feedback on unmet objectives and iteratively modifies optimisation parameters to improve plan quality.38 As shown in figure 1, this concept shift reframes treatment planning as an iterative reasoning process rather than a single-shot prediction task. In this context, emerging approaches that incorporate large language models (LLMs) represent a further step in this evolution of planning AI. These systems are able to evaluate quality metrics, assess dosimetric trade-offs and guide optimisation adjustments in a manner analogous to expert human planners.59At present, however, reasoning-driven planning systems remain largely exploratory. Early multimodal planning agents have demonstrated promising feasibility but validation has generally been limited to small test cohorts, narrow disease settings and single-institution environments.38 Similarly, the NRG Oncology working group has emphasised that integrating AI into clinical trials presents ongoing methodological, regulatory, and implementation challenges requiring further development before widespread adoption.44 As such, reasoning-driven systems should presently be viewed as investigational decision-support concepts rather than clinically established planning technologies, with prospective evaluation, external validation, and regulatory clarity remaining active areas of development.Deep learning-based dose prediction and reasoning-driven planning assistance should be viewed as complementary rather than competing paradigms. Neural networks provide the quantitative foundation for modelling achievable dose trade-offs and spatial dose distributions, whereas LLM-based systems operate at a higher level of abstraction, supporting navigation of these trade-offs through higher-level reasoning.60 A near-term and clinically realistic application of reasoning-based planning AI lies in the use of AI agents as information synthesisers rather than autonomous decision makers.61 By integrating plan quality metrics, dose-toxicity relationships, prior cases, and institutional preferences, such systems could support clinicians in making more informed, transparent, and consistent planning decisions.59 62 Together, these developments outline a plausible and technically grounded pathway towards advanced planning AI capable of addressing the optimisation decisions that most directly influence outcomes in head and neck radiotherapy.Ultimately, these developments outline not a linear succession in which each approach replaces its predecessor but a layered evolution in which each generation addresses recognised limitations of the prior one while introducing new challenges of its own. This framing may more accurately position reasoning-driven planning AI not as a mature endpoint but as an emerging extension of a broader planning intelligence ecosystem.Designing safe and trustworthy planning AI As planning and dosimetry AI systems move closer to supporting clinical decision-making, their success will depend not only on technical performance but also on how effectively they are designed to operate safely, transparently and in partnership with human expertise. As outlined in figure 2 (Framework for Safe and Trustworthy Planning and Dosimetry AI), safety in planning AI requires an integrated framework that combines failure awareness, clinician oversight, uncertainty modelling, and structured recovery pathways. Unlike auto-segmentation errors, which are often visually apparent and readily corrected, failures in planning and dosimetric reasoning may be subtle yet clinically consequential.63 This distinction is particularly important in head and neck radiotherapy, where small deviations in dose distribution can meaningfully affect toxicity profiles and long-term quality of life.Figure 2Framework for safe and trustworthy planning and dosimetry AI. AI, artificial intelligence. Created in BioRender. Choi, S. (https://BioRender.com/svxqq07) is licensed under CC BY 4.0.One common failure mode is misprioritisation of competing objectives, such as overemphasising target coverage or a single OAR constraint at the expense of other clinically relevant structures. A second failure mode involves convergence to locally optimal but clinically undesirable solutions. These solutions happen when gradient-based or heuristic optimisation settles on solutions that are mathematically efficient yet misaligned with clinical intent or functional preservation goals.64 A further critical risk is overconfidence and silent failure, whereby AI systems generate recommendations without communicating uncertainty or recognising when they are operating outside their domain of reliability. In such cases, errors may propagate downstream without triggering clinician review.65These failure modes underscore the need for clinician-in-the-loop feedback mechanisms as a central design principle for safe and effective planning AI. By enabling clinicians to review, modify, and annotate AI-generated planning suggestions, such systems preserve clinical accountability while supporting bidirectional learning and continuous improvement. Complementary to this approach, uncertainty modelling and out-of-distribution detection are essential components of trustworthy planning AI, enabling systems to flag low-confidence recommendations, defer to human judgement when appropriate, and avoid overconfident extrapolation.66 These design principles are also aligned with emerging regulatory expectations for clinical AI, which emphasise human oversight, auditability and accountability in high-risk medical decision support systems. In addition, rigorous commissioning and clinical evaluation testing should be viewed as integral components of this safety framework, ensuring that system performance, failure modes, and boundary conditions are well characterised prior to clinical deployment.43 67Beyond failure detection, safe planning AI should incorporate explicit recovery pathways triggered by predefined signals.43 68 Such signals could be elevated predictive uncertainty, deviation from learnt population-level feasibility bounds or detection of out-of-distribution anatomies or constraint combinations.69 Recovery mechanisms may include automated rollback to previously validated plans, generation of Pareto-optimal solutions and structured escalation to clinician review when algorithmic confidence is low or competing objectives cannot be reliably reconciled automatically.70 Importantly, such recovery should be understood as a graceful degradation rather than system failure.A commonly cited concern surrounding advanced automation is the potential for clinical deskilling. In the context of treatment planning, however, this risk is not inevitable but instead reflects design choices. Planning AI that operates as a black-box optimiser may erode engagement, whereas systems that make trade-offs explicit, expose underlying rationale and require clinician arbitration can reinforce expertise rather than replace it.71 Human–AI collaborative frameworks that emphasise transparency, interpretability, and active clinician participation are therefore more likely to preserve skill development while improving consistency and safety in planning workflows.72 73At a broader systems level, the safe and equitable deployment of planning and dosimetry AI depends on the representativeness of the data and environments in which these systems are developed and validated. Broader analyses of clinical AI have demonstrated that training data sets, model development and authorship are disproportionately concentrated in high-income countries and large academic centres, raising concerns that AI systems may underperform in underrepresented populations or lower-resource settings.74 75 In radiation oncology, these risks may be amplified by variation in imaging protocols, contouring practices, treatment planning systems and available hardware across institutions and geographic regions. Without diverse external validation and adaptation to local workflows, planning AI may inadvertently perpetuate or exacerbate existing disparities in access to high-quality radiotherapy rather than reduce them. In addition, implementation in resource-limited settings may be constrained by limited digital infrastructure, fragmented data systems and lack of local technical expertise, even when algorithmic performance is strong. Addressing these structural and global limitations will be essential if planning AI is to fulfil its potential as a scalable and equitable component of modern radiotherapy delivery.System-level value and strategic importance Beyond individual patient outcomes, investment in planning and dosimetry AI offers system-level advantages across the radiotherapy care continuum. Whereas auto-segmentation primarily improves efficiency at a single upstream step, AI-driven treatment planning has the capacity to streamline the entire planning process. 76 77 By reducing the number of optimisation cycles, shortening time from simulation to treatment and improving the predictability of plan quality, AI-driven treatment planning has the potential to translate into meaningful improvements in workflow efficiency and departmental throughput. These benefits may be especially relevant in complex disease sites such as head and neck cancer, where planning is highly iterative and resource-intensive.78 In a multicentre evaluation across three institutions and five disease sites, a hybrid deep learning-based automated planning system integrating dose prediction with clinical-goal-guided optimisation generated directly deliverable plans within 5 min, with over 80% meeting institutional clinical criteria and 60% preferred over manually optimised plans in blinded review.45Planning and dosimetry AI may also reduce dependence on highly specialised local expertise. Head and neck treatment planning requires extensive experience to balance competing dosimetric objectives across numerous critical structures, and substantial interinstitutional and interplanner variability persists when standardised guidelines are followed.79 80 Chen et al found that beyond planning variability, interobserver differences in target volume and OAR delineation among radiation oncologists have been shown to reduce prescription dose coverage and increase doses to OARs, with relative volume differences in target delineation strongly correlating with decreased tumour control probability in nasopharyngeal carcinoma.80 By embedding expert-level planning strategies into automated or semiautomated optimisation frameworks, planning AI has potential to improve consistency and quality of care across diverse practice settings, thereby addressing inequities in access to high-quality head and neck radiotherapy.81From a strategic perspective, planning AI is also central to the scalability of emerging radiotherapy paradigms. Adaptive radiotherapy has been increasingly relevant in head and neck cancer due to pronounced anatomical and volumetric changes over the course of treatment. However, widespread adoption of adaptive workflows has been limited by the time and expertise required for repeated plan re-optimisation. AI-driven planning and dosimetry infrastructure enables rapid, reproducible reoptimisation while preserving target coverage and OAR sparing, positioning planning AI not merely as an efficiency tool but as an enabling infrastructure for next-generation radiotherapy delivery.82Towards integrated and predictive radiotherapy: digital twins and future directions Sustained investment in planning and dosimetry AI enables a broader reconceptualisation of radiotherapy as a predictive and adaptive process, rather than a static, one-time intervention. A long-term vision emerging from this trajectory is the concept of a digital twin: a continuously updated computational representation of a patient that integrates anatomy, delivered dose, biological response and functional outcomes over time.83 Although still aspirational, head and neck cancer represents a particularly compelling use case given the sensitivity of outcomes. Importantly, digital twins cannot arise from segmentation or prognostication alone; they depend on robust planning infrastructure capable of iteratively adjusting treatment. Chaudhuri et al demonstrated a predictive digital twin framework for high-grade gliomas that integrates longitudinal imaging with mechanistic modelling and Bayesian calibration to optimise radiotherapy regimens under uncertainty, illustrating the type of robust computational infrastructure needed to support adaptive and personalised treatment planning.84 In this context, planning and dosimetry AI function as enabling technologies, providing the decision-making substrate on which more comprehensive predictive models may eventually be built.71Complementing this vision, emerging evidence supports the use of unsupervised and self-supervised methods to reveal latent relationships between anatomy, dose and function that are not explicitly encoded in current planning paradigms. Such methods may reveal previously unrecognised patterns of toxicity or recovery, further informing adaptive and personalised treatment strategies. For example, unsupervised clustering of radiotherapy dose data has been shown to identify distinct patient subgroups with differing toxicity profiles that are not explained by standard dose–volume parameters or clinical grading alone.85Conclusion Planning and dosimetry AI have the potential to influence multiple aspects of head and neck radiotherapy due to their ability to synthesise complex clinical, anatomical and dosimetric information. Emerging applications of planning AI in head and neck cancer range from dose prediction and optimisation support to adaptive replanning and decision-support tools. When coupled with advances in multimodal learning and integration within compound AI frameworks, planning AI systems are increasingly positioned to address long-standing clinical and translational challenges.However, several challenges remain to the development of planning and dosimetry AI in routine clinical practice. These include ensuring reliability across heterogeneous patient anatomies, maintaining transparency and interpretability in optimisation decisions, safeguarding patient data and addressing ethical and medicolegal considerations associated with AI-assisted decision-making. Validation of planning AI systems therefore requires demonstration of meaningful benefit in real-world clinical settings, with evaluation frameworks that extend beyond geometric accuracy. While current evidence does not support the replacement of clinician expertise in head and neck treatment planning, AI may serve as a valuable decision-support tool under clinician oversight, augmenting rather than supplanting human judgement in complex radiotherapy workflows.Ultimately, the value of planning AI lies not in replacing clinical expertise but in operationalising it; translating the dosimetric trade-offs, prognostic insights and adaptive strategies discussed throughout this review into reproducible, efficient workflows. By compressing planning timelines from days to minutes and enabling rapid reoptimisation, these systems can shorten time to treatment initiation while freeing clinical staff to focus on plan review, quality assurance and patient care. This consistency also creates a foundation for meaningful quality assurance metrics, because when plans are generated through uniform optimisation frameworks rather than variable manual approaches, institutions can establish benchmarks, track performance across patients and planners, and identify deviations that warrant review. In this way, planning and dosimetry AI can tie together the advances in segmentation, prognostication and adaptive delivery discussed in this review, moving head and neck radiotherapy closer to a more consistent, accurate, accessible and individualised standard of care.