Abstract
Artificial intelligence (AI) is rapidly expanding across the radiation oncology workflow, with applications spanning imaging, contouring, treatment planning, quality assurance, outcome prediction, workflow automation, and clinical decision support. Although technical progress has accelerated substantially, successful clinical translation remains inconsistent. Many of the challenges limiting implementation are not unique to radiation oncology and have previously emerged across healthcare and other high-stakes industries. In this narrative review, we examine radiation oncology AI through the broader lens of cross-industry AI development and deployment. We first summarize the current landscape of AI applications in radiation oncology and then analyze representative examples of successful and unsuccessful AI implementation from healthcare and other sectors. These experiences reveal recurring themes that strongly influence clinical AI success, including data representativeness, robust validation, workflow-centered design, human-AI collaboration, uncertainty management, bias mitigation, operational boundaries, and continuous performance monitoring. We discuss how these lessons apply directly to radiation oncology, where AI systems must function within complex clinical workflows involving imaging, planning, adaptive treatment, quality assurance, and longitudinal patient management. Emerging agentic and multimodal AI systems further amplify both opportunities and risks associated with deployment. Ultimately, the future impact of AI in radiation oncology will likely depend less on isolated algorithmic performance than on the development of trustworthy clinical AI ecosystems. Successful implementation will require rigorous validation, seamless workflow integration, human oversight, regulatory governance, and continuous adaptation. Lessons from healthcare and other industries suggest that the greatest and most durable clinical value may arise from AI systems that augment human expertise, cognitive workflows, and multidisciplinary decision-making rather than replace clinical decision-makers.
Introduction
Artificial intelligence (AI) is increasingly reshaping modern radiation oncology. Rapid advances in machine learning, deep learning, multimodal data integration, large language models (LLMs), and computational infrastructure have enabled AI systems to participate in nearly every stage of the radiotherapy workflow. Current applications include image synthesis and enhancement, auto-segmentation, treatment planning optimization, outcome prediction, quality assurance (QA), workflow automation and orchestration, and clinical decision support ().
The growing interest in radiation oncology AI reflects both technological opportunity and clinical necessity. Yet experience from other industries suggests that technical capability alone rarely guarantees successful adoption. Radiation oncology workflows are highly data-intensive, operationally complex, and increasingly dependent on integration across imaging, planning, delivery, adaptation, and longitudinal outcome assessment (, ). Simultaneously, escalating treatment complexity, workforce pressures, and demands for personalization continue to challenge conventional clinical paradigms. AI, therefore, offers the potential not only to automate repetitive tasks but also to augment decision-making, improve efficiency, enhance consistency, and ultimately enable more adaptive and biologically informed radiotherapy paradigms ().
Figure 1 illustrates the broad penetration of AI across the radiation oncology care continuum, spanning imaging, contouring, treatment planning, treatment delivery, treatment adaptation, QA, and outcome assessment. Current applications already include synthetic imaging, auto-segmentation, dose prediction, toxicity modeling, adaptive planning, and workflow automation. However, despite substantial progress, many efforts remain fragmented and are variably integrated across the clinical workflow (). This has also been a persistent disparity between AI innovations developed and validated at large academic centers and the broader adoption, implementation, and scalability of AI tools across diverse real-world clinical radiation oncology practices (). Importantly, the central challenge facing radiation oncology is no longer whether AI can perform individual technical tasks. Rather, the emerging challenge is whether and how AI can be deployed safely, robustly, and meaningfully within real-world clinical ecosystems.
Figure 1
Much of the existing literature in radiation oncology AI focuses on model architectures, benchmark performance metrics, or disease-specific applications (). Although valuable, these discussions often underemphasize broader translational issues, including workflow integration, human oversight, bias propagation, model drift, clinician trust, regulatory governance, and long-term adaptation. These are precisely the challenges that have already emerged across numerous AI-enabled industries.
Outside of radiation oncology, and indeed across medicine and other high-stakes industries, AI systems have experienced both transformative successes and highly visible failures (). Transformative examples include AlphaFold, which revolutionized protein structure prediction, and Paige AI, which successfully translated foundation-model pathology into FDA-cleared clinical deployment (, ). In contrast, IBM Watson for Oncology struggled to deliver reliable treatment recommendations in real-world practice, while Amazon’s recruiting algorithm reproduced historical hiring biases embedded within its training data (–). Some platforms improved operational efficiency and clinical outcomes, whereas others collapsed under real-world complexity despite impressive developmental performance. These experiences provide a valuable opportunity for radiation oncology to learn from prior successes and mistakes rather than repeating them independently.
Unlike prior reviews that primarily catalog AI applications across individual stages of the radiation oncology workflow, this narrative review examines radiation oncology AI through the broader perspective of cross-industry AI development and deployment. Drawing on successes and failures from healthcare and other high-stakes industries, we identify recurring challenges involving data quality, workflow integration, human oversight, bias propagation, uncertainty management, regulatory governance, and long-term adaptation. We then synthesize these lessons into a practical framework for developing trustworthy clinical AI systems and increasingly integrated AI ecosystems.
This article is intended as a narrative review rather than a systematic evidence synthesis. Instead of comprehensively cataloging all AI applications in radiation oncology, we selected representative radiation oncology and cross-industry examples that illustrate recurring principles governing successful AI development, validation, implementation, and long-term deployment. Priority was given to influential publications, landmark implementation successes and failures, major regulatory developments, and examples with direct translational relevance to radiation oncology. Accordingly, this review emphasizes conceptual synthesis and practical implementation lessons rather than exhaustive literature retrieval or formal quality assessment of the selected AI examples.
Current landscape of AI in radiation oncology
AI applications now span nearly every component of the radiotherapy workflow, although maturity and clinical adoption vary substantially across domains. Figure 2 conceptualizes the evolving AI landscape in key radiation oncology applications according to developmental maturity and projected future integration. Regulatory maturity has progressed alongside technical development. Most commercially deployed radiation oncology AI systems are regulated as Software as a Medical Device (SaMD) and have entered clinical practice through FDA 510(k), De Novo, or similar international regulatory pathways. To date, the majority of cleared radiation oncology AI products focus on relatively bounded tasks such as auto-segmentation, image processing, treatment planning support, and quality assurance, reflecting areas where performance can be objectively validated and human oversight remains straightforward. Emerging FDA initiatives, including the Predetermined Change Control Plan framework for adaptive AI, may further facilitate safe deployment of continuously evolving clinical AI systems ().
Figure 2
Imaging
AI has become increasingly important in radiotherapy imaging workflows. Deep learning approaches have been applied to image denoising, artifact reduction, deformable image registration, image reconstruction, and synthetic image generation (). In particular, synthetic computed tomography (sCT) generation from magnetic resonance (MR) imaging and cone beam CT (CBCT) has facilitated adaptive radiotherapy workflows and enabled streamlined online adaptive radiotherapy (oART) platforms (, , ). Despite promising early results, substantial challenges remain involving reproducibility, imaging standardization, generalizability, and biological interpretability. In the future, accurate and robust real-time multimodal image fusion or synthetic image generation would be desirable.
Contouring
Auto-segmentation is one of the most clinically mature AI applications in radiation oncology (). Deep learning-based contouring systems can substantially reduce contouring time and improve consistency for many organs at risk (OARs) and even for target structures (, ). Clinical deployment, overtaking previous atlas-based algorithms, has accelerated in recent years, particularly for relatively standardized disease sites. However, contouring performance remains highly dependent on anatomy, imaging quality, institutional contouring philosophy, and disease complexity. AI models may fail in cases of uncommon anatomy, low-quality imaging, or anatomically distorted cases (). Furthermore, large spurious errors may still happen, such as generating contours in the wrong part of the anatomy or grossly wrong contours. Even seemingly small deviations in segmentation may carry clinically meaningful dosimetric implications (). Consequently, auto-segmentation currently has varying degrees of penetration across different clinics (). While it has been shown to improve clinical efficiency, it still requires careful review and often manual edits (, ). Nevertheless, auto-segmentation represents one of the clearest examples of successful clinical AI translation in radiation oncology, with several commercially available FDA-cleared platforms now routinely reducing contouring workload and improving workflow efficiency across many disease sites.
Treatment planning
Machine learning and knowledge-based planning (KBP) systems have increasingly been incorporated into treatment planning workflows. These tools aim to improve planning consistency, reduce inter-planner variability, and accelerate optimization processes. Current approaches include automated rule implementation and reasoning, multicriteria optimization, dose-volume histogram (DVH) prediction models, dose prediction models, generative AI- or reinforcement learning-based planning, and AI-assisted agentic treatment planning (). Although these systems may improve efficiency, they remain strongly influenced by the quality and characteristics of historical training data. Planning models frequently learn institutional planning preferences rather than universally optimal treatment strategies (). As technologies and planning paradigms evolve, static planning models may also become increasingly vulnerable to performance drift. Furthermore, training these models requires significant expertise and time, and direct integration into commercial clinical treatment planning systems (TPS) is frequently unavailable. Consequently, wide clinical integration remains limited and relies on simple methods, while newer, more advanced methods typically remain in-house at large academic centers.
Outcome prediction
Predictive AI models have been developed to estimate toxicity, local control, survival, treatment response, and other clinically relevant outcomes using combinations of imaging, dosimetric, genomic, and clinical variables (). These approaches hold promise for individualized treatment adaptation and outcome-driven, biologically guided radiotherapy (). However, outcome prediction remains challenging due to heterogeneous datasets, inconsistent endpoint definitions, evolving systemic therapies, and relatively limited sample sizes compared with many non-medical AI domains. Many published models remain retrospective, lack robust external validation, and demonstrate limited generalizability (, ). Consequently, their prospective evaluation in clinical trials remains highly limited.
Quality assurance and workflow automation
AI applications in these areas are rapidly expanding. Proposed or investigated applications include machine QA and monitoring, patient-specific QA, automated plan checking, contour QA, delivery verification, continuous machine performance surveillance, incident learning, scheduling optimization, and treatment workflow orchestration (, ). Reported applications include EPID-based delivery verification, log-file analysis, machine performance prediction, automated image quality assessment, predictive maintenance, and anomaly detection using longitudinal machine QA data (). Several of these applications have reached clinical implementation and illustrate that AI can augment not only planning workflows but also treatment delivery and equipment quality management (). LLMs and natural language processing are also leveraged to further support documentation, communication, and clinical decision support (, ). A rapidly emerging application area involves physician-facing workflow support. Recent advances in LLMs have enabled ambient documentation systems to generate consultation notes, on-treatment visit documentation, treatment summaries, referral letters, and patient instructions directly from clinical encounters (). Beyond documentation, these systems can assist with chart summarization, extraction of staging and treatment history from unstructured records, tumor board preparation, inbox management, and guideline retrieval (, ). For radiation oncologists, who routinely integrate pathology, imaging, multidisciplinary recommendations, prior therapies, dose-fractionation decisions, and longitudinal toxicity assessments, these tools may substantially reduce administrative burden and cognitive workload. Notably, documentation-focused AI represents a distinct paradigm from many traditional AI applications in radiation oncology, as its primary value lies in workflow augmentation rather than in autonomous clinical decision-making. Although these technologies promise to improve operational efficiency, widespread clinical integration remains limited. Furthermore, they introduce critical concerns, including automation fatigue, alert overload, hidden failure modes, and excessive reliance on AI outputs.
Collectively, these developments suggest that radiation oncology is transitioning from isolated AI tools toward increasingly interconnected and adaptive AI ecosystems. Yet the broader challenge remains not merely a matter of technical capability but of successful integration into complex human clinical systems.
Cross-industry AI successes and failures as a framework for radiation oncology
As radiation oncology increasingly adopts AI, many of the emerging challenges resemble those already encountered in other industries. Cross-industry AI experiences, therefore, offer valuable insights into both successful deployment strategies and common failure modes. Table 1 summarizes representative examples and their implications for radiation oncology.
Table 1
| Lesson domain | Industry failure | Industry success | Radiation oncology parallel & R&D guidance |
|---|---|---|---|
| Data Integrity & Validation | IBM Watson for Oncology and Epic’s sepsis prediction model performed poorly in real-world settings because the models were trained on limited, retrospective, or nonrepresentative datasets, leading to unsafe recommendations and excessive false alarms (, , ). | Paige AI achieved FDA breakthrough designation through validation on large, diverse pathology datasets across multiple tissue types (). | Auto-segmentation and planning models may fail on uncommon anatomy, evolving workflows, or institutional variations if these are absent from training data. Robust multicenter validation and inclusion of “edge-case” anatomy are essential to prevent clinically unusable “shelfware.” |
| Algorithmic Reasoning | Google Health COVID-19 models demonstrated “shortcut learning,” identifying confounding features (e.g., patient positioning) rather than true pathology (, ). LLM chatbots may generate medically plausible but unsafe oncology recommendations (). | AlphaFold achieved robust protein structure prediction through biologically grounded model design, confidence estimation, and rigorous external validation against experimentally determined structures, helping ensure predictions reflected meaningful biological relationships rather than superficial correlations (). | RT AI may learn physician-specific habits or geometric shortcuts rather than biologically meaningful relationships. Development should incorporate explainable AI, saliency mapping, counterfactual testing, and iterative expert review to ensure clinically meaningful reasoning rather than pattern mimicry. |
| Operational Boundaries | NYC’s “MyCity” chatbot provided inaccurate legal advice because it lacked mechanisms to recognize uncertainty or defer to authoritative sources (). | Klarna implemented “trust thresholds,” requiring AI-to-human handoff when confidence falls below predefined limits (). | AI deployment should incorporate explicit operational boundaries and uncertainty-triggered human intervention. Examples include mandatory physician review for high-risk anatomy, large PTV-OAR overlap, or low-confidence segmentation outputs. |
| Contextual Metrics | Zillow’s pricing algorithms failed despite high predictive accuracy because the models ignored contextual and subjective real-world factors (). | Waymo evaluates autonomous driving systems using real-world safety performance, edge-case behavior, and intervention rates rather than isolated perception metrics alone (). | Standard AI metrics such as Dice similarity coefficient and DVH indices may fail to capture clinically meaningful contouring errors or plan quality pitfalls. RT AI should incorporate context-aware endpoints, including dosimetric impact, toxicity relevance, functional status, and longitudinal clinical outcomes. |
| Integration & Infrastructure | Google Thailand’s diabetic retinopathy program failed partly because algorithms could not accommodate low-quality images and fragile local infrastructure (). Volkswagen Cariad failed due to poorly integrated “black-box” software architecture (). | Viz.ai improved stroke workflow efficiency by integrating imaging AI directly into clinical communication and logistical coordination pathways (, ). | AI tools requiring manual DICOM transfer, isolated workflows, or poor interoperability are unlikely to scale clinically. Successful RT AI should integrate natively with R & V and TPS while minimizing workflow friction. |
| Model Bias & Fairness | Amazon’s recruiting AI reproduced historical gender bias embedded within training data (). | Paige AI reduced institutional bias by training foundation models on diverse pathology data spanning hundreds of institutions and dozens of countries (). | RT AI trained on historical institutional data may perpetuate disparities in treatment selection, access, or planning philosophy. Diverse datasets, fairness auditing, and continuous retraining are critical. |
| Human Agency & Oversight | Overreliance on opaque automation systems contributed to workforce disruption, automation complacency, and silent system failures (). Oncology chatbots have demonstrated partially non-concordant or potentially harmful treatment recommendations (). | Paige AI and emerging ambient clinical documentation platforms have demonstrated successful human-AI collaboration by functioning as supervised amplifiers that augment expert performance while preserving clinician accountability (, , ). | AI should function as an assistive “agentic resident” or clinical copilot rather than an autonomous decision-maker. Successful deployment will likely emphasize workflow augmentation, information synthesis, documentation support, and decision assistance while preserving clinician supervisory authority, override capability, and active engagement. |
Cross-industry AI lessons and implications for AI in radiation oncology.
DVH, dose volume histogram; FDA, Food and Drug Administration; LLM, large language model; NYC, New York City; OAR, organ at risk; PTV, planning target volume; R&D, research and development; R&V, record and verify system; RT, radiotherapy; TPS, treatment planning system.
Data integrity and validation
One of the clearest lessons from prior AI deployment failures is that model performance is fundamentally constrained by the quality, diversity, and representativeness of the underlying training data. IBM Watson for Oncology became an early cautionary example after reports revealed clinically inappropriate and potentially unsafe treatment recommendations during deployment ().
Similarly, Epic’s widely implemented sepsis prediction algorithm demonstrated poor sensitivity and excessive false-positive alerts when evaluated in independent, real-world clinical settings, despite promising retrospective performance during development (, ). This discrepancy occurred largely because the original model was trained on curated historical datasets that relied heavily on retrospective administrative billing codes rather than real-time clinical definitions of sepsis onset. When exposed to heterogeneous, fragmented, and noisy electronic health record data, it failed to replicate its development accuracy, triggering massive alert fatigue. In both cases, models trained on curated or retrospective datasets struggled when exposed to the messiness of live clinical workflows.
Conversely, successful systems achieved regulatory milestones by fundamentally shifting how clinical training data is sourced and scaled. Paige AI secured the first-ever FDA De Novo clearance for an AI tool in pathology (Paige Prostate) by moving away from traditional, manually annotated training sets, which are highly prone to local institutional bias (). Instead, they built a massive foundation model trained on millions of unannotated whole-slide images sourced across hundreds of global institutions and dozens of countries. This cross-institutional diversity inherently accounted for variations in tissue staining, scanner types, and clinical populations.
Radiation oncology faces highly analogous dataset-driven challenges. Deep learning-based auto-segmentation and automated planning algorithms, trained predominantly on idealized anatomy, frequently degrade when confronted with the physical realities of the clinic. For example, auto-segmentation models can fail when they encounter anatomy unseen during training, such as postoperative changes, rectal spacers, abdominal compression, institutional contouring variability, or image artifacts (, ). For example, a head-and-neck auto-segmentation model trained predominantly on intact anatomy may perform poorly in patients with extensive postoperative reconstruction, bulky nodal disease, dental artifacts, or re-irradiation changes. Similarly, automated planning models developed using historical datasets may degrade when confronted with anatomy or clinical scenarios that were absent or underrepresented during training. Cases involving extensive postoperative reconstruction, metal implants, dental artifacts, hydrogel spacers, abdominal compression, severe anatomical distortion, re-irradiation, or rare anatomical presentations such as prune belly syndrome may represent out-of-distribution scenarios in which dose prediction and automated optimization may produce clinically unacceptable outputs, generating unrealistic objectives or clinically unacceptable plans despite excellent benchmark performance on standard datasets. Models that perform well under benchmark testing may become clinically unusable when exposed to these complexities of routine practice. Accordingly, ensuring the safety and generalizability of radiation oncology AI requires a shift toward multi-institutional development that prioritizes edge-case enrichment, rigorous external validation across disparate scanner types and acquisition protocols, and continuous, post-deployment performance auditing.
A critical gap in the radiation oncology AI literature is the limited use of prospective implementation studies that measure real-world clinical impact. Most published evidence remains retrospective and single-institutional, which risks optimistic performance estimates and limited generalizability (). Future work should increasingly incorporate silent-mode validation, pragmatic clinical trials, cluster-randomized implementation studies, and standardized reporting of workflow outcomes, efficiency gains, and patient-centered endpoints. Demonstrating clinical utility under real-world conditions will be essential for sustainable adoption.
Algorithmic reasoning and shortcut learning
AI systems can achieve exceptional predictive accuracy on paper while relying entirely on clinically irrelevant, confounding features. During the COVID-19 pandemic, multi-institutional studies revealed that many deep-learning models developed to detect pulmonary pathology on chest radiographs were derailed by “shortcut learning.” (, ) Instead of evaluating actual consolidated lung tissue, these computer vision models learned to exploit spurious correlations, such as the unique font styles of institutional text annotations, the presence of chest drains, or whether an image was captured portably in a supine position (highly correlated with severe disease) versus an outpatient upright scan. AI systems optimized strictly for a loss function will naturally exploit hidden mathematical shortcuts, which inevitably collapse under real-world distributional shifts (). Similarly, contemporary LLMs can generate highly fluent, structurally flawless recommendations that are medically unsafe, masking severe omissions or pharmacological hallucinations behind an authoritative tone. For example, researchers found that 33.3% of the chatbot’s treatment recommendations were at least partially non-concordant with the National Comprehensive Cancer Network (NCCN) guidelines (). It blended incorrect radiation and chemo advice with highly accurate text, making the errors nearly invisible to non-experts. These failures underscore the critical distinction between statistical correlation and clinically meaningful reasoning and warn against AI hallucinations. Conversely, successful AI systems often incorporate mechanisms that encourage learning of meaningful underlying relationships rather than superficial correlations. For example, AlphaFold achieved unprecedented protein structure prediction performance through biologically grounded model design, confidence estimation, and rigorous validation against experimentally determined structures, helping to ensure that predictions reflected fundamental structural biology rather than dataset-specific artifacts (). Together, these examples illustrate that robust AI requires not only predictive accuracy but also evidence that model reasoning aligns with the underlying domain of interest.
Radiation oncology AI models face nearly identical systemic vulnerabilities. RT AI may learn physician-specific contouring habits, institution-specific planning conventions, or spurious correlations with imaging features such as immobilization devices, rather than clinically meaningful biological or anatomical relationships. In automated planning algorithm development or dose prediction, deep reinforcement learning or convolutional neural networks can easily memorize geometric regularities, standard couch angles, or a single institution’s non-protocol scripting habits rather than learning the underlying radiobiological and physical principles (). Likewise, radiomic and clinical outcome prediction models are highly susceptible to historical treatment selection biases and confounding factors (). A practical example is toxicity prediction. A model may identify feeding-tube placement or treatment breaks as powerful predictors of toxicity, not because they are biologically causal, but because they act as proxies for patients who already experienced severe treatment-related complications. Without careful causal reasoning, such models may achieve excellent predictive performance while offering limited guidance for treatment optimization. For instance, a model might misinterpret a clinician’s defensive decision to prescribe lower doses to a frail patient as a causal biological relationship between reduced dose and increased toxicity. Outcome models can also misinterpret planning trade-offs as biological phenomena. A model might flag a specific OAR DVH metric as ‘protective’ due to a negative correlation with toxicity, failing to realize this is a mathematical artifact of dose redistribution. The lower toxicity is actually driven by a reduction in a separate, more critical volume constraint, not a biological benefit inherent to the metric itself (). To prevent these models from acting as mere superficial pattern matchers, future development must look beyond standard performance metrics and enforce explainable AI frameworks, such as saliency mapping to verify anatomical attention, counterfactual stress testing to evaluate model bounds, and iterative, expert-in-the-loop clinical validation.
Operational boundaries and human oversight
Another recurring lesson from cross-industry AI failures is the danger of deploying systems without clearly defined operational limits. New York City’s “MyCity” chatbot, designed to answer questions about municipal regulations and public services, provided users with inaccurate and sometimes legally incorrect guidance on topics such as housing and employment regulations because it lacked reliable mechanisms to recognize uncertainty, verify information against authoritative sources, or escalate complex questions to human experts (). In contrast, financial technology companies such as Klarna have implemented explicit “trust thresholds,” whereby AI systems can autonomously handle routine customer interactions but must transfer cases to human representatives whenever confidence falls below predefined thresholds or when requests involve ambiguity, exceptions, or potential financial risk ().
Trustworthy clinical AI similarly requires transparent operational boundaries. Radiation oncology workflows contain numerous high-risk scenarios in which AI uncertainty may substantially increase patient risk, such as patients with unusual anatomy, postoperative changes, large overlap between targets and OARs, poor-quality or artifact-degraded imaging, and uncommon disease presentations that are underrepresented in training datasets. In such cases, inaccurate auto-segmentation, treatment planning, image registration, or outcome prediction may propagate errors directly into clinical decision-making and treatment delivery. Accordingly, AI systems should incorporate predefined escalation triggers, quantitative confidence estimation, and mandatory human review mechanisms whenever uncertainty exceeds acceptable thresholds. In practice, these triggers may include gross contour-volume outliers, unexpected target shifts during adaptive radiotherapy, unusually high predicted OAR doses, substantial disagreement between independent AI models, or major deviations from institutional planning objectives. Such scenarios should automatically prompt enhanced physician or physicist review rather than silent acceptance of AI outputs. Importantly, clinician trust is not established merely by high average performance metrics but by demonstrated robustness to rare but consequential failures, predictable behavior under uncertainty, transparent acknowledgment of limitations, and the ability to reliably identify situations that require human intervention.
Contextual metrics and clinical relevance
Cross-industry experiences also demonstrate that conventional quantitative performance metrics may fail to capture real-world utility. For example, real estate company Zillow’s home valuation algorithms achieved strong predictive accuracy when assessed against historical sales data, yet the company’s home-buying initiative suffered substantial losses because the models inadequately accounted for local market dynamics, rapidly changing economic conditions, property-specific characteristics, and subjective factors influencing buyer behavior like “curb appeal” and local sentiment (). Conversely, autonomous driving companies such as Waymo evaluate system performance using broader measures of real-world safety, including rare-event handling, intervention rates, and operational reliability, rather than relying solely on isolated perception or object-detection metrics (). Together, these examples illustrate that technically impressive benchmark performance does not necessarily translate into successful real-world decision-making when important contextual factors are omitted.
Radiation oncology faces analogous challenges when evaluating AI performance with quantitative metrics that may not fully capture clinical impact. For example, the Dice similarity coefficient remains one of the most widely reported metrics in auto-segmentation studies, yet it may be insensitive to clinically meaningful errors (). A segmentation of the optic chiasm, brachial plexus, cochlea, or other small structures may achieve a seemingly acceptable Dice score while still containing millimeter-level boundary deviations that substantially alter dose estimates (). Likewise, two target contours may achieve high volumetric overlap despite clinically important differences in shape, extent, or inclusion of critical regions that could affect treatment coverage. Similarly, treatment-planning AI is often evaluated using limited DVH metrics or composite plan-score functions. For example, two lung SBRT plans may demonstrate nearly identical DVH statistics, yet differ substantially in chest-wall dose distribution, hotspot location, interplay robustness, or physician-perceived deliverability. These nuances are often immediately apparent to experienced planners but remain difficult to capture using conventional benchmark metrics. While useful, these metrics reduce complex 3D dose distributions to a limited set of numerical summaries and may fail to capture clinically relevant features such as dose spillage into adjacent tissues, hotspot location, dose conformity around critical structures, robustness to anatomical variation, or tradeoffs that experienced clinicians consider when judging plan quality. Furthermore, outcome prediction models may demonstrate strong discrimination, or calibration statistics yet provide little practical benefit if their predictions do not change treatment decisions, improve patient selection, reduce toxicity, or enhance clinical outcomes.
These limitations highlight the distinction between technical performance and clinical value. An AI model may outperform existing methods on benchmark datasets yet be neither clinically viable nor deliver measurable benefit to patients or clinicians in routine practice, either because it misses nuanced but important clinical considerations or because it is clinically viable but delivers little measurable benefit. Future evaluation frameworks should therefore prioritize clinically meaningful endpoints, including dosimetric consequences of AI-generated contours, efficiency gains across the treatment workflow, reduction in toxicity and treatment-related complications, improvements in adaptive decision-making, consistency across institutions and patient populations, and ultimately longitudinal patient outcomes. Such evaluation strategies more closely align AI assessment with the goals of clinical care rather than relying solely on surrogate technical metrics.
Integration, infrastructure, and workflow compatibility
Many AI failures arise not from poor algorithms alone, but from inadequate integration into operational ecosystems. For example, Google Thailand’s diabetic retinopathy screening initiative developed an AI system capable of accurately detecting diabetic eye disease from retinal photographs, achieving 90% laboratory accuracy and reducing diagnostic confirmation time from 10 weeks to 10 minutes (). However, for this AI tool created for those who lacked immediate access to eye specialists, real-world deployment proved challenging because images acquired in rural clinics were often lower quality than those used during development, necessitating repeated image acquisition and creating workflow bottlenecks. Limited internet connectivity and speed, staffing constraints, and mismatches between AI requirements and local clinical processes also reduced practical effectiveness despite strong technical performance. Similarly, Volkswagen’s Cariad software initiative sought to create a unified software platform across multiple vehicle brands and functions. Although individual components were technologically sophisticated, insufficient interoperability, integration complexity, and delays in coordinating software across existing vehicle systems contributed to major implementation challenges (). Both examples illustrate that even highly capable AI systems may fail when they are not designed around the realities of operational infrastructure and end-user workflows.
In contrast, the stroke detection platform Viz.ai achieved clinical impact not merely through accurate image interpretation but by embedding AI directly into existing stroke-care pathways (, ). When the system identified a suspected large-vessel occlusion on imaging, it automatically alerted the appropriate specialists, shared relevant information with care teams, and accelerated coordination across emergency, radiology, and neurointerventional services. By reducing communication delays rather than simply generating predictions, the platform translated algorithmic performance into measurable clinical benefit. The closest analogue in radiation oncology may not be a contouring model or planning algorithm, but future AI systems that orchestrate adaptive workflows by automatically coordinating image review, contour approval, plan generation, QA, and treatment delivery.
Radiation oncology similarly depends on highly interconnected workflows involving imaging acquisition, contouring, treatment planning, plan review, QA, treatment delivery, adaptive replanning, and longitudinal follow-up. These processes span TPSs, oncology information systems, imaging platforms, delivery systems, and multidisciplinary clinical teams. An AI contouring tool that requires manual DICOM export and import, a planning model that operates outside the treatment planning environment, or an adaptive workflow that requires numerous additional user interactions may create friction that outweighs its theoretical benefits. The challenge is particularly relevant for adaptive radiotherapy, where the value of AI-generated contours or plan adaptation may disappear if piecemeal software solutions handle different workflow steps and require cumbersome data export, import, and manual verification, thereby prolonging treatment sessions and increasing on-table time. Similar considerations apply to AI-assisted applicator reconstruction, auto-segmentation for image-guided brachytherapy, and automatic planning for intraoperative brachytherapy, where highly accurate algorithms may provide limited clinical benefit if they remain fragmented across multiple platforms rather than functioning within a seamless end-to-end workflow. In these settings, workflow efficiency, usability, and seamless integration become as important as algorithmic accuracy. Likewise, AI-generated recommendations that are not visible within existing clinical workspaces are less likely to influence decision-making. Successful deployment will therefore require seamless interoperability, native integration, automated data exchange, and minimal workflow disruption. In practice, proprietary software architectures and vendor-specific interfaces, especially in highly streamlined systems such as integrated adaptive radiotherapy platforms, remain major barriers to scalable third-party AI deployment despite algorithmic performance.
Human agency, bias, and the future clinical role of AI
Perhaps the most important cross-industry lesson is that successful AI deployment augments rather than replaces human expertise. A widely cited example is Amazon’s experimental recruiting algorithm, which was trained using historical hiring data (). Because the underlying data reflected existing workforce imbalances, the system learned to systematically favor characteristics associated with previously hired candidates and penalize some indicators associated with female applicants. The issue was not malicious programming, but rather the algorithm’s replication of biases embedded within historical data. Similarly, studies evaluating LLM-based oncology chatbots have shown that while these systems can generate fluent, confident, and often helpful responses, they may also produce partially incorrect, outdated, or potentially harmful recommendations that are difficult for non-experts to recognize (, ). These examples highlight that AI systems can inherit limitations from their training data and may appear more reliable than they actually are. This concern is particularly relevant in oncology, where substantial disparities in cancer outcomes and survival already exist across racial, socioeconomic, geographic, and other underserved populations (). AI systems trained on historical practice patterns risk perpetuating or even amplifying these inequities if fairness is not explicitly considered during development, validation, and deployment. For example, AI may learn disparities in referral patterns and access to advanced technologies such as SBRT, adaptive radiotherapy, MR-guided radiotherapy, or proton therapy. Without deliberate fairness evaluation, these systems risk reinforcing historical inequities under the appearance of objective decision-making.
Conversely, successful healthcare AI deployments have generally functioned as supervised decision-support systems rather than autonomous decision-makers. For example, Paige AI has been successfully deployed in digital pathology to assist pathologists by identifying suspicious regions, prioritizing cases, and reducing oversight errors, while leaving final interpretation and diagnosis under physician control (). Similar patterns have emerged across healthcare, where AI has demonstrated value by supporting cognitive workflows, reducing the information-processing burden, synthesizing complex data, coordinating tasks, and supporting clinical decision-making rather than replacing expert judgment ().
The rapid adoption of ambient documentation platforms and clinical AI copilots provides a contemporary example of this paradigm (, ). These systems generate draft documentation, summarize relevant clinical information, and streamline workflows, thereby reducing clerical burden while preserving physician accountability for clinical decisions. Their success suggests that healthcare AI may achieve broader acceptance when deployed as a cognitive amplifier rather than an autonomous clinician.
This framework may be particularly relevant to radiation oncology, where treatment decisions require continuous integration of imaging, pathology, treatment history, multidisciplinary recommendations, dosimetric tradeoffs, and longitudinal outcomes. Rather than functioning as an autonomous clinician, radiation oncology AI may be most effective as an assistive “agentic resident” that supports complex workflows while remaining under human supervision. Such systems could generate contours, propose treatment plans, identify anatomical changes during adaptive therapy, prioritize quality assurance tasks, summarize clinical information, or estimate toxicity risk. For example, an AI “agentic resident” might assemble prior radiation records, summarize cumulative dose constraints, identify candidate adaptive triggers, draft treatment recommendations, and prepare documentation before physician review. Importantly, such systems could synthesize information from current literature, institutional protocols, and individual physician practice patterns while integrating data distributed across years of clinical records, imaging studies, treatment courses, and multidisciplinary encounters. This capability may be particularly valuable in radiation oncology, where clinicians must routinely reconcile large volumes of longitudinal information and evolving evidence. By continuously aggregating and contextualizing data across patients and time, AI systems may help overcome inherent human cognitive limitations in information retrieval, pattern recognition, and longitudinal decision support while preserving clinician oversight and final accountability. In this model, AI reduces cognitive burden while final responsibility for treatment decisions remains with the clinical team. However, physicians, physicists, and dosimetrists must retain ultimate responsibility for contextual interpretation, safety oversight, and ethical judgment. In addition, clinicians must remain vigilant to ensure that AI systems do not perpetuate or amplify disparities embedded within the datasets used for model development.
Preserving meaningful human engagement is therefore essential not only for patient safety but also for mitigating automation complacency, preventing progressive workflow deskilling, and detecting system failures that may otherwise go unnoticed when clinicians become overly reliant on automated outputs (). Accordingly, the greatest impact of AI in radiation oncology may lie not in replacing expertise, but in augmenting it: redistributing cognitive workload, accelerating routine tasks, expanding access to specialized capabilities, and enhancing decision support while preserving human judgment, accountability, and professional responsibility.
The path forward for AI in radiation oncology
The next phase of AI development in radiation oncology will likely move beyond isolated task automation toward increasingly integrated, adaptive, agentic, and multimodal clinical ecosystems. Future systems may increasingly integrate radiation oncology data into continuously learning platforms that support real-time clinical decision-making across contouring, treatment planning, quality assurance, adaptive radiotherapy, and clinical documentation. Advances in large language models, multimodal foundation models, reinforcement learning, and agentic AI may further enable coordinated workflow orchestration across these domains. However, these same capabilities may also introduce new failure modes. Foundation models and agentic systems can propagate hallucinations across interconnected workflows, exhibit prompt sensitivity, misuse external tools, or amplify errors through autonomous task delegation (). Rigorous evaluation frameworks, sandbox testing environments, uncertainty estimation, and human-in-the-loop safeguards will therefore be essential before widespread clinical deployment. Prospective failure mode and effects analysis (FMEA) may provide a useful framework for systematically identifying AI-associated failure modes, evaluating their clinical consequences, and implementing mitigation strategies before routine clinical deployment ().
Long-term clinical impact will depend less on isolated algorithmic sophistication than on the coordinated development of trustworthy clinical AI ecosystems. Lessons from healthcare, aviation, finance, and other AI-enabled industries suggest that successful deployment requires simultaneous progress across data quality, model robustness, workflow integration, human oversight, regulatory governance, and continuous performance monitoring (67). Failure in any one of these domains may limit clinical impact, even when algorithmic accuracy is high. Future AI development should therefore follow a multilayer roadmap that prioritizes technical excellence alongside operational feasibility, clinical adoption, safety, equity, and long-term sustainability. Sustainable adoption will also require demonstration of economic value through improvements in efficiency, resource utilization, access to care, or clinical outcomes sufficient to justify implementation, maintenance, and governance costs.
Regulatory and institutional governance will be central to this roadmap. In the United States, the FDA’s Predetermined Change Control Plan framework represents an important step toward regulating adaptive AI systems while maintaining safety and effectiveness (68). Similarly, the European Union AI Act introduces risk-based oversight requirements for healthcare AI systems (). In radiation oncology, these external regulatory frameworks will likely need to be complemented by local governance structures, including AI oversight committees, predefined performance thresholds, audit trails, version control, incident reporting mechanisms, and routine drift monitoring (69, 70). Such infrastructure will be especially important for managing continuously learning systems whose behavior may evolve over time. Model performance may also drift following routine clinical changes unrelated to the AI itself. Examples include treatment planning system algorithm or beam model upgrades, implementation of new CT simulation or CBCT reconstruction protocols, introduction of new immobilization devices, adoption of online adaptive workflows, or shifts in patient populations. These changes may gradually alter the clinical data distribution and degrade AI performance despite unchanged model parameters. This underscores the need for continuous local performance monitoring, periodic revalidation, and version-aware governance. Trustworthy AI deployment ultimately depends on multidisciplinary collaboration rather than any single individual or profession. While governance structures should be tailored to local practice, successful implementation requires coordinated clinical, technical, operational, and organizational oversight. Table 2 presents representative stakeholder roles that together support the safe, effective, and sustainable integration of AI into radiation oncology workflows.
Table 2
| Stakeholder | Primary role in trustworthy AI deployment |
|---|---|
| Radiation Oncologist | Clinically validate AI outputs, approve AI-assisted decisions, and monitor clinical outcomes. |
| Medical Physicist | Commission and validate AI tools, perform QA, monitor performance drift, and investigate failures. |
| Dosimetrist/Radiation Therapist | Verify AI-assisted planning and workflow outputs, identify operational issues, and provide user feedback. |
| Clinical Informatics/IT | Integrate AI with clinical systems, ensure interoperability, cybersecurity, version control, and data governance. |
| Department Leadership/AI Champion | Lead implementation, user education, change management, and resource planning. |
| AI Oversight Committee | Establish governance policies, oversee auditing and drift monitoring, define escalation pathways, and ensure regulatory compliance. |
Example multidisciplinary roles supporting trustworthy AI deployment in radiation oncology.
Representative examples only; responsibilities may vary across institutions.
Importantly, future AI systems should be designed not only to improve technical efficiency but also to enhance clinically meaningful outcomes, reduce disparities in access to high-quality radiotherapy, and support more personalized and biologically informed care. As radiotherapy increasingly incorporates online adaptation, functional imaging, biologically guided treatment strategies, and longitudinal response assessment, AI may serve as a coordinating infrastructure that synthesizes growing clinical complexity into actionable decision support.
Ultimately, the future role of AI in radiation oncology will likely be that of a collaborative clinical partner rather than a fully autonomous replacement for human expertise. The most successful systems may be those that preserve clinician agency, enhance multidisciplinary decision-making, improve workflow resilience, and augment human judgment while maintaining transparency, accountability, and patient-centered care.
A practical roadmap for this future is illustrated in Figure 3. Diverse multicenter datasets, standardized benchmarking, and bias mitigation form the foundation for robust AI development. These models must then undergo external validation, uncertainty quantification, and failure-mode analysis before clinical deployment. Successful translation will also require seamless integration, human oversight, confidence-triggered escalation, and interoperability with existing clinical systems. Finally, regulatory governance, continuous monitoring for model drift, and institutional oversight structures will be necessary to ensure safe long-term operation. Together, these interconnected layers provide a framework for achieving meaningful clinical impact while preserving safety, transparency, equity, and clinician accountability.
Figure 3
This review has several limitations. First, as a narrative review, it was designed to provide conceptual synthesis rather than a comprehensive systematic assessment of all published literature. Representative examples from radiation oncology and other high-stakes industries were intentionally selected to illustrate recurring implementation principles, and other relevant examples may also exist. Second, because AI is evolving rapidly, some technologies discussed, such as multimodal foundation models, clinical AI copilots, and agentic systems, remain early in clinical translation and should be interpreted as emerging future directions rather than established standards of care. Finally, the review emphasizes implementation science, governance, and translational considerations rather than detailed technical comparison of individual AI algorithms.
Conclusion
Artificial intelligence is rapidly transforming radiation oncology, but experiences across healthcare and other high-stakes industries demonstrate that successful deployment depends on far more than algorithmic performance alone. The most durable AI successes have emerged from systems that combine robust validation, workflow integration, human oversight, and continuous monitoring, whereas many highly publicized failures resulted from nonrepresentative data, shortcut learning, workflow incompatibility, hidden bias, or inadequate governance. Importantly, some of the most impactful healthcare AI deployments have achieved value not through autonomous clinical decision-making, but through workflow augmentation, cognitive support, information synthesis, and reduction of administrative burden.
As radiation oncology evolves toward increasingly integrated clinical ecosystems, future development should prioritize prospective implementation, fairness auditing, uncertainty-aware deployment, and preservation of clinician supervisory authority. The goal is not autonomous replacement of clinicians but trustworthy AI that augments human expertise while improving efficiency, consistency, and patient care.
As radiation oncology enters an era of biologically guided, data-intensive, and increasingly adaptive treatment paradigms, the ultimate impact of AI will likely depend on how effectively the field integrates technological innovation with clinical judgment, multidisciplinary collaboration, and patient-centered care. By learning from prior cross-industry successes and failures, radiation oncology can move beyond viewing AI as a collection of isolated algorithms and instead build trustworthy clinical ecosystems that safely amplify human expertise, improve patient care, and realize the full potential of data-driven radiotherapy.
Statements
Author contributions
MZ: Writing – review & editing, Data curation, Investigation, Software, Methodology, Resources, Visualization, Writing – original draft, Formal analysis, Validation. LG: Data curation, Formal analysis, Validation, Investigation, Writing – review & editing. CZ: Formal analysis, Writing – review & editing, Resources, Methodology, Conceptualization, Visualization, Validation. DZ: Project administration, Visualization, Resources, Formal analysis, Validation, Conceptualization, Writing – review & editing, Methodology, Data curation, Supervision, Writing – original draft, Investigation, Software.
Funding
The author(s) declared that financial support was not received for this work and/or its publication.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that generative AI was used in the creation of this manuscript. The authors used ChatGPT (OpenAI) to assist with language editing and improvement of manuscript readability. Some figures were conceptually designed by the authors, generated using AI-assisted illustration tools, and subsequently refined and scientifically verified by the authors. The authors reviewed and edited all AI-generated content and take full responsibility for the accuracy and integrity of the manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
References
1
VandewinckeleLClaessensMDinklaABrouwerCCrijnsWVerellenDet al. Overview of artificial intelligence-based applications in radiotherapy: Recommendations for implementation and quality assurance. Radiotherapy Oncol. (2020) 153:55–66. doi: 10.1016/j.radonc.2020.09.008
2
WebsterMPodgorsakALiFZhouYJungHYoonJet al. New approaches in radiotherapy. Cancers. (2025) 17:1980. doi: 10.3390/cancers17121980
3
ChandraRAKeaneFKVonckenFEThomasCR. Contemporary radiotherapy: present and future. Lancet. (2021) 398:171–84. doi: 10.1016/s0140-6736(21)00233-6
4
Dona LemusOMCaoMCaiBCummingsMZhengD. Adaptive radiotherapy: next-generation radiotherapy. Cancers. (2024) 16:1206. doi: 10.3390/cancers16061206
5
HuynhEHosnyAGuthierCBittermanDSPetitSFHaas-KoganDAet al. Artificial intelligence in radiation oncology. Nat Rev Clin Oncol. (2020) 17:771–81. doi: 10.1038/s41571-020-0417-8
6
IstasyPLeeWSIansavicheneAUpshurRGyawaliBBurkellJet al. The impact of artificial intelligence on health equity in oncology: scoping review. J Med Internet Res. (2022) 24:e39748. doi: 10.2196/39748
7
BernstamEVShiremanPKMeric‐BernstamFZozusNJiangXBrimhallBBet al. Artificial intelligence in clinical and translational science: Successes, challenges and opportunities. Clin Transl Sci. (2022) 15:309–21. doi: 10.1111/cts.13175
8
CampanellaGHannaMGGeneslawLMiraflorAWerneck Krauss SilvaVBusamKJet al. Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nat Med. (2019) 25:1301–9. doi: 10.1038/s41591-019-0508-1
9
JumperJEvansRPritzelAGreenTFigurnovMRonnebergerOet al. Highly accurate protein structure prediction with AlphaFold. Nature. (2021) 596:583–9. doi: 10.1038/s41586-021-03819-2
10
Act EAI. The eu artificial intelligence act. In: European Union. Luxembourg: Publications Office of the European Union (2024).
11
DastinJ. Amazon scraps secret AI recruiting tool that showed bias against women. In: Ethics of Data and Analytics. Boca Raton, FL: Auerbach Publications (2022). p. 296–9.
12
StricklandE. IBM Watson, heal thyself: How IBM overpromised and underdelivered on AI health care. IEEE Spectr. (2019) 56:24–31. doi: 10.1109/mspec.2019.8678513
13
CarvalhoEMascarenhasMPinheiroFCorreiaRBalseiroSBarbosaGet al. Predetermined change control plans: Guiding principles for advancing safe, effective, and high-quality AI-ML technologies. JMIR AI. (2025) 4:e76854. doi: 10.2196/76854
14
SahinerBPezeshkAHadjiiskiLMWangXDrukkerKChaKHet al. Deep learning in medical imaging and radiation therapy. Med Phys. (2019) 46:e1–e36. doi: 10.1002/mp.13264
15
JohnstoneEWyattJJHenryAMShortSCSebag-MontefioreDMurrayLet al. Systematic review of synthetic computed tomography generation methodologies for use in magnetic resonance imaging–only radiation therapy. Int J Radiat Oncol Biol Phys. (2018) 100:199–217. doi: 10.1016/j.ijrobp.2017.08.043
16
LiangXChenLNguyenDZhouZGuXYangMet al. Generating synthesized computed tomography (CT) from cone-beam computed tomography (CBCT) using CycleGAN for adaptive radiation therapy. Phys Med Biol. (2019) 64:125002. doi: 10.1088/1361-6560/ab22f9
17
CardenasCEYangJAndersonBMCourtLEBrockKB. (2019). “ Advances in auto-segmentation”, in: Seminars in Radiation Oncology ( Elsevier).
18
MaCYZhouJYXuXTGuoJHanMFGaoYZet al. Deep learning‐based auto‐segmentation of clinical target volumes for radiotherapy treatment of cervical cancer. J Appl Clin Med Phys. (2022) 23:e13470. doi: 10.1002/acm2.13470
19
MenKZhangTChenXChenBTangYWangSet al. Fully automatic and robust segmentation of the clinical target volume for radiotherapy of breast cancer using big data and deep learning. Physica Med. (2018) 50:13–9. doi: 10.1016/j.ejmp.2018.05.006
20
WangBDohopolskiMBaiTWuJHannanRDesaiNet al. Performance deterioration of deep learning models after clinical deployment: a case study with auto-segmentation for definitive prostate cancer radiotherapy. Mach Learning: Sci Technol. (2024) 5:025077. doi: 10.1088/2632-2153/ad580f
21
KieselmannJPKamerlingCPBurgosNMentenMJFullerCDNillSet al. Geometric and dosimetric evaluations of atlas-based segmentation methods of MR images in the head and neck region. Phys Med Biol. (2018) 63:145007. doi: 10.1088/1361-6560/aacb65
22
RusanovBEbertMASabetMRowshanfarzadPBarryNKendrickJet al. Guidance on selecting and evaluating AI auto-segmentation systems in clinical radiotherapy: insights from a six-vendor analysis. Phys Eng Sci Med. (2025) 48:301–16. doi: 10.1007/s13246-024-01513-x
23
DoolanPJCharalambousSRoussakisYLeczynskiAPeratikouMBenjaminMet al. A clinical evaluation of the performance of five commercial artificial intelligence contouring systems for radiotherapy. Front Oncol. (2023) 13:1213068. doi: 10.3389/fonc.2023.1213068
24
LucidoJJDeWeesTALeavittTRAnandABeltranCJBrookeMDet al. Validation of clinical acceptability of deep-learning-based automated segmentation of organs-at-risk for head-and-neck radiotherapy treatment planning. Front Oncol. (2023) 13:1137803. doi: 10.3389/fonc.2023.1137803
25
WangCZhuXHongJCZhengD. Artificial intelligence in radiotherapy treatment planning: present and future. Technol Cancer Res Treat. (2019) 18:1533033819873922. doi: 10.1177/1533033819873922
26
ChungCVKhanMSOlanrewajuAPhamMNguyenQTPatelTet al. Knowledge-based planning for fully automated radiation therapy treatment planning of 10 different cancer sites. Radiotherapy Oncol. (2025) 202:110609. doi: 10.1016/j.radonc.2024.110609
27
BeraKBramanNGuptaAVelchetiVMadabhushiA. Predicting cancer outcomes with radiomics and artificial intelligence in radiology. Nat Rev Clin Oncol. (2022) 19:132–46. doi: 10.1038/s41571-021-00560-7
28
ZhengDEl NaqaIQiXSSethiAAlongiF. Imaging biomarkers in radiotherapy. Cancers. (2026) 18:1232. doi: 10.3390/cancers18081232
29
KangJSchwartzRFlickingerJBeriwalS. Machine learning approaches for predicting radiation therapy outcomes: a clinician's perspective. Int J Radiat Oncol Biol Phys. (2015) 93:1127–35. doi: 10.1016/j.ijrobp.2015.07.2286
30
ThompsonRFValdesGFullerCDCarpenterCMMorinOAnejaSet al. Artificial intelligence in radiation oncology: a specialty-wide disruptive transformation? Radiotherapy Oncol. (2018) 129:421–6. doi: 10.1016/j.radonc.2018.05.030
31
ChanMFWitztumAValdesG. Integration of AI and machine learning in radiotherapy QA. Front Artif Intell. (2020) 3:577620. doi: 10.3389/frai.2020.577620
32
BleaseCKaptchukTJBernsteinMHMandlKDHalamkaJDDesRochesCM. Artificial intelligence and the future of primary care: exploratory qualitative study of UK general practitioners’ views. J Med Internet Res. (2019) 21:e12802. doi: 10.2196/12802
33
WangPLiuZLiYHolmesJShuPZhangLet al. Fine‐tuning open‐source large language models to improve their performance on radiation oncology tasks: A feasibility study to investigate their potential clinical applications in radiation oncology. Med Phys. (2025) 52:e17985. doi: 10.1002/mp.17985
34
NgJJWWangEZhouXZhouKXGohCXLSimGZNet al. Evaluating the performance of artificial intelligence-based speech recognition for clinical documentation: a systematic review. BMC Med Inf Decis Making. (2025) 25:236. doi: 10.1186/s12911-025-03061-0
35
LiuFZhouHGuBZouXHuangJWuJet al. Application of large language models in medicine. Nat Rev Bioeng. (2025) 3:445–64. doi: 10.1038/s44222-025-00279-5
36
ShoolSAdimiSSaboori AmleshiRBitarafEGolpiraRTaraM. A systematic review of large language model (LLM) evaluations in clinical medicine. BMC Med Inf Decis Making. (2025) 25:117. doi: 10.1186/s12911-025-02954-4
37
HabibARLinALGrantRW. The epic sepsis model falls short—the importance of external validation. JAMA Intern Med. (2021) 181:1040–1. doi: 10.1001/jamainternmed.2021.3333
38
WongAOtlesEDonnellyJPKrummAMcCulloughJDeTroyer-CooleyOet al. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern Med. (2021) 181:1065–70. doi: 10.1001/jamainternmed.2021.2626
39
DeGraveAJJanizekJDLeeS-I. AI for radiographic COVID-19 detection selects shortcuts over signal. Nat Mach Intell. (2021) 3:610–9. doi: 10.1038/s42256-021-00338-7
40
RobertsMDriggsDThorpeMGilbeyJYeungMUrsprungSet al. Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans. Nat Mach Intell. (2021) 3:199–217. doi: 10.1038/s42256-021-00307-0
41
ChenSKannBHFooteMBAertsHJSavovaGKMakRHet al. Use of artificial intelligence chatbots for cancer treatment information. JAMA Oncol. (2023) 9:1459–62. doi: 10.1001/jamaoncol.2023.2954
42
BellAFonsecaJ. Output scouting: auditing large language models for catastrophic responses. In: Arxiv Preprint Arxiv:241005305. Ithaca: Cornell University (2024).
43
YuzhongY. AI agents as institutional actors: toward a sociology of agentic governance. Digital Soc Virtual Gov. (2025) 1:43–61. doi: 10.6914/dsvg.010203
44
GudigantalaNMehrotraV. Teaching case: When strength turns into weakness: Exploring the role of AI in the closure of Zillow offers. J Inf Syst Educ. (2024) 35:67–72. doi: 10.62273/trcf3655
45
SunPKretzschmarHDotiwallaXChouardAPatnaikVTsuiPet al. (2020). “ Scalability in perception for autonomous driving: Waymo open dataset”, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.
46
RuamviboonsukPTiwariRSayresRNganthaveeVHemaratKKongprayoonAet al. Real-time diabetic retinopathy screening by deep learning in a multisite national screening programme: a prospective interventional cohort study. Lancet Digital Health. (2022) 4:e235–44. doi: 10.1016/s2589-7500(22)00017-6
47
CallecchiaLGC. Navigating the digital highway: A case study of volkswagen group’s cariad and a make-or-buy decision. In: Universidade Catolica Portuguesa (Portugal). Lisbon: Universidade Católica Portuguesa. (2023).
48
ChatterjeeASomayajiNRKabakisIM. Abstract WMP16: artificial intelligence detection of cerebrovascular large vessel occlusion-nine month, 650 patient evaluation of the diagnostic accuracy and performance of the Viz. ai LVO algorithm. Stroke. (2019) 50:AWMP16–AWMP. doi: 10.1161/str.50.suppl_1.wmp16
49
KaramchandaniRRHelmsAMSatyanarayanaSYangHClementeJDDefilippGet al. Automated detection of intracranial large vessel occlusions using Viz. ai software: experience in a large, integrated stroke network. Brain Behav. (2023) 13:e2808. doi: 10.1002/brb3.2808
50
Rinta-KahilaTPenttinenESalovaaraASolimanWRuissaloJ. The vicious circles of skill erosion: A case study of cognitive automation. J Assoc For Inf Syst. (2023) 24:1378–412. doi: 10.17705/1jais.00829
51
DugganMJGervaseJSchoenbaumAHansonWHowellJTIIISheinbergMet al. Clinician experiences with ambient scribe technology to assist with documentation burden and efficiency. JAMA Netw Open. (2025) 8:e2460637. doi: 10.1001/jamanetworkopen.2024.60637
52
TierneyAAGayreGHobermanBMatternBBallescaMKipnisPet al. Ambient artificial intelligence scribes to alleviate the burden of clinical documentation. NEJM Catalyst Innov Care Delivery. (2024) 5:CAT. 23.0404. doi: 10.1056/cat.23.0404
53
RoperJLinMHRongY. Extensive upfront validation and testing are needed prior to the clinical implementation of AI‐based auto‐segmentation tools. J Appl Clin Med Phys. (2022) 24:e13873. doi: 10.1002/acm2.13873
54
ZhaoHWangBDohopolskiMBaiTJiangSNguyenD. Deep unsupervised clustering for prostate auto-segmentation with and without hydrogel spacer. Mach Learning: Sci Technol. (2025) 6:015015. doi: 10.1088/2632-2153/ada8f3
55
FengMValdesGDixitNSolbergTD. Machine learning in radiation oncology: opportunities, requirements, and needs. Front Oncol. (2018) 8:110. doi: 10.3389/fonc.2018.00110
56
GeirhosRJacobsenJ-HMichaelisCZemelRBrendelWBethgeMet al. Shortcut learning in deep neural networks. Nat Mach Intell. (2020) 2:665–73. doi: 10.1038/s42256-020-00257-z
57
Barragan-MonteroABibalADastaracMHDraguetCValdesGNguyenDet al. Towards a safe and efficient clinical implementation of machine learning in radiation oncology by exploring model interpretability, explainability and data-model dependency. Phys Med Biol. (2022) 67:11TR01. doi: 10.1088/1361-6560/ac678a
58
LuLAhmedFSAkinOLukLGuoXYangHet al. Uncontrolled confounders may lead to false or overvalued radiomics signature: a proof of concept using survival analysis in a multicenter cohort of kidney cancer. Front Oncol. (2021) 11:638185. doi: 10.3389/fonc.2021.638185
59
BowenSRFlynnRTBentzenSMJerajR. On the sensitivity of IMRT dose optimization to the mathematical form of a biological imaging-based prescription function. Phys Med Biol. (2009) 54:1483–501. doi: 10.1088/0031-9155/54/6/007
60
ShererMVLinDElguindiSDukeSTanL-TCacicedoJet al. Metrics to evaluate the performance of auto-segmentation for radiation treatment planning: A critical review. Radiotherapy Oncol. (2021) 160:185–91. doi: 10.1016/j.radonc.2021.05.003
61
GuoHWangJXiaXZhongYPengJZhangZet al. The dosimetric impact of deep learning-based auto-segmentation of organs at risk on nasopharyngeal and rectal cancer. Radiat Oncol. (2021) 16:113. doi: 10.21203/rs.3.rs-328649/v1
62
TillerNBMarconARZenoneMKiddKEJeukendrupAEMasterZet al. Generative artificial intelligence-driven chatbots and medical misinformation: an accuracy, referencing and readability audit. BMJ Open. (2026) 16:e112695. doi: 10.1136/bmjopen-2025-112695
63
ShaversVLBrownML. Racial and ethnic disparities in the receipt of cancer treatment. J Natl Cancer Ins. (2002) 94:334–57. doi: 10.1093/jnci/94.5.334
64
FahimYAHasaniIWKabbaSRagabWM. Artificial intelligence in healthcare and medicine: clinical applications, therapeutic advances, and future perspectives. Eur J Med Res. (2025) 30:848. doi: 10.1186/s40001-025-03196-w
65
AdabaraISadiqBOShuaibuANDanjumaYIManintiV. Trustworthy agentic AI systems: a cross-layer review of architectures, threat models, and governance strategies for real-world deployment. F1000Research. (2025) 14:905. doi: 10.12688/f1000research.169927.1
66
HuqMSFraassBADunscombePBGibbonsJPJrIbbottGSMundtAJet al. The report of Task Group 100 of the AAPM: Application of risk analysis methods to radiation therapy quality management. Med Phys. (2016) 43:4209–62. doi: 10.1118/1.4947547
67
SellenAHorvitzE. The rise of the ai co-pilot: Lessons for design from aviation and beyond. Commun ACM. (2024) 67:18–23. doi: 10.1145/3637865
68
BrownNACareyCHGerryEI. FDA releases action plan for artificial intelligence/machine learning-enabled software as a medical device. J Robot Artif Intell Law. (2021) 4:255.
69
BradyAPAllenBChongJKotterEKottlerNMonganJet al. Developing, purchasing, implementing and monitoring AI tools in radiology: practical considerations. A multi-society statement from the ACR, CAR, ESR, RANZCR & RSNA. Can Assoc Radiol J. (2024) 75:226–44. doi: 10.1177/08465371231222229
70
HurkmansCBibaultJ-EBrockKKvan ElmptWFengMFullerCDet al. A joint ESTRO and AAPM guideline for development, clinical validation and reporting of artificial intelligence models in radiation therapy. Radiotherapy Oncol. (2024) 197:110345. doi: 10.1016/j.radonc.2024.110345
Summary
Keywords
AI, clinical decision support, human-AI collaboration, radiation oncology, trustworthy AI, workflow integration
Citation
Zhu M, Gou L, Zhang C and Zheng D (2026) Trustworthy artificial intelligence in radiation oncology: cross-industry lessons for development, validation, and deployment. Front. Oncol. 16:1912362. doi: 10.3389/fonc.2026.1912362
Received
18 June 2026
Revised
09 July 2026
Accepted
13 July 2026
Published
23 July 2026
Volume
16 - 2026
Edited by
Chunhao Wang, Duke University, United States
Reviewed by
Maria F. Chan, Memorial Sloan Kettering Cancer Center, United States
Bhavna Singla, Erie County Medical Center, United States
Updates
Copyright
© 2026 Zhu, Gou, Zhang and Zheng.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Dandan Zheng, dandan_zheng@urmc.rochester.edu
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.