HYPOTHESIS AND THEORY article

Front. Toxicol.

Sec. Computational Toxicology and Informatics

A Pragmatic Workflow for Integrating Large Language Models into Toxicological Risk Assessment

  • 1. Innovatune srl, Padova, Italy

  • 2. Preclinical Safety, Sanofi, Frankfurt, Germany

  • 3. In silico Drug Discovery Lab (I2D Lab), Department of Pharmaceutical Sciences, University of Perugia, Perugia, Italy

  • 4. Paul Bradley consulting, Royston, United Kingdom

  • 5. Merck Healthcare KGaA, Darmstadt, Germany

  • 6. Department of Pharmaceutical Sciences, University of Basel, Basel, Switzerland

  • 7. Global Toxicology Services, KaLy-Cell, a company of GBA Pharma, Plobsheim, France

  • 8. Global Regulatory Affairs, Sciences & Strategy, FMC, Geneva, Switzerland

  • 9. Global Regulatory & Computational Toxicology Team, Hazard Communication and Chemicals Regulations, Merck Life Science S.r.l., Milan, Italy

  • 10. Stantec Consulting Services Inc., Irvine, CA, United States

  • 11. GlaxoSmithKline Medicines Research Centre, GSK, Stevenage, United Kingdom

  • 12. Global Pharmacology and Toxicology, Viatris, Morgantown, WV, United States

  • 13. Environment and Health Department, Istituto Superiore di Sanità, Rome, Italy

  • 14. Pfizer Worldwide Research, Development and Medical, Sandwich, United Kingdom

  • 15. Department of Environmental and Public Health Sciences, University of Cincinnati, Cincinnati, OH, United States

  • 16. Ron Steigerwalt Regulatory Tox Consulting LLC, Palm Springs, CA, United States

  • 17. Strategic Health Sciences, TRC Companies, Pittsburgh, PA, United States

  • 18. BIBRA Toxicology Advice and Consulting Ltd, Wallington, United Kingdom

  • 19. Informence Labs LLC, St. Petersburg, FL, United States

The final, formatted version of the article will be published soon.

Abstract

Artificial intelligence (AI), particularly large language models (LLMs), offers promising support for toxicological risk assessment (TRA), which integrates multiple lines of evidence to support health-based conclusions, including those used in regulatory decision-making. Effective use of LLMs in TRA requires a clear understanding of their strengths and limitations, especially in the context of regulatory expectations for transparency, reproducibility, and accountability. This work explores the emerging role of LLMs in TRA through three integrated components: an overview of LLMs in the context of TRA, an exemplar case study examining LLM-assisted derivation (mainly using ChatGPT) of the Permitted Daily Exposure (PDE) for acetyl tributyl citrate under the draft ICH Q3E guideline, and a workflow hypothesis for the integration of LLMs into TRA. Observations from the case study suggest that LLMs may support pattern recognition, targeted information retrieval, and rapid summarization of large bodies of text, facilitating literature triage, extraction of study details, and evidence summaries. At the same time, the case study highlights important limitations for LLMs. Observed variability in LLM outputs reflects both inherent model stochastic behavior, model limitations, and the intrinsic complexity of TRA, where different experts may reasonably reach distinct, yet scientifically defensible conclusions based on the available evidence. The observations from the case study, interpreted in the context of the available literature, are used to inform a hypothesis that LLMs may be most effectively integrated into TRA through structured, modular, human-supervised workflows. This hypothesis reflects the view that the suitability of LLM-assisted TRA depends not only on LLM capabilities, but also on how LLMs are integrated into the assessment process to address both LLM variability and the interpretive nature of TRA. The proposed pragmatic workflow separates data search, information extraction, and integrated summarization while maintaining human-in-the-loop control over interpretation, expert judgment, and final conclusions.

Summary

Keywords

AI, human-in-the-loop, LLM, PDE, Pragmatic workflow, Toxicological risk assessment

Received

09 May 2026

Accepted

12 August 2026

Copyright

© 2026 Bassan, Amberg, Barreca, Bradley, Bringezu, De Paula Souza, Guicheney, Lo Piparo, Kovarich, Massarsky, Maundrill, Molnar, Parenti, Parris, Pavan, Reichard, Steigerwalt, Tcheremenskaia, Unice, Waine, Whelan, White and Myatt. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.

*Correspondence: Arianna Bassan

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Share article

Article metrics