The field of medical natural language processing (NLP) and information retrieval is undergoing a rapid transformation fueled by advances in large language models (LLMs). While these technologies have dramatically improved the accessibility and summarization of medical knowledge, their integration into high-stakes domains like healthcare remains limited by persistent concerns around reliability and trustworthiness. Current systems struggle with well-documented limitations, including hallucinated content, inconsistent generalization across medical subdomains, difficulty expressing uncertainty, and challenges maintaining provenance. Such issues are particularly acute in clinical settings, where errors may have serious implications for patient safety, clinical decision-making, and public trust. Despite emerging research in retrieval-augmented generation and evidence-grounded outputs, there is still a pressing need for principled, holistic approaches that address not just technical capability, but also transparency, usability, and alignment with real-world workflows.
This Research Topic aims to advance the methodological and practical foundations of trustworthy medical LLMs and evidence-centered medical information systems. We seek to catalyze research that rigorously tackles the multifaceted challenges of reliability, calibration of uncertainty, and evidence grounding, while accounting for local contexts such as those shaped by European governance and regulatory perspectives. The objective is to foster new evaluation frameworks that accurately capture factuality, safety, and clinical utility in realistic settings, bridge robust retrieval and citation mechanisms with principled algorithmic improvements and illuminate human-centered trade-offs through applied studies. Key goals include developing methods that align generated statements with verifiable medical sources, promoting interoperability and generalizability across institutions and languages, and supporting safe, effective use through transparent system design and meaningful evaluation metrics.
To gather further insights into the intersection of trustworthy medical LLMs, robust information retrieval, and human-centered design, we welcome articles addressing, but not limited to, the following themes:
• Trustworthy medical LLMs: factuality, uncertainty, calibration, robustness, and safety
• Domain-specific pretraining and tokenization for medical language
• Retrieval models, indexing techniques, and query understanding in medical information retrieval
• Evidence-grounded generation: managing citation correctness and controlled summarization
• Benchmarks and datasets for medical NLP and information retrieval
• Robustness and generalization across clinical institutions, subdomains, and multilingual or low-resource settings
• Human-centered design: usability, explainability, user interaction, and stakeholder needs
• Human-in-the-loop evaluation and clinically meaningful performance metrics
• Privacy-preserving, governance-aware development and deployment workflows
• Real-world case studies, including systems deployment, failure analysis, and post-deployment monitoring
We welcome a broad range of submissions, including new algorithms, reproducible pipelines, theoretical analyses, benchmark or dataset articles, and real-world applied studies focusing on the reliability and utility of medical NLP and IR in clinical, research, and public health contexts.
Article types and fees
This Research Topic accepts the following article types, unless otherwise specified in the Research Topic description:
Brief Research Report
Case Report
Clinical Trial
Editorial
FAIR² Data
General Commentary
Hypothesis and Theory
Methods
Mini Review
Articles that are accepted for publication by our external editors following rigorous peer review incur a publishing fee charged to Authors, institutions, or funders.
Article types
This Research Topic accepts the following article types, unless otherwise specified in the Research Topic description:
Brief Research Report
Case Report
Clinical Trial
Editorial
FAIR² Data
General Commentary
Hypothesis and Theory
Methods
Mini Review
Opinion
Original Research
Perspective
Policy and Practice Reviews
Policy Brief
Review
Study Protocol
Systematic Review
Technology and Code
Keywords: trustworthy medical AI, medical language models, evidence-grounded retrieval, human-centered design, clinical decision support, retrieval-augmented generation (RAG), evaluation
Important note: All contributions to this Research Topic must be within the scope of the section and journal to which they are submitted, as defined in their mission statements. Frontiers reserves the right to guide an out-of-scope manuscript to a more suitable section or journal at any stage of peer review.