Language understanding is inherently multimodal. Whether we read, listen, or converse, our brains go beyond words to draw on visual scenes, prosody, prior context, and background knowledge. This ability to merge linguistic and perceptual information allows us to resolve ambiguities, follow conversations, and construct rich mental representations of meaning. Research in cognitive neuroscience shows that the brain networks supporting language are not confined to traditional language areas; they also include systems for vision, motor control, and social cognition. In parallel, computational and corpus-based studies demonstrate that multimodal data—involving text, images, and audio—encode structural and semantic patterns that purely linguistic approaches cannot capture. With the recent rise of vision-language models (VLMs) and large language models (LLMs), researchers now have new computational tools that simulate how linguistic and perceptual information interact. However, these models also highlight crucial open questions: How do machine representations compare to those in the human brain? Where do they succeed in mirroring human meaning-making, and where do they fall short? A shared framework linking neural and computational perspectives is still missing.
This Research Topic aims to bring together work at the interface of neuroscience, psychology, linguistics, and artificial intelligence to understand how humans and machines process and represent language in multimodal contexts. It invites researchers to explore how the brain integrates verbal and non-verbal signals, and how AI systems approximate these processes. Key goals include identifying the brain areas and time scales involved in multimodal integration, assessing how computational models achieve or fail to achieve semantic grounding, and examining the ways context and perceptual information shape discourse and pragmatic inference. Ultimately, the aim is to build a more complete, mechanistic understanding of multimodal language processing—one that advances human cognitive science and supports the creation of AI systems that engage with language in genuinely human-compatible ways.
To gather further insights into the shared principles of language processing across the brain and computational systems, we welcome contributions from a range of disciplines and methods. Submissions may address, but are not limited to, the following themes: - Neural systems supporting the integration of visual, auditory, and linguistic information during comprehension and production - Predictive processing and representational similarity between human neural responses and multimodal AI models - Mechanisms of semantic grounding in LLMs and VLMs, and comparisons with perceptual grounding in the brain - Discourse, reference, and pragmatic inference across multimodal and interactive contexts - Methodological advances for benchmarking AI models against human behavior and neural data - Cross-linguistic and cultural variation in multimodal language integration - Applications in clinical communication, digital discourse, and the detection of multimodal misinformation
This Research Topic welcomes original research, reviews, theoretical perspectives, methods papers, and interdisciplinary commentaries. Studies may use behavioral, neuroimaging, or computational approaches, and should demonstrate clear relevance to the neural or computational mechanisms of multimodal language processing in humans, AI systems, or both.
Article types and fees
This Research Topic accepts the following article types, unless otherwise specified in the Research Topic description:
Brief Research Report
Case Report
Clinical Trial
Community Case Study
Conceptual Analysis
Curriculum, Instruction, and Pedagogy
Data Report
Editorial
FAIR² Data
Articles that are accepted for publication by our external editors following rigorous peer review incur a publishing fee charged to Authors, institutions, or funders.
Article types
This Research Topic accepts the following article types, unless otherwise specified in the Research Topic description:
Brief Research Report
Case Report
Clinical Trial
Community Case Study
Conceptual Analysis
Curriculum, Instruction, and Pedagogy
Data Report
Editorial
FAIR² Data
General Commentary
Hypothesis and Theory
Methods
Mini Review
Opinion
Original Research
Perspective
Policy and Practice Reviews
Policy Brief
Registered Report
Review
Study Protocol
Systematic Review
Technology and Code
Keywords: multimodal communication, language and vision, AI interaction, multimodal cognition, healthcare communication, human-AI interaction
Important note: All contributions to this Research Topic must be within the scope of the section and journal to which they are submitted, as defined in their mission statements. Frontiers reserves the right to guide an out-of-scope manuscript to a more suitable section or journal at any stage of peer review.