Abstract
Introduction:
Artificial intelligence (AI) is becoming central to genomics and multi-omics, but its concepts, architectures, applications, evaluation standards, and translational requirements remain fragmented. This scoping review mapped how AI is defined and operationalized in genomic science, including machine learning, deep learning, graph-based methods, foundation models, and large language models, and synthesized their data modalities, applications, evaluation practices, interpretability strategies, and governance challenges.
Methods:
We conducted a PRISMA-ScR scoping review with Joanna Briggs Institute guidance. Eligible studies applied AI to genomics or closely allied omics in research, clinical, or public health contexts. MEDLINE/PubMed, Embase, and supplementary registers were searched from January 2001 to 3 September 2025 without language restrictions. Records were screened in duplicate, and standardized items were extracted, including AI concept or method family, omics modality, task, metrics, interpretability, governance, and deployment considerations. Methodological reporting and quality were appraised using design-appropriate JBI tools and summarized descriptively as a normalized 0%–100% checklist-fulfillment index.
Results:
From 3,785 records, 1,040 studies were included. Publication remained sparse until 2017 and then expanded steeply, with more than 90% appearing from 2018 onward. The normalized JBI checklist-fulfillment index was modest overall (mean 35.3%, SD 20.1; range 7.5%–87.5%) and was interpreted descriptively, not as a directly comparable quality score across designs. Conceptually, the field has moved from feature-engineered statistical learning toward representation learning systems modeling nucleotide sequences, regulatory context, single-cell states, multi-omics profiles, biomedical text, and clinical-genomic knowledge. Applications concentrated on variant interpretation, regulatory genomics, multi-omics integration, single-cell analysis, pathology/radiology-genomics fusion, and genomic decision support, with increasing use of deep learning, graph models, foundation models, and LLMs. Calibration, external validation, mechanistic interpretability, ancestry-aware fairness, privacy protection, and deployment models for sensitive genomic data were unevenly reported; prospective multisite evaluations were rare.
Discussion:
AI in genomics has scaled rapidly since 2017–2018, but translation remains constrained by heterogeneous concepts, inconsistent benchmarks, incomplete reporting, and limited governance. Priorities include biologically meaningful benchmarks; calibrated uncertainty for genomic decision support; mechanism-linked interpretability; ancestry- and site-aware validation; privacy-preserving analysis of sensitive genomic data; and human oversight for variant interpretation, precision medicine, and public health genomics.
Systematic Review Registration:
1 Introduction
Artificial intelligence (AI) has become an increasingly important methodological layer in genomic science, supporting tasks that range from variant calling, regulatory annotation, and gene expression prediction to multi-omics integration, single-cell analysis, biomarker discovery, and clinical decision support. Earlier applications in genomics were largely built on feature-engineered statistical learning, including support vector machines, random forests, regularized regression, and probabilistic sequence models, which remain useful for structured genomic data, smaller cohorts, and interpretable clinical pipelines (; ; ; ). Over the last decade, however, the field has shifted toward deep representation learning, with convolutional, recurrent, attention-based, and transformer architectures learning regulatory motifs, variant effects, chromatin signals, and long-range sequence dependencies directly from biological data (; ; ; ). This transition has expanded the analytical scope of genomics from task-specific prediction pipelines toward systems capable of integrating sequence, epigenomic, transcriptomic, proteomic, imaging, literature-based, and clinical information for discovery and translational inference (; ; ; ).
More recently, foundation models and large language models (LLMs) have introduced a new conceptual and technical layer to genomic AI. DNA and RNA language models such as the Nucleotide Transformer and GENA-LM exemplify the use of large-scale pretraining to learn generalizable representations from genomic sequences (; ). Single-cell and multi-omics foundation models, including scGPT and related large-scale architectures, extend this paradigm to cellular states and cross-modal biological representation learning (; ; ). Tool-augmented and retrieval-augmented LLMs, such as GeneGPT and variant summarization systems grounded in curated databases, illustrate a complementary direction in which language models are connected to external biomedical resources to support genomic question answering, literature curation, variant interpretation, and evidence-grounded reporting (; ; ). These developments are promising, but they also introduce unresolved challenges related to factuality, calibration, hallucination, benchmark comparability, sensitive genomic data governance, and human oversight in high-stakes clinical or public health contexts (; ; ).
Several prior reviews have addressed important portions of this landscape. Reviews of deep learning in computational biology and genomics have clarified how neural networks can learn regulatory elements, noncoding variant effects, chromatin accessibility, and transcriptional regulation directly from sequence or epigenomic data (; ; ; ). Other syntheses have focused on AI in cancer genomics, precision oncology, histopathology-genomics integration, and radiogenomics, highlighting applications in biomarker discovery, patient stratification, survival prediction, and molecularly guided treatment selection (; ; ; ). A further body of work has reviewed multi-omics integration, single-cell analysis, and machine learning across heterogeneous molecular modalities, emphasizing the value of deep generative models, graph-based methods, and multimodal fusion for complex biological systems (; ; ; ). More recent papers have begun to synthesize LLMs, foundation models, code-generation benchmarks, and biomedical model evaluation, but these discussions often remain focused on specific model classes, datasets, or technical benchmarks rather than the full genomic research and translation ecosystem (; ; ; ; ).
Despite these contributions, the literature remains fragmented in ways that limit a panoramic understanding of AI in genomic science. First, prior reviews often focus on one methodological family, such as traditional machine learning, deep learning, graph models, or LLMs, rather than comparing how these families coexist across genomic tasks and evidence types. Second, many reviews are organized around a single data modality or application area, such as cancer, rare disease diagnosis, single-cell omics, imaging-genomics fusion, or variant interpretation, making it difficult to evaluate cross-cutting patterns in data modalities, model adaptation strategies, evaluation metrics, and deployment readiness (; ; ; ). Third, key translational dimensions are unevenly integrated into existing syntheses, particularly probability calibration, external validation, ancestry and site bias, interpretability linked to biological mechanisms, privacy-preserving computation, protected health information handling, and governance for clinical implementation (; ; ; ; ). Consequently, it remains unclear whether the field is moving toward reproducible and clinically actionable genomic AI or whether many applications remain proof-of-concept systems with limited validation, comparability, and governance.
A scoping review is therefore appropriate because AI in genomics is methodologically broad, conceptually heterogeneous, and not suitable for a narrow meta-analysis of pooled effectiveness. The field includes classical machine learning, convolutional and recurrent neural networks, graph neural networks, transformer-based sequence models, foundation models, LLMs, retrieval-augmented systems, and tool-using agents. These methods are applied across DNA and RNA sequences, epigenomics, bulk and single-cell transcriptomics, proteomics, spatial assays, imaging-derived molecular signals, biomedical literature, and clinical narratives when these are used to support genomic inference. They are also evaluated with heterogeneous outcomes, including AUROC, AUPRC, F1 score, Matthews correlation coefficient, calibration metrics, expert review, factuality assessment, safety evaluation, and task-specific biological benchmarks (; ; ; ). Mapping this diversity requires a design that can organize concepts, architectures, applications, evaluation practices, and governance issues without assuming that the included studies are directly comparable in design, population, task, or outcome.
The present review addresses this gap by systematically mapping the conceptual, methodological, and translational landscape of AI in genomics and closely allied omics domains. Specifically, we examine how AI methods are defined and operationalized across the literature; which model families and architectures are most frequently used; which genomic data modalities and application domains are prioritized; how performance, interpretability, calibration, and uncertainty are evaluated; and how governance issues, including privacy, fairness, protected health information, and deployment context, are reported. By integrating these dimensions, this review contributes a panoramic synthesis of AI in genomic science, with particular attention to the transition from classical machine learning and deep learning toward graph-based, transformer-based, foundation-model, and LLM-enabled systems. The aim is not to rank models or make pooled claims of comparative effectiveness, but to clarify how the field is structured, where evidence is strongest, and which methodological and governance gaps must be addressed before AI systems can be responsibly translated into genomic research, clinical genomics, precision medicine, and public health genomics.
2 Methods
2.1 Protocol and registration
This review was designed as a scoping review and conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) guidelines (). The protocol was prospectively registered in the Open Science Framework (OSF; https://osf.io/uexzh) to ensure transparency and reproducibility. The review was guided by the following research question: How has artificial intelligence been applied to genomics and allied omics domains, and what are the main concepts, architectures, applications, evaluation practices, and open challenges reported in literature? Because the field of AI in genomics is methodologically broad, rapidly evolving, and heterogeneous in study design, we complemented the scoping review framework with descriptive quantitative mapping, bibliometric summaries, exploration topic modeling, and contextual methodological appraisal. These components were used to support evidence charting and interpretation, not to convert the review into a comparative effectiveness review or meta-analysis. Specifically, bibliometric analyses were used to describe temporal and geographic publication patterns; topic modeling was used as an exploratory tool to corroborate thematic patterns identified during charting; and critical appraisal was used to contextualize reporting and methodological variability across sources. No pooled effect estimates were calculated, and no included study was excluded based on methodological appraisal.
2.2 Eligibility criteria
We considered studies that explicitly applied artificial intelligence to genomics or closely allied omics domains across human and non-human systems and in research, clinical, public health, agricultural, or translational settings. For this review, “genomics” referred to analyses involving DNA or RNA sequence data, genome structure, gene regulation, genetic variation, genome annotation, gene or variant prioritization, genome-wide association, comparative genomics, functional genomics, or genome-informed decision support. “Closely allied omics” was operationally defined as omics modalities directly connected to genomic inference or genome-enabled interpretation, including epigenomics, transcriptomics, single-cell transcriptomics, spatial transcriptomics, proteomics, metagenomics, microbiome profiles, and multi-omics datasets in which genomic or transcriptomic information was a central component.
Eligible designs included original investigations, methodological and benchmarking papers, conference proceedings, reviews, scoping reviews, systematic reviews, perspectives, and preprints, provided that they contained substantive methodological, empirical, evaluative, or conceptual content relevant to AI in genomics. This broad inclusion was intentional and consistent with the aim of mapping a rapidly evolving field rather than estimating pooled intervention effects. To account for heterogeneity, studies were charted and synthesized according to study type, AI method family, omics modality, application domain, evaluation strategy, interpretability approach, and governance or deployment content. Reviews and perspectives were used primarily to map concepts, frameworks, challenges, and reported trends, whereas original, methodological, and benchmarking studies were used to characterize concrete models, datasets, tasks, and evaluation practices.
Text sources and clinical narratives were eligible only when they were explicitly used to support genomic inference, genomic decision support, variant interpretation, gene prioritization, literature-grounded genomic reporting, bioinformatics code generation, or integration of genomic findings with clinical or phenotypic information. For example, LLM studies using biomedical literature, clinical reports, phenotypic descriptions, or curated variant databases were included only when the task was directly linked to genomic or omics interpretation. Studies focused exclusively on general clinical NLP, radiology, pathology, or electronic health records without a genomic, transcriptomic, epigenomic, proteomic, metagenomic, or multi-omics linkage were excluded.
Large language and foundation models were included when they were trained on, evaluated with, or deployed for genomics-relevant tasks, including DNA/RNA/protein sequence modeling, variant interpretation, regulatory annotation, single-cell or multi-omics analysis, genomic question answering, literature curation for genomic evidence, or bioinformatics pipeline/code generation. We excluded studies whose AI applications were unrelated to genomics or allied omics, records lacking explicit omics content, and editorials or opinion pieces without methodological, empirical, evaluative, or conceptual substance. When both a preprint and a peer-reviewed version were available, the peer-reviewed article was retained. When only a preprint existed but provided sufficient methodological detail, including data, model, task, and evaluation description, it was eligible because of the rapid development cycle of AI and foundation-model research. Duplicate reports were consolidated, and inaccessible full texts were excluded after reasonable retrieval efforts.
2.3 Information sources
We searched MEDLINE/PubMed and Embase using database specific controlled vocabulary (MeSH/Emtree) and harmonized free text terms tailored to each platform. To capture gray literature and fast moving developments, we ran targeted searches in bioRxiv, medRxiv, and arXiv and screened conference abstracts and technical reports identified through society websites and institutional repositories. Reference lists of key reviews and benchmark papers were hand searched to maximize recall. No language restrictions were applied. We did not contact study authors to obtain additional data or pursue literature via direct outreach; all sources were obtained exclusively from bibliographic databases and public repositories.
All sources were searched from January 2001 through 03 September 2025. This time window (2001–2025) was selected because it encompasses the emergence, maturation, and consolidation of modern machine learning, deep learning, and genomics pipelines, capturing the full evolution of AI driven genomic analysis from early computational frameworks to contemporary foundation models. Full reproducible search strategies for each source, including exact queries and applied filters, are provided in Supplementary Material 1. Duplicates across databases and sources were removed before screening using automated matching followed by manual verification.
2.4 Search strategy
We combined controlled vocabulary, including MeSH and Emtree terms, with harmonized free-text terms and Boolean operators searched in titles and abstracts. No language restrictions were applied, and the prespecified search window covered January 2001 to 3 September 2025. The last search was executed on 3 September 2025, in accordance with PRISMA-ScR recommendations to report the most recent search date.
The search strategy was intentionally broad because the terminology used to describe AI in genomics has changed substantially over time. Earlier studies often used terms such as computational intelligence, machine intelligence, neural networks, computer reasoning, or knowledge representation, whereas more recent studies use terms such as machine learning, deep learning, transformer architectures, foundation models, and large language models. Similarly, genomic applications may be indexed under genomics, functional genomics, comparative genomics, structural genomics, transcriptomics, multi-omics, bioinformatics, or precision medicine. Therefore, we prioritized sensitivity during database searching to avoid missing older terminology, emerging model classes, or interdisciplinary studies that may not yet be consistently indexed.
To balance recall with specificity, broad retrieval was followed by duplicate removal, independent title/abstract screening, full-text eligibility assessment, and explicit exclusion of records that did not contain a genomics or closely allied omics component. During full-text review, studies were retained only when the AI method was directly linked to genomic or omics data, genomic inference, genome-informed clinical or public health decision support, variant interpretation, gene or pathway prioritization, or bioinformatics tasks. This staged approach allowed us to maximize capture of a rapidly evolving and inconsistently indexed field while removing records where AI was unrelated to genomic science.
Database-specific syntax, including explosion of subject headings, truncation, and adjacency operators where supported, was applied in MEDLINE/PubMed and Embase. Complete reproducible strategies for all sources are provided in Supplementary File 1. To ensure reproducibility, the complete PubMed strategy as executed is provided below. PubMed (MEDLINE) Search Strategy: (((((((Genomics [Title/Abstract]) OR (Comparative Genomics [Title/Abstract])) OR (Genomics, Comparative [Title/Abstract])) OR (Structural Genomics [Title/Abstract])) OR (Genomics, Structural [Title/Abstract])) OR (Functional Genomics [Title/Abstract])) OR (Genomics, Functional [Title/Abstract])) AND (((((((((((((((((((((Artificial Intelligence [Title/Abstract]) OR (Intelligence, Artificial [Title/Abstract])) OR (Computer Reasoning [Title/Abstract])) OR (Reasoning, Computer[Title/Abstract])) OR (AI [Title/Abstract])) OR (Machine Intelligence [Title/Abstract])) OR (Intelligence, Machine [Title/Abstract])) OR (Computational Intelligence [Title/Abstract])) OR (Intelligence, Computational [Title/Abstract])) OR (Computer Vision Systems [Title/Abstract])) OR (Computer Vision System [Title/Abstract])) OR (System, Computer Vision [Title/Abstract])) OR (Systems, Computer Vision [Title/Abstract])) OR (Vision System, Computer [Title/Abstract])) OR (Vision Systems, Computer [Title/Abstract])) OR (Knowledge Acquisition (Computer) [Title/Abstract])) OR (Acquisition, Knowledge (Computer) [Title/Abstract])) OR (Knowledge Representation (Computer) [Title/Abstract])) OR (Knowledge Representations (Computer) [Title/Abstract])) OR (Representation, Knowledge (Computer) [Title/Abstract])).
2.5 Selection of sources of evidence
All records were imported into Rayyan (Qatar Computing Research Institute) for systematic management. Duplicates were removed automatically and verified manually. Title and abstract screening were conducted independently by two reviewers (WFR and AGN), with disagreements resolved by a third reviewer. Eligible studies underwent full text review, also in duplicate, with reasons for exclusion documented. The process is summarized in a PRISMA 2020 flow diagram (Figure 1).
FIGURE 1
2.6 Data charting process
We extracted data using a structured form refined in a pilot on 20 studies to ensure consistent field definitions and coding rules. Data charting was performed in Covidence, which supported standardized form management, duplicate entry, discrepancy resolution, and controlled updates across reviewers. Two reviewers independently charted each included record, documented uncertainties, and reconciled discrepancies through discussion; a senior reviewer adjudicated any remaining disagreements.
The form captured bibliographic details, AI method family, genomic or omics modality, study type, application domain, evaluation metrics, interpretability approaches, governance and deployment considerations, and reported limitations. Multi label fields were allowed where appropriate; missing or unclear items were coded as not reported. Periodic internal audits were conducted to verify consistency and completeness before locking the dataset for analysis.
2.7 Data items
We charted a common set of variables for every included study to support consistent synthesis and comparison. For identification, we recorded authors, year of publication, country or region, and venue. Methodologically, we classified the primary AI approach into traditional machine learning, deep learning, graph-based models, or large language and other foundation models. We captured the genomic or omics modality addressed, including DNA or RNA sequence, epigenomics, bulk and single cell transcriptomics, proteomics, imaging derived signals, and text sources such as biomedical literature and clinical narratives when used to support genomic inference. We coded the application focus, spanning variant calling and interpretation, regulatory genomics, gene and variant prioritization, multi omics integration, clinical decision support, and program synthesis or code generation for bioinformatics tasks. Evaluation practices were abstracted as intrinsic metrics, such as accuracy, area under the ROC or precision recall curves, F1 score, Brier score, and probability calibration, and extrinsic assessments that included expert review, factuality checks, and safety or risk evaluations. We documented interpretability and governance elements, noting the use of attribution or explanation methods, explicit citation of evidence, handling of protected health information, de identification procedures, federated or on premises training, and references to regulatory frameworks. Finally, we recorded limitations and failure modes reported by the authors, including dataset or ancestry bias, hallucination or factual errors, reproducibility concerns, and computational or resource constraints. Multi label entries were permitted where applicable to reflect the heterogeneity of designs and outcomes.
2.8 Critical appraisal of individual sources of evidence
Critical appraisal is optional in scoping reviews; however, we included a structured methodological appraisal to contextualize the reporting quality and methodological variability of the included evidence. The purpose of this appraisal was descriptive and interpretive, not exclusionary. Appraisal results were not used to determine eligibility, to weight studies statistically, or to make direct claims about comparative effectiveness.
We used the Joanna Briggs Institute (JBI) Critical Appraisal Tools, selecting the checklist most appropriate to each study design. Two reviewers appraised each study independently and reconciled disagreements by consensus. Because the review included heterogeneous source types, including original studies, methodological papers, benchmarking studies, reviews, conference papers, perspectives, and preprints, we did not interpret JBI scores as directly interchangeable across designs. Instead, appraisal was used to identify broad patterns in reporting completeness, methodological transparency, validation practices, and potential limitations within each relevant evidence category.
For descriptive transparency, we summarized JBI appraisal results using a normalized percentage index from 0 to 100, calculated as the number of “Yes” items divided by the number of applicable items on the relevant checklist and multiplied by 100. Items marked “Not applicable” were excluded from the denominator, and items marked “Unclear” were treated as not meeting the criterion. This normalized index was used only as a pragmatic descriptive summary of checklist fulfillment within the constraints of each design-specific tool. It should not be interpreted as a universal quality score, as proof of methodological equivalence across study designs, or as a quantitative weight in the synthesis.
Where appraisal results are reported by strata, these strata are intended to contextualize the maturity and reporting completeness of the evidence base, not to re-rank studies or pool findings across designs. The full appraisal spreadsheet, including item-level decisions and applicable checklist domains for each study, is available in the public OSF repository linked to the protocol.
2.9 Synthesis of results
We synthesized the evidence using descriptive mapping, structured narrative synthesis, and exploration thematic analysis aligned with the scoping review objective. Given the heterogeneity of source types, model families, omics modalities, tasks, datasets, and evaluation metrics, no meta-analysis or pooled comparative effectiveness analysis was planned or conducted. Instead, synthesis was organized around five analytic lenses; concepts and definitions used to operationalize AI in genomics; method families and architectures, including traditional machine learning, deep learning, graph-based methods, foundation models, and LLMs; data modalities and application areas, including genomic sequence, epigenomics, bulk and single-cell transcriptomics, proteomics, metagenomics, spatial assays, imaging-derived molecular signals, and text used for genomic inference; evaluation, interpretability, and reliability practices, including discrimination metrics, calibration, external validation, attribution, explainability, factuality, and uncertainty; and governance and deployment considerations, including privacy, protected health information, de-identification, federated or on-premises training, fairness, and regulatory framing.
To account for heterogeneity, findings were charted and summarized by study type, AI method family, omics modality, application domain, and evidence function. Original investigations, methods papers, and benchmarking studies were used primarily to summarize implemented models, datasets, tasks, and evaluation practices. Reviews, perspectives, and conceptual papers were used primarily to map definitions, frameworks, emerging challenges, and translational considerations. Conference papers and preprints were interpreted cautiously and were used to capture emerging developments only when they reported sufficient methodological detail. This stratified synthesis prevented narrative claims from treating all source types as equivalent evidence.
Quantitatively, we conducted descriptive and bibliometric analyses to map temporal trends, geographic distribution, study type composition, and cumulative method adoption. These analyses were used to characterize the structure and evolution of literature rather than to infer causality or comparative effectiveness. Topic modeling of study objectives was used as an exploratory adjunct to support the thematic synthesis. Specifically, latent Dirichlet allocation was applied to identify recurring terms and topic patterns, which were then compared with manually charted themes. Topic modeling results were not treated as independent proof of thematic importance; rather, they served as a reproducibility-oriented check on whether machine-assisted text patterns were consistent with reviewer-coded themes.
All analytical workflows were implemented in R 4.3.2 (x86_64-w64-mingw32). Core packages included tidyverse 2.0.0 for data management, ggplot2 3.5.2 and ggrepel 0.9.6 for graphics, and the topicmodels package for latent Dirichlet allocation, with the exact version recorded in the repository sessionInfo output. Random seeds were fixed and reported in the analytical code. De-identified summary tables, CSV files, and figure scripts were exported for reproducibility, with exact queries and processing steps cross-referenced to Supplementary Material 1 and to the public OSF repository linked to the protocol.
3 Results
3.1 Scale of the literature and quality gradient across AI–genomics studies
A total of 3,785 records were identified across MEDLINE/PubMed and Embase. After removal of 1,518 duplicates, 2,267 records were screened by title and abstract, of which 1,183 were excluded. We sought 1,084 reports for full-text retrieval; 19 were unavailable. Among 1,065 full texts assessed for eligibility, 25 were excluded because they were not directly related to AI in genomics, resulting in 1,040 included studies (Figure 1).
Most exclusions occurred during title/abstract screening (1,183/2,267; 52.2%). The non-retrieval rate was low (19/1,084; 1.8%), and only 25 of 1,065 retrieved full texts were excluded at eligibility (2.3%). Methodological quality scores ranged from 7.5% to 87.5%, with a mean of 35.31% (SD 20.06%), corresponding to an overall low-to-moderate methodological quality profile with substantial dispersion. Most studies clustered in the low-to-intermediate range, whereas only a minority reached scores ≥70%. Per-study and domain-level methodological quality values are available in the OSF repository associated with this review.
3.2 Thematic distribution of AI methods, genomic modalities, and application domains
For each included study, we charted bibliographic information, study design, principal application area, AI method family, genomic or multiomic modality, target task, evaluation metrics, interpretability strategy, and governance or deployment considerations. Aggregated findings from these extracted variables are summarized below, while the complete per-study dataset is available in the Open Science Framework repository associated with the review protocol.
Across the included corpus, AI applications in genomics were distributed across five broad method families: classical machine learning, deep learning, graph-based models, foundation models, and LLM/RAG-based systems. Classical machine learning was most often associated with structured or feature-engineered data, including variant annotations, tabular omics, polygenic risk-related features, and clinical-genomic reporting tasks. Deep learning was mainly used in sequence-to-function modeling, regulatory variant prediction, protein representation learning, histopathology-genomics fusion, biomarker discovery, and multimodal prediction. Graph-based models were concentrated in molecular networks, regulatory connectivity, cell–cell relationships, and multiomic integration. Foundation models were increasingly represented in DNA, RNA, protein, methylome, single-cell, and multiomic applications. LLM/RAG-based systems appeared mainly in variant summarization, genomic question answering, clinical reporting, literature curation, and gene or variant prioritization.
The genomic and multiomic modalities most frequently represented included DNA and RNA sequences, variant data, gene expression matrices, epigenomic profiles, single-cell RNA/ATAC data, protein sequences, imaging-genomics data, and integrated multiomic datasets. Application domains were similarly diverse, including regulatory annotation, gene prediction, variant effect prediction, pathogenicity classification, cancer genomics, rare disease interpretation, single-cell annotation, biomarker discovery, risk prediction, clinical decision support, and biomedical knowledge extraction.
Evaluation strategies varied according to task. Classification and risk prediction studies commonly reported accuracy, AUROC, AUPRC, F1 score, Matthews correlation coefficient, sensitivity, specificity, and calibration metrics. Regulatory annotation and gene-prediction studies frequently used PR-AUC, MCC, and task-specific benchmark scores. Single-cell and multiomic integration studies commonly reported clustering and integration metrics, including normalized mutual information, adjusted Rand index, adjusted mutual information, silhouette, and batch-mixing scores. LLM/RAG-oriented studies used heterogeneous metrics, including accuracy, expert agreement, factuality assessment, retrieval quality, summarization metrics, and Pass@K for code-generation tasks.
When studies directly compared method families, reported performance advantages were task dependent. Deep learning and foundation models were more often reported as advantageous in studies involving long-range sequence context, sparse single-cell data, complex nonlinear interactions, or multimodal integration. However, classical machine learning remained competitive in smaller, tabular, cleaner, or more interpretable settings. Therefore, the evidence does not support a universal hierarchy of model superiority; rather, it indicates that performance depends on the interaction among data modality, sample size, task structure, validation design, and reporting quality. These patterns are summarized in Table 1, which provides a structured overview of method families, genomic modalities, application domains, evaluation metrics, and recurrent limitations across the included corpus.
TABLE 1
| Method family | Main genomic modality | Common applications | Typical metrics | Recurrent limitation |
|---|---|---|---|---|
| Classical ML | Tabular omics, variant features, PRS, clinical-genomic variables | Risk prediction, variant classification, expression classification, benchmarking baselines | AUROC, accuracy, F1-score, MCC, sensitivity/specificity | Dependence on feature engineering; limited modeling of complex nonlinear structure |
| Deep learning | DNA/RNA sequence, imaging, protein data, multiomics | Regulatory prediction, variant effect prediction, histopathology-genomics fusion, biomarker discovery | AUROC, AUPRC, F1-score, correlation, calibration metrics | Larger data requirements; interpretability and external validation challenges |
| Graph models | Molecular networks, regulatory graphs, cell–cell relationships | Multiomics integration, regulatory inference, biomarker discovery, pathway-level modeling | AUROC, clustering metrics, pathway-level validation, integration scores | Dependence on graph construction and prior biological knowledge |
| Foundation models | DNA, RNA, protein, methylome, single-cell, multiomics | Transfer learning, annotation, gene prediction, single-cell integration, regulatory modeling | MCC, AUPRC, NMI, ARI, task-specific benchmark scores | Compute demand, limited external validation, calibration gaps |
| LLM/RAG systems | Genomic text, variants, ClinVar/ClinGen, biomedical literature, EHR-linked data | Variant summarization, Q&A, clinical reporting, gene prioritization, code generation | Accuracy, expert agreement, factuality, retrieval metrics, Pass@K | Hallucination, retrieval coverage, governance, uncertainty quantification |
Summary of AI method families, genomic modalities, applications, evaluation patterns, and recurrent limitations across included studies.
AI, artificial intelligence; ML, machine learning; PRS, polygenic risk score; AUROC, area under the receiver operating characteristic curve; AUPRC, area under the precision-recall curve; F1-score, harmonic mean of precision and recall; MCC, matthews correlation coefficient; DNA, deoxyribonucleic acid; RNA, ribonucleic acid; LLM, large language model; RAG, retrieval-augmented generation; ClinVar, Clinical Variation database; ClinGen, Clinical Genome Resource; EHR, electronic health record; NMI, normalized mutual information; ARI, adjusted Rand index; Q&A, question answering; Pass@K, pass at k.
3.3 Bibliometric growth and geographic distribution of AI–genomics publications
A complementary bibliometric analysis was conducted using MEDLINE/PubMed, Embase, Web of Science, and SciELO to characterize temporal and geographic publication patterns in AI–genomics research from 2001 to August 2025. This analysis included 1,131 records. Annual output increased from 2 records in 2001 to 260 records during January–August 2025, corresponding to a 130-fold increase. Most records were published from 2018 onward (1,068/1,131; 94.4%).
Trend analyses supported a two-phase temporal pattern. Across the full period, annual publication rate was strongly correlated with calendar year (Spearman ρ = 0.823, 95% CI 0.621–0.923; p < 0.0001). Before 2017, the trend was not statistically significant (Spearman ρ = 0.396; p = 0.129; slope = 0.0179 percentage points/year). In contrast, the 2017–2025 period showed a near-monotonic increase (Spearman ρ = 1.000; p < 0.0001), with a steeper slope of 2.720 percentage points/year (95% CI 2.201–3.240; R2 = 0.956; p < 0.0001). These results identify 2017–2018 as the main inflection period in AI–genomics publication growth (Figure 2).
FIGURE 2

Temporal growth of AI–genomics publications, 2001–August 2025. Annual publication rates were estimated from MEDLINE/PubMed, Embase, Web of Science, and SciELO. The dashed vertical line marks the 2017–2018 inflection period identified by segmented trend analyses. Labels were restricted to selected reference years to improve readability. The 2025 estimate reflects January–August only.
Geographically, AI–genomics publications were concentrated in three major regions (Figure 3A). North America accounted for the largest share of records (38.16%), followed by Asia (32.42%) and Europe (24.53%). Together, these three regions contributed approximately 95.12% of the global AI–genomics bibliometric record set. Smaller shares were observed for Oceania (3.59%), South America (0.72%), and Africa (0.57%), indicating a marked geographic imbalance in scientific production.
FIGURE 3

Geographic distribution of AI–genomics publications. (A) Continental distribution of records, showing concentration in North America, Asia, and Europe. (B) Top contributing countries, with countries outside the top 15 grouped as “Other”. Percentages represent shares of the global AI–genomics bibliometric record set. The 2025 estimate reflects January–August only. Percentages may not sum to exactly 100% because of rounding.
At the national level, the distribution was also highly concentrated and right-skewed (Figure 3B). The United States accounted for approximately one-third of all records (34.15%), followed by China (14.49%) and India (8.32%); together, these three countries represented approximately 56.96% of the total output. A second group of contributors included Italy, the United Kingdom, Canada, Australia, and Germany, each contributing between approximately 3% and 5% of the global record set. Additional countries contributed smaller shares, generally close to or below 2%, forming a long tail of lower-frequency national contributions. Overall, the top 10 countries accounted for approximately 80% of global output, and the top 20 for approximately 90%, indicating that AI–genomics publication activity was concentrated in a relatively small number of national producers.
Percentages represent relative shares of the global AI–genomics bibliometric record set and may not sum to exactly 100% because of rounding.
3.4 Temporal thematic patterns and text-mining findings
The thematic trajectory of the studies included three broad phases. First, studies published around 2015–2016 were mainly conceptual or infrastructural, focusing on EMR architecture, secure genomic data access, sequence-sharing infrastructure, and early ML applications in omics-scale biomarker discovery. Second, from 2017 to 2020, the literature shifted toward more concrete translational applications, including tumor sequencing interpretation, radiogenomics, regulatory genomics, variant calling, and early clinical decision-support systems. Third, from 2021 onward, the corpus increasingly emphasized multimodal integration, single-cell and spatial omics, graph-based learning, foundation models, explainability, federated learning, and governance. Representative high-volume or clinically oriented studies included AI-supported interpretation of 1,018 NG oncology cases, exome-based psychiatric genomics involving 5,090 samples, and pan-cancer multimodal prognostic modeling across 14 tumor types.
Text mining of study objectives identified high-frequency bigrams related to EMR infrastructure, mobile/PHR ecosystems, stakeholder-oriented architecture, and personalized medicine. The most frequent bigrams included “next-generation,” “generation architecture,” “electronic medical,” “medical records,” “records EMR,” “mobile apps,” “apps PHRs,” “serves users,” “incorporates trends,” and “personalized medicine” (Figure 4A).
FIGURE 4

Text-mining and topic-modeling patterns extracted from study objectives. (A) Most frequent bigrams identified after tokenization and Portuguese/English stop-word and domain-filler removal. The leading bigrams were mainly related to next-generation EMR architecture, mobile applications, personal health records, stakeholder-oriented systems, and personalized medicine. (B) LDA topic modeling of study objectives using k = 8 topics. Top terms by β weight were grouped into four recurring motifs: EMR modernization, mobile/PHR ecosystem, architecture and stakeholders, and personalized medicine. LDA was trained on study objectives using terms appearing in at least 10 documents; β represents term weight within each topic and γ represents topic proportion per document.
LDA topic modeling with k = 8 yielded topics that could be grouped into four recurring motifs: EMR modernization, mobile/PHR data capture, architecture and stakeholders, and personalized medicine. The leading terms across these topics included “emr,” “electronic,” “records,” and “incorporates” for EMR modernization; “apps,” “mobile,” “phrs,” and “users” for mobile/PHR ecosystems; “architecture,” “serves,” “stakeholders,” and “propose” for system design; and “personalized,” “medicine,” “trends,” and “next/generation” for personalized medicine (Figure 4B).
Together, these text-mining patterns indicate that a subset of the corpus framed AI–genomics integration through health informatics infrastructure, data capture systems, and personalized medicine architectures, rather than through model development alone.
3.5 Historical distribution of AI method milestones in genomics
A milestone-based synthesis of 41 peer-reviewed records published between 2004 and 2024 showed a temporal shift in the dominant AI approaches used in genomics and closely related omics domains. Earlier milestones were mainly represented by traditional machine learning or other/unspecified statistical learning approaches, including HMMs, support vector machines, random forests, and related probabilistic or feature-engineered methods. From the mid-2010s onward, deep learning milestones accumulated more rapidly, particularly in regulatory genomics, variant calling, proteomics, computational pathology, and multimodal prediction.
Figure 5A positions selected peer-reviewed milestones by year and highlights representative methodological transitions, including early neural-network applications in proteomics, DeepSEA for regulatory variant prediction, DeepVariant for germline variant calling, weakly supervised whole-slide imaging models, multimodal histology–genomics fusion, Vision Transformers in pathology, and emerging foundation-model applications in genomics. Figure 5B summarizes cumulative adoption by method family through 2024. Cumulative milestone counts were highest for Other/Unspecified methods (n = 19) and Deep Learning (n = 17), followed by Foundation/XAI approaches (n = 3) and Traditional ML (n = 2).
FIGURE 5

Milestones and cumulative adoption of AI in genomics (2004–2024). (A) Milestone timeline. Each dot marks a peer-reviewed milestone (n = 41) positioned by year; the vertical jitter has no quantitative meaning and is used only to avoid overlap. Color/shape encode the method family: Traditional ML, Deep Learning, Foundation/XAI, and Other/Unspecified (e.g., HMM/kNN/unspecified or mixed ML). Canonical items are labeled, including early deep nets for proteomics (2006), DeepSEA for regulatory variant prediction (2015), DeepVariant for germline calling (2018), weakly supervised WSI cancer detection (2019), multimodal histology–genomics fusion (PathomicFusion, 2021), Vision Transformers in pathology (2021), and the emergence of foundation models in genomics (2022), alongside reports of FDA cleared AI tools entering clinical workflows (2023). (B) Cumulative counts by method family. Step curves show the accumulation of milestones per year. By 2024 the totals are: Other/Unspecified (n = 19), Deep Learning (n = 17), Foundation/XAI (n = 3), and Traditional ML (n = 2). The trajectory highlights a mid-2010s inflection driven by deep learning, followed by early but growing activity around foundation models and explainability/federated approaches. Abbreviations: CNN, convolutional neural network; ViT, Vision Transformer; WSI, whole slide imaging; FM, foundation model; XAI, explainable AI; NGS, next-generation sequencing.
Overall, this milestone mapping indicates a shift from earlier feature-engineered and probabilistic approaches toward deep learning, multimodal modeling, and emerging foundation/XAI-oriented methods. These patterns should be interpreted as a descriptive synthesis of selected methodological milestones rather than as evidence of universal superiority of any model family across genomic tasks (Figures 5A,B).
3.6 Method-family-specific patterns: traditional machine learning
Traditional machine learning studies were mainly associated with structured, feature-engineered genomic and omics datasets, including variant annotations, gene expression profiles, GWAS/PRS-derived variables, microbiome features, epigenomic markers, regulatory annotations, and clinical-genomic variables. The most frequently reported algorithmic families included support vector machines, random forests, gradient boosting frameworks, regularized regression models, decision trees, k-nearest neighbors, Naive Bayes, hidden Markov models, and Bayesian networks. Across the corpus, these approaches were generally used when predictors could be represented as tabular, sequence-derived, pathway-derived, or biologically annotated features.
The main application domains included phenotype and clinical risk prediction, variant pathogenicity and functional-effect classification, GWAS/PRS-related modeling, gene expression-based classification, somatic mutation calling, microbiome and 16 S rRNA-based classification, epigenomic state classification, regulatory modeling, and rare-disease diagnostic workflows. Traditional ML was also frequently used as a benchmarking strategy against more complex architectures, particularly in studies involving modest sample sizes, structured omics data, or settings in which feature-level interpretability was important.
Reported evaluation practices most often included AUROC, AUPRC, accuracy, sensitivity, specificity, F1-score, Matthews correlation coefficient, R2, RMSE, MAE, and, less frequently, calibration metrics. Cross-validation, out-of-bag estimation for random forests, and external or held-out validation were reported in studies with more rigorous evaluation designs. Recurrent methodological considerations included feature engineering, dimensionality reduction, feature selection, normalization, missing-data handling, class-imbalance correction, ancestry or population-structure adjustment, and prevention of data leakage.
Across the studies included, traditional ML was most often positioned as a competitive baseline or preferred approach for smaller, tabular, high-dimensional, or feature-engineered genomic datasets. Its recurrent advantages were interpretability, data efficiency, and compatibility with biologically meaningful predictors, whereas recurrent limitations included dependence on feature engineering, sensitivity to preprocessing choices, limited capacity to model complex nonlinear or long-range biological structure, and reduced generalizability when validation was restricted or population structure was insufficiently addressed.
3.7 Method-family-specific patterns: deep learning
Deep learning studies were concentrated in sequence-to-function modeling, regulatory variant prediction, chromatin accessibility, transcription factor binding, histone-mark prediction, gene expression prediction, splicing-related tasks, methylation detection, variant calling, single-cell analysis, multimodal integration, histopathology–genomics fusion, and other genomic signal-annotation tasks. The main data modalities included raw DNA/RNA sequence, one-hot nucleotide matrices, k-mer or learned embeddings, NGS and ONT read-derived representations, pileup images, epigenomic tracks, gene expression matrices, single-cell omics, and multimodal combinations of sequence, chromatin, expression, imaging, and clinical data.
CNN-based architecture was frequently reported for motif discovery, local sequence features, chromatin accessibility modeling, transcription factor binding prediction, and regulatory variant interpretation. Recurrent architecture, including LSTM- and GRU-based models, were used to capture broader contextual dependencies across genomic windows. Attention-based and Transformer-derived models were increasingly associated with long-range regulatory interactions, enhancer–promoter inference, gene expression prediction from extended genomic contexts, and multi-task outputs. Autoencoders and variational autoencoders were mainly represented in unsupervised representation learning, denoising, dimensionality reduction, and single-cell or multiomic integration, whereas graph-aware variants were used to model cellular or molecular relationships.
Across the corpus, representative deep learning systems included DeepSEA, Basset, ExPecto, Enformer, and DeepVariant, reflecting recurrent applications in sequence-to-chromatin prediction, variant-effect estimation, long-range regulatory modeling, and variant calling. Training practices reported across studies included one-hot or binary encoding of DNA sequences, k-mer embeddings, pileup-image representations for variant calling, reverse-complement augmentation or shared-weight designs, dropout, early stopping, weight decay, class weighting, undersampling, synthetic examples, and augmentation strategies for noisy genomic labels. Several studies also emphasized train/validation/test splitting strategies designed to avoid coordinate leakage or overoptimistic performance estimates.
Evaluation most often relied on AUROC, AUPRC, accuracy, F1-score, correlation metrics for continuous genomic signals, chromosome-held-out validation, k-fold validation, platform-specific benchmarking for variant calling, and, less frequently, calibration assessment. Interpretability strategies included in silico mutagenesis, Integrated Gradients, DeepLIFT, saliency or CAM-based approaches, SHAP, motif extraction from convolutional filters, and inspection of attention maps. Recurrent limitations included larger data requirements, sensitivity to class imbalance and noisy labels, risk of data leakage, limited external validation, limited calibration reporting, and variability in the biological validation of model-derived explanations.
3.8 Method-family-specific patterns: foundation models and LLM/RAG systems
Foundation-model and LLM/RAG studies were concentrated in DNA/RNA sequence modeling, epigenomic prediction, single-cell annotation, variant summarization, genomic question answering, clinical reporting, literature curation, gene and variant prioritization, perturbation reasoning, pathway inference, drug sensitivity prediction, and automated code generation for bioinformatics pipelines. The models reported across this subset included general-purpose LLMs such as GPT-4/4o, Llama variants, and Qwen, as well as domain-specific genomic models, including DNABERT-2, Nucleotide Transformer, GENA-LM, Omni-DNA, EpiGePT, GeneGPT, and Geneverse. These systems were applied to both discriminative and generative tasks, using DNA and RNA sequences, protein sequences, single-cell transcriptomics, epigenomic tracks, structured variant resources, biomedical literature, EHR-linked clinical text, spatial transcriptomics, and imaging-related data.
Architecturally, Transformer-based models predominated, including encoder, decoder, and encoder–decoder designs. Long-context and domain-adapted architectures were used for extended genomic sequences and regulatory modeling, while tool-augmented and retrieval-augmented systems were used for tasks requiring access to external genomic or clinical knowledge bases. Adaptation strategies included domain-specific pretraining, supervised fine-tuning, instruction tuning, reinforcement learning from human feedback, LoRA/PEFT-based parameter-efficient tuning, retrieval-augmented generation over curated resources such as ClinVar and ClinGen, and API-based tool integration. In the included studies, RAG and tool-augmented approaches were primarily used to improve factual grounding, provenance, and access to curated genomic evidence.
Reported performance varied substantially by task, benchmark, and evaluation design. Selected studies reported 17% top-50 accuracy for GPT-4 in rare-disease gene prioritization, average MCC of 0.767 for Omni-DNA across 18 of 26 benchmark tasks, F1 ≈ 93.7 for GENA-LM in promoter prediction, PCC ≈0.787 for EpiGePT in cross-cell-type epigenomic prediction, recall of 97.85% for fine-tuned models in causative variant identification, near-perfect expert agreement for RAG-grounded variant summaries, and Pass@20 ≈ 50% for GPT-4 in code-generation tasks. These values indicate strong performance in selected settings, but they do not establish uniform superiority across genomic tasks because benchmarks, datasets, input modalities, and validation designs differ widely.
Evaluation metrics included accuracy, F1-score, MCC, Pearson correlation, AUROC, AUPRC, BLEU/ROUGE, BERTScore, distributional distances for single-cell annotation, Pass@K for code generation, retrieval quality measures, factuality checks, expert review, and safety scoring. Benchmarking efforts included Genomic Touchstone for DNA/RNA/protein modeling, BioCoder for code generation, CARDBiomedBench for biomedical Q&A and safety, and specialized datasets for variant summarization and enhancer prediction. However, heterogeneity in preprocessing, benchmark design, task definitions, and reporting practices limited direct comparability across studies.
Interpretability and grounding strategies were mainly based on evidence citation, attribution to retrieved sources, RAG pipelines, saliency mapping for sequence models, motif-level relevance analysis, and API-based tool use for traceability. Recurrent limitations included hallucination, incomplete retrieval coverage, overclassification bias, dataset and domain shifts, inconsistent safety and fairness evaluation, limited calibration, limited uncertainty quantification, compute demands, reproducibility gaps, and incomplete biomedical knowledge coverage. Zero-shot performance was reported as insufficient for several structured omics or clinical genomics tasks, with fine-tuning, retrieval augmentation, or human expert review often required for more reliable outputs.
Governance and deployment considerations were unevenly reported. Although many genomic LLMs and foundation models were released through open repositories or model hubs, explicit discussion of protected health information, de-identification, federated training, on-premises deployment, clinical accountability, and human-in-the-loop oversight was limited. Overall, the reviewed studies show rapid expansion of foundation-model and LLM/RAG approaches in genomics, but also highlight persistent gaps in standardized evaluation, calibration, interpretability, factual grounding, governance, and clinical validation.
4 Discussion
This scoping review mapped 1,040 studies on artificial intelligence (AI) in genomics and multiomics, covering classical machine learning, deep learning, graph-based models, foundation models, and LLM/RAG-based systems. The revised Results show that the field is broad, rapidly expanding, and methodologically heterogeneous. Three findings are particularly important for interpreting the literature. First, the evidence base is large but uneven in methodological quality, with scores ranging from 7.5% to 87.5% and a mean of 35.31% (SD 20.06%), indicating an overall low-to-moderate quality profile. Second, AI applications are distributed across distinct method families, genomic modalities, application domains, and evaluation practices, rather than forming a single homogeneous field. Third, bibliometric and thematic analyses indicate that AI–genomics research accelerated sharply after 2017–2018 and increasingly shifted toward multimodal integration, single-cell and spatial omics, foundation models, explainability, and governance. Together, these findings directly address the aims of this review: to map the main AI method families, characterize their genomic and multiomic applications, and examine how evaluation, interpretability, translation, and governance have been addressed across the literature.
A central implication of these results is that AI methods in genomics should be interpreted according to task, data modality, sample size, validation design, and reporting quality. The evidence does not support a universal hierarchy in which one model family is consistently superior across genomic tasks. Instead, the extracted patterns indicate a division of labor among method families. Classical ML was most frequently associated with structured or feature-engineered variables, including tabular omics, variant annotations, polygenic risk score-related features, microbiome profiles, epigenomic markers, and clinical-genomic variables. This pattern is consistent with studies showing that support vector machines, random forests, regularized regression, and gradient-boosting models remain useful in high-dimensional but structured genomic settings, particularly when interpretability, smaller sample sizes, and feature-level explanations are important (
In contrast, deep learning and foundation models were more often reported as advantageous in settings involving long-range sequence context, nonlinear interactions, sparse single-cell data, or multimodal integration. DeepSEA and related sequence-to-function models illustrate the use of deep neural networks for regulatory prediction and variant-effect estimation from DNA sequence features (
The same task-dependent interpretation applies to foundation models and LLM/RAG systems. The Results showed strong performance in selected benchmark settings, including MCC = 0.767 for Omni-DNA across 18 of 26 tasks, F1 ≈ 93.7 for GENA-LM in promoter prediction, PCC ≈0.787 for EpiGePT in cross-cell-type epigenomic prediction, and Pass@20 ≈ 50% for GPT-4 in code-generation tasks. These findings demonstrate the promise of large pretrained models in specific tasks, but they also highlight the heterogeneity of benchmarks, input modalities, and validation strategies. Therefore, the main implication is not that foundation models replace classical ML or deep learning, but that they expand the methodological repertoire for transfer learning, sequence modeling, variant summarization, single-cell annotation, biomedical question answering, and retrieval-grounded genomic reporting (
The application landscape identified in the Results was also structured rather than diffuse. AI–genomics applications were concentrated in recurrent domains, including regulatory annotation, variant calling and interpretation, cancer genomics, biomarker discovery, risk prediction, single-cell annotation, multiomic integration, clinical reporting, and biomedical knowledge extraction. This concentration suggests that AI is being repeatedly applied to areas where genomic data are high dimensional, heterogeneous, and difficult to interpret using conventional analytical pipelines. Cancer genomics and precision oncology remain especially prominent. AI-supported interpretation of 1,018 NG oncology cases illustrates the role of AI in clinical variant interpretation and trial matching (
Variant interpretation and genotype–phenotype mapping also emerged as central clinical genomics use cases. Benchmarking and review studies emphasize variant calling, curation concordance, explainability, and natural language processing for genomic interpretation (
One finding is that part of the AI–genomics literature is concerned not only with model development, but also with the information systems required to make AI-enabled genomics usable. The text-mining analysis identified high-frequency bigrams related to “electronic medical,” “medical records,” “records EMR,” “mobile apps,” “apps PHRs,” “serves users,” and “personalized medicine”. The LDA analysis similarly grouped study objectives into motifs related to EMR modernization, mobile/PHR data capture, architecture and stakeholders, and personalized medicine. These findings suggest that health informatics infrastructure represents an important translational layer for AI-enabled genomics. Genomic prediction alone is insufficient for clinical implementation; genomic and phenotypic information must be captured, harmonized, interpreted, and returned within usable clinical or public health workflows.
In this context, EMR-centered architectures, PHRs, mobile data capture, and stakeholder-oriented systems may function as integration points where genomic interpretation, risk stratification, decision support, and clinical reporting can be connected. However, the Results also show that governance and deployment considerations remain unevenly reported, particularly for foundation models and LLM/RAG systems. This creates a gap between model capability and implementation readiness. The implication is that future translation of AI-enabled genomics will depend not only on higher AUROC, AUPRC, MCC, or F1-score, but also on interoperability, provenance, privacy protection, auditability, and clinical usability. Legal and ethical analyses similarly emphasize consent, data-use governance, group risk, and GDPR-aligned data sharing in genomic AI contexts (
The most consistent translational bottlenecks identified across the review were limited external validation, heterogeneous benchmarks, incomplete calibration reporting, uneven interpretability, and underdeveloped governance. These limitations appeared across method families, but they were especially consequential for deep learning, graph-based models, foundation models, and LLM/RAG systems. The Results showed that evaluation strategies varied widely by task, including AUROC, AUPRC, F1-score, MCC, clustering metrics, integration scores, factuality assessment, retrieval metrics, expert agreement, and Pass@K. This diversity is appropriate given the range of tasks, but it also limits direct comparability across studies. For clinical or translational applications, internal discrimination metrics are not enough. Models used for risk prediction, variant interpretation, or clinical reporting also require calibration, external validation, and decision-relevant endpoints. Several studies reported cross-validation or held-out validation, but the broader evidence base remains weighted toward retrospective evaluations, benchmarking studies, and methodological demonstrations. This pattern explains why the review can map promising applications but should not be interpreted as establishing comparative clinical effectiveness across model families.
Interpretability remains similarly uneven. Classical ML offers feature-level interpretability through coefficients, tree importance, or simpler model structures, whereas deep learning and foundation models often require post-hoc methods such as saliency maps, in silico mutagenesis, DeepLIFT, Integrated Gradients, SHAP, attention inspection, or motif extraction. These methods can help connect model behavior to biological signals, but their clinical meaning is not always validated. In this respect, interpretability should not be treated as a visual add-on; it should be evaluated against biological plausibility, known regulatory mechanisms, experimental evidence, and clinical usability requirements. In silico mutagenesis and attribution-based methods have been used to interrogate sequence models and regulatory predictions (
Governance is the least mature dimension of the literature. The Results showed that LLM/RAG and foundation-model studies frequently discussed hallucination, retrieval coverage, uncertainty quantification, reproducibility, and safety, but explicit treatment of protected health information, de-identification, federated training, on-premises deployment, clinical accountability, and human-in-the-loop oversight was limited. For LLM/RAG systems, this is especially important because factuality and traceability depend not only on model architecture but also on retrieval quality, database coverage, prompt design, and evidence attribution. RAG and tool-augmented systems may improve provenance and access to curated genomic evidence, but they do not eliminate the need for expert review, uncertainty communication, and governance standards.
The complementary bibliometric analysis provides additional context for these methodological and translational patterns. AI–genomics publication output increased sharply after 2017–2018: annual output rose from 2 records in 2001 to 260 records during January–August 2025, and 1,068 of 1,131 bibliometric records were published from 2018 onward. This acceleration helps contextualize the methodological diversity observed in the scoping review, because the field expanded during the same period in which deep learning, multimodal integration, single-cell analysis, and foundation models became more prominent. However, this bibliometric trend should be interpreted as evidence of publication growth, not as proof of clinical maturity.
The geographic analysis also has important implications. The Results showed that AI–genomics publications were concentrated in North America, Asia, and Europe, with these three regions accounting for approximately 95% of the bibliometric record set. At the country level, the United States, China, and India accounted for more than half of the total output. This concentration suggests that the global evidence base may be shaped disproportionately by a limited set of national research ecosystems. Such asymmetry matters because genomic datasets, health infrastructures, ancestry composition, computing resources, and regulatory environments differ across regions. As a result, models trained and validated primarily in high-output countries may not generalize equally well to underrepresented populations or lower-resource health systems. This geographic imbalance also intersects with the governance and fairness concerns identified in the thematic synthesis. If AI–genomics tools are developed primarily in settings with greater computational resources, richer genomic databases, and stronger research infrastructure, then their deployment in less represented regions may reproduce or amplify existing inequities. Addressing this issue will require more inclusive datasets, transparent reporting of ancestry and site composition, regionally diverse validation cohorts, and collaborative governance models that include researchers and health systems from underrepresented regions.
The nature of the evidence base and the scoping approach entail several limitations. Although the search strategy covered MEDLINE/PubMed, Embase, Web of Science, and SciELO, and was supplemented by targeted searches in bioRxiv, medRxiv, and arXiv, the AI literature evolves rapidly. Important materials such as model cards, benchmark documentation, code repositories, technical reports, and updated checkpoints may fall outside conventional indexing, generating ascertainment bias toward peer-reviewed publications and highly visible preprints. In addition, consistent with the objectives of a scoping review, this study mapped concepts, methods, applications, and evidence patterns rather than estimating pooled model effects. Therefore, the review should not be interpreted as a meta-analysis of comparative model performance.
The methodological quality profile of the included studies also limits the strength of translational conclusions. The mean methodological quality score was 35.31%, with substantial dispersion, and only a minority of studies reached high-quality thresholds. This finding supports caution when interpreting broad claims about clinical utility, deployment readiness, or superiority of one method family over another. Heterogeneity in preprocessing, model families, benchmark datasets, split strategies, and evaluation metrics further complicates direct comparison. Studies varied in whether they used random, chromosome-held-out, site-held-out, or external validation designs; such differences can substantially affect estimated performance and generalizability. Reporting on interpretability, calibration, reproducibility, and governance was also uneven. Many studies mentioned interpretability tools, but fewer connected explanations to predefined biological or clinical criteria. Similarly, privacy, fairness, and protected health information were often acknowledged but less often operationalized through concrete governance procedures, federated designs, de-identification workflows, or auditing frameworks.
Reproducibility signals were also mixed. Code, model weights, dataset versions, preprocessing scripts, and curation procedures were not uniformly available. This issue is particularly important in foundation-model and LLM/RAG research, where model behavior may change with checkpoint versions, tokenizer updates, retrieval corpora, and prompt design. These limitations indicate constructive priorities for future research in the field. Future studies should use standardized and biologically meaningful benchmarks, report calibration and decision relevant outcomes in addition to discrimination metrics, adopt validation strategies that account for study site, technological platform, population ancestry and data source, and, whenever possible, make available versioned datasets, analytical code, model cards and evaluation protocols. For clinical and public health translation, prospective multicenter studies are particularly needed in areas such as variant interpretation, precision oncology, rare disease genomics, pharmacogenomics, infectious disease genomics and public health genomics. In the case of foundation models and large language model systems supported by retrieval augmented generation, future work should prioritize factual grounding, uncertainty quantification, transparency of retrieval sources, governance of protected data and oversight by human experts. These priorities arise directly from the main findings of this review, which show rapid growth, expanding methodological capacity, substantial heterogeneity, uneven validation and persistent gaps in interpretability, reporting and governance.
In summary, artificial intelligence in genomics has evolved from feature engineering and task specific models toward deep learning, multimodal systems, graph-based approaches and foundation model architectures. The field now demonstrates broad methodological diversity and promising performance in specific tasks. However, the evidence base remains heterogeneous and unevenly validated. Therefore, the most defensible conclusion is cautious: artificial intelligence is becoming increasingly central to genomic and multiomic research, but its responsible translation into clinical, research and public health contexts will depend on stronger external validation, clearer reporting, better calibration, biologically grounded interpretability, more inclusive datasets and explicit governance frameworks.
5 Final considerations
This scoping review mapped 1,040 studies on artificial intelligence in genomics and multiomics, providing a structured synthesis of method families, genomic modalities, application domains, evaluation practices, bibliometric trends, geographic distribution, topic-modelling patterns, and historical method trajectories. The findings show that AI–genomics research has expanded rapidly since 2017–2018 and now spans classical machine learning, deep learning, graph-based models, foundation models, and LLM/RAG-based systems. Rather than supporting a universal hierarchy of model superiority, the evidence indicates that model performance and suitability are task dependent, varying according to data modality, sample size, biological context, validation design, and reporting quality.
The main contribution of this review is to organize a highly heterogeneous and fast-moving field into interpretable evidence layers. Classical machine learning remains relevant for structured, feature-engineered, and interpretable genomic tasks; deep learning and graph-based approaches are prominent in sequence-to-function modeling, regulatory prediction, single-cell analysis, and multimodal integration; and foundation models and LLM/RAG systems are emerging in transfer learning, variant summarization, genomic question answering, clinical reporting, and retrieval-grounded knowledge synthesis. The review also highlights that AI–genomics research is geographically concentrated, methodologically uneven, and variable in evaluation practices, with recurrent gaps in external validation, calibration, interpretability, reproducibility, and governance.
Taken together, these findings position AI as an increasingly central component of genomic and multiomic research, but they also support a cautious interpretation of its translational readiness. The key take-home message is that the future value of AI in genomics will depend less on isolated performance gains and more on robust validation, transparent reporting, biologically meaningful benchmarks, interpretable outputs, inclusive datasets, and governance frameworks capable of supporting responsible clinical and public health use. By consolidating current applications, methodological patterns, and evidence gaps, this review provides a foundation for more reproducible, accountable, and clinically relevant AI-enabled genomics.
Statements
Data availability statement
The original contributions presented in this study are publicly available. All data underlying the analyses of methodological quality and results of the eligible studies included in this scoping review are deposited in the Open Science Framework (OSF) and can be accessed at: https://osf.io/uexzh. The full search strategies used for each database are provided in Supplementary File 1: Search Strategies, which accompanies this article as Supplementary Material.
Author contributions
WR: Conceptualization, Data curation, Formal Analysis, Investigation, Methodology, Software, Visualization, Writing – original draft, Writing – review and editing. MP: Data curation, Investigation, Validation, Visualization, Writing – review and editing. DP: Data curation, Investigation, Visualization, Writing – review and editing. LdS: Data curation, Formal Analysis, Software, Visualization, Writing – review and editing. PrR: Investigation, Resources, Validation, Writing – review and editing. PaR: Investigation, Resources, Validation, Writing – review and editing. MC: Investigation, Resources, Validation, Writing – review and editing. VA: Resources, Supervision, Writing – review and editing. SS: Methodology, Supervision, Writing – review and editing. RM: Formal Analysis, Methodology, Software, Validation, Writing – review and editing. AG-N: Conceptualization, Funding acquisition, Methodology, Project administration, Supervision, Writing – review and editing.
Funding
The author(s) declared that financial support was received for this work and/or its publication. This research was supported by the Fundação de Amparo à Pesquisa do Estado de Minas Gerais (FAPEMIG) through the Science, Technology, and Innovation Development Grant program, under the “Call 012/2023: Structuring Networks for Scientific Research or Technological Development”. The project “Applications of Genomics in the Context of One Health” (Grant No. RED-00181-23).
Acknowledgments
The authors gratefully acknowledge the support and institutional infrastructure provided by the Molecular and Computational Biology of Fungi Laboratory and the Department of Microbiology at the Universidade Federal de Minas Gerais; the Department of Computer Science at the Universidade Federal de Minas Gerais; the Expertise Centre for Leptospirosis/WOAH Reference Laboratory for Leptospirosis, Amsterdam University Medical Centers; the Laboratory of Cellular and Molecular Genetics at the Universidade Federal de Minas Gerais; and the Department of Microbiology, Immunology and Parasitology at the Federal University of Triângulo Mineiro. Their collective commitment to fostering high quality scientific research was essential for the development of this study.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Supplementary material
The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fbinf.2026.1829576/full#supplementary-material
References
1
AbdelwahabO.BelzileF.TorkamanehD. (2023). Performance analysis of conventional and AI-based variant callers using short and long reads. BMC Bioinforma.24 (24), 472. 10.1186/s12859-023-05596-3
2
AngermuellerC.PärnamaaT.PartsL.StegleO. (2016). Deep learning for computational biology. Mol. Syst. Biol.12, 878. 10.15252/msb.20156651
3
AvsecŽ.AgarwalV.VisentinD.LedsamJ. R.Grabska-BarwinskaA.TaylorK. R.et al (2021). Effective gene expression prediction from sequence by integrating long-range interactions. Nat. Methods18, 1196–1203. 10.1038/s41592-021-01252-x
4
BaiD.MoS.ZhangR.LuoY.GaoJ.YangJ. P.et al (2026). scLong: a billion-parameter foundation model for capturing long-range gene context in single-cell transcriptomics. Nat. Commun.17, 2380. 10.1038/s41467-026-69102-y
5
BaiãoA. R.CaiZ.PoulosR. C.RobinsonP. J.ReddelR. R.ZhongQ.et al (2025). A technical review of multi-omics data integration methods: from classical statistical to deep generative approaches. Briefings Bioinforma.26, bbaf355. 10.1093/bib/bbaf355
6
BallardJ.WangZ.LiW.ShenL.LongQ. (2024). Deep learning-based approaches for multi-omics data integration and analysis. BioData Min.17, 38. 10.1186/s13040-024-00391-z
7
BarshaiM.TriptoE.OrensteinY. (2020). Identifying regulatory elements via deep learning. Annu. Rev. Biomed. Data Sci.3, 315–338. 10.1146/annurev-biodatasci-022020-021940
8
BianchiO.WilleyM.AlvaradoC. X.DanekB.KhaniM.KuznetsovN.et al (2026). CARDBiomedBench: a benchmark for evaluating the performance of large language models in biomedical research. Lancet Digital Health8 (1), 100943. 10.1016/j.landig.2025.100943
9
CalvinoG.PeconiC.StrafellaC.TrastulliG.MegalizziD.AndreucciS.et al (2024). Federated learning: breaking Down barriers in global genomic research. Genes (Basel). Dec22, 15. 10.3390/genes15121650
10
ChapmanC. R.QuinnG. P.NatriH. M.BerriosC.DwyerP.OwensK.et al (2025). Consideration and disclosure of group risks in genomics and other data-centric research: does the common rule need revision?Am. J. Bioeth.25, 47–60. 10.1080/15265161.2023.2276161
11
ChenX.IshwaranH. (2012). Random forests for genomic data analysis. Genomics. 2012/06/01/99, 323–329. 10.1016/j.ygeno.2012.04.003
12
ChenR. J.LuM. Y.WilliamsonD. F. K.ChenT. Y.LipkovaJ.NoorZ.et al (2022). Pan-cancer integrative histology-genomic analysis via multimodal deep learning. Cancer Cell40, 865–878.e866. 10.1016/j.ccell.2022.07.004
13
CoghlanS.GyngellC.VearsD. F. (2024). Ethics of artificial intelligence in prenatal and pediatric genomic medicine. J. Community Genet. Feb15, 13–24. 10.1007/s12687-023-00678-4
14
CuiH.WangC.MaanH.PangK.LuoF.DuanN.et al (2024). scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nat. Methods21, 1470–1480. 10.1038/s41592-024-02201-0
15
Dalla-TorreH.GonzalezL.Mendoza-RevillaJ.Lopez CarranzaN.Henryk GrzywaczewskiA.OteriF.et al (2024). The nucleotide transformer: building and evaluating robust foundation models for. Hum. Genomics. Biorxiv.2023, 523679. 10.1101/2023.01.11.523679
16
DiasR.TorkamaniA. (2019). Artificial intelligence in clinical and genomic diagnostics. Genome Med.11, 70. 10.1186/s13073-019-0689-8
17
DingemansA. J. M.HinneM.JansenS.van ReeuwijkJ.de LeeuwN.PfundtR.et al (2022). “Phenotype based prediction of exome sequencing outcome using machine learning for neurodevelopmental disorders,” Genet. Med.24. 645–653. 10.1016/j.gim.2021.10.019
18
DuX.NagyA.OatesM. F.WangY.WangX.PlasekJ. M.et al (2026). Precision grounding: augmenting large language models with evidence-based databases for trustworthy genetic variant summarization. Int. J. Med. Inform.214, 106424. 10.1016/j.ijmedinf.2026.106424
19
FishmanV.KuratovY.ShmelevA.PetrovM.PenzarD.ShepelinD.et al (2025). GENA-LM: a family of open-source foundational DNA language models for long sequences. Nucleic Acids Res.53 (2), gkae1310. 10.1093/nar/gkae1310
20
FrascaM.La TorreD.RepettoM.De NicolòV.PravettoniG.CuticaI. (2024). Artificial intelligence applications to genomic data in cancer research: a review of recent trends and emerging areas. Discov. Anal.2 (2024/07/29), 10. 10.1007/s44257-024-00017-y
21
FritzscheM. C.AkyüzK.Cano AbadíaM.McLennanS.MarttinenP.MayrhoferM. T.et al (2023). Ethical layering in AI-driven polygenic risk scores-new complexities, new challenges. Front. Genet.14, 1098439. 10.3389/fgene.2023.1098439
22
GaoQ.YangL.LuM.JinR.YeH.MaT. (). The artificial intelligence and machine learning in lung cancer immunotherapy. J. Hematol. Oncol.16, 55. 10.1186/s13045-023-01456-y
23
GoldsteinB. A.PolleyE. C.BriggsF. B. (2011). Random forests for genetic association studies. Stat. Appl. Genet. Mol. Biol.10, 32. 10.2202/1544-6115.1691
24
HaddawayN. R.PageM. J.PritchardC. C.McGuinnessL. A. (2022). PRISMA2020: an R package and shiny app for producing PRISMA 2020‐compliant flow diagrams, with interactivity for optimised digital transparency and open synthesis. Campbell Systematic Reviews18, e1230. 10.1002/cl2.1230
25
HarfoucheA. L.JacobsonD. A.KainerD.RomeroJ. C.HarfoucheA. H.Scarascia MugnozzaG.et al (2019). Accelerating climate resilient plant breeding by applying next-generation artificial intelligence. Trends Biotechnol. Nov.37, 1217–1235. 10.1016/j.tibtech.2019.05.007
26
HuangS.CaiN.PachecoP. P.NarrandesS.WangY.XuW. (2018). Applications of support vector machine (SVM) learning in cancer genomics. Cancer Genomics - proteomics.15, 41–51. 10.21873/cgp.20063
27
JinQ.YangY.ChenQ.LuZ. (2024). GeneGPT: augmenting large language models with domain tools for improved access to biomedical information. Bioinformatics40, btae075. 10.1093/bioinformatics/btae075
28
Kan-TorY.DanzigerM. M.ZoharE.NinioM.ShimoniY. (2024). Does your model understand genes? A benchmark of gene properties for biological and text models. ArXiv.abs/2412.04075. 10.48550/arXiv.2412.04075
29
KuoT. T.GabrielR. A.CidambiK. R.Ohno-MachadoL. (2020). EXpectation propagation LOgistic REgRession on permissioned blockCHAIN (ExplorerChain): decentralized online healthcare/genomics predictive model learning. J. Am. Med. Inf. Assoc.27, 747–756. 10.1093/jamia/ocaa023
30
NicoraG.VitaliF.DagliatiA.GeifmanN.BellazziR. (2020). Integrated multi-omics analyses in oncology: a review of machine learning methods and tools. Front. Oncol.10, 1030. 10.3389/fonc.2020.01030
31
NovakovskyG.FornesO.SaraswatM.MostafaviS.WassermanW. W. (2023). ExplaiNN: interpretable and transparent neural networks for genomics. Genome Biol.24, 154. 10.1186/s13059-023-02985-y
32
OkserS.PahikkalaT.AirolaA.SalakoskiT.RipattiS.AittokallioT. (2014). Regularized machine learning in the genetic prediction of complex traits. PLoS Genet. Nov.10, e1004754. 10.1371/journal.pgen.1004754
33
PageM. J.McKenzieJ.BossuytP.BoutronI.HoffmannT.MulrowC.et al (2020). Updating Guidance for Reporting Systematic Reviews: Development of the PRISMA 2020 Statement.
34
PatelN. M.MicheliniV. V.SnellJ. M.BaluS.HoyleA. P.ParkerJ. S.et al (2018). Enhancing next-generation sequencing-guided cancer care through cognitive computing. Oncologist23, 179–185. 10.1634/theoncologist.2017-0170
35
PoplinR.ChangP.-C.AlexanderD.SchwartzS.ColthurstT.KuA.et al (2018). A universal SNP and small-indel variant caller using deep neural networks. Nat. Biotechnol.36, 983–987. 10.1038/nbt.4235
36
PrelajA.GanzinelliM.ProvenzanoL.MazzeoL.ViscardiG.MetroG.et al (2024). “APOLLO 11 Project, consortium in advanced lung cancer patients treated with innovative therapies: integration of real-world data and translational research,” Clin. Lung Cancer25. 190–195. 10.1016/j.cllc.2023.12.012
37
RahimzadehV.LawsonJ.RushtonG.DoveE. S. (2022). Leveraging algorithms to improve decision-making workflows for genomic data access and management. Biopreserv Biobank. Oct.20, 429–435. 10.1089/bio.2022.0042
38
ReelP. S.ReelS.PearsonE.TruccoE.JeffersonE. (2021). Using machine learning approaches for multi-omics data analysis: a review. Biotechnol. Adv.49 (2021/07/01/), 107739. 10.1016/j.biotechadv.2021.107739
39
RizviS. A.LevineD.PatelA.ZhangS.WangE.PerryC. J.et al (2025). Scaling large language models for next-generation single-cell analysis. bioRxiv2014, 648850. 10.1101/2025.04.14.648850
40
RohaimM. A.ClaytonE.SahinI.VilelaJ.KhalifaM. E.Al-NatourM. Q.et al (2020). Artificial intelligence-assisted loop mediated isothermal amplification (AI-LAMP) for rapid detection of SARS-CoV-2. Viruses. 1;12. 10.3390/v12090972
41
Roman-NaranjoP.Parra-PerezA. M.Lopez-EscamezJ. A. (2023). A systematic review on machine learning approaches in the diagnosis and prognosis of rare genetic diseases. J. Biomed. Inf. Jul143, 104429. 10.1016/j.jbi.2023.104429
42
ShrikumarA.GreensideP.KundajeA. (2017). Reverse-complement parameter sharing improves deep learning models for genomics. bioRxiv, 103663. 10.1101/103663
43
SokolovaK.ChenK. M.HaoY.ZhouJ.TroyanskayaO. G. (2024). Deep learning sequence models for transcriptional regulation. Annu. Rev. Genomics Hum. Genet.25, 105–122. 10.1146/annurev-genom-021623-024727
44
TangX.QianB.GaoR.ChenJ.ChenX.GersteinM. B. (2024). BioCoder: a benchmark for bioinformatics code generation with large language models. Bioinformatics40, i266–i276. 10.1093/bioinformatics/btae230
45
TelentiA.LippertC.ChangP. C.DePristoM. (2018). Deep learning of genomic variation and regulatory network data. Hum. Mol. Genet.27 (27), R63–r71. 10.1093/hmg/ddy115
46
ThudiM.PalakurthiR.SchnableJ. C.ChitikineniA.DreisigackerS.MaceE.et al (2021). Genomic resources in plant breeding for sustainable agriculture. J. Plant Physiol.257, 153351. 10.1016/j.jplph.2020.153351
47
TopçuoğluB. D.LesniakN. A.RuffinM. T.WiensJ.SchlossP. D. (2020). A framework for effective application of machine learning to microbiome-based classification problems, mBio11 (3), e00420–00434. 10.1128/mbio.00420-00434
48
TownendD. (2018). “Conclusion: harmonisation in genomic and health data sharing for research: an impossible dream?,” Hum. Genet. 137. 657–664. 10.1007/s00439-018-1924-x
49
UngerM.KatherJ. N. (2024). A systematic analysis of deep learning in genomics and histopathology for precision oncology. BMC Med. Genomics17, 48. 10.1186/s12920-024-01796-9
50
WangY.CaiZ.ZengQ.GaoY.OuyangJ.XuY.et al (2025). Genomic touchstone: benchmarking genomic language models in the context of the central dogma. bioRxiv. 2025.2006.2025.661622. 10.1101/2025.06.25.661622
51
Warnat-HerresthalS.PerrakisK.TaschlerB.BeckerM.BaßlerK.BeyerM.et al (2020). Scalable prediction of acute myeloid leukemia using high-dimensional machine learning and blood transcriptomics. Iscience. Jan.23, 100780. 10.1016/j.isci.2019.100780
52
WhalenS.SchreiberJ.NobleW. S.PollardK. S. (2022). Navigating the pitfalls of applying machine learning in genomics. Nat. Rev. Genet.23, 169–181. 10.1038/s41576-021-00434-9
53
XuJ.YangP.XueS.SharmaB.Sanchez-MartinM.WangF.et al (2019). Translating cancer genomics into precision medicine with artificial intelligence: applications, challenges and future perspectives. Hum. Genet.138, 109–124. 10.1007/s00439-019-01970-5
54
ZhouJ.TroyanskayaO. G. (2015). Predicting effects of noncoding variants with deep learning–based sequence model. Nat. Methods12, 931–934. 10.1038/nmeth.3547
55
ZhouL.SuominenH.GedeonT. (2019). Adapting state-of-the-art deep language models to clinical information extraction systems: potentials, challenges, and solutions. JMIR Med. Inf.25 (7), e11499. 10.2196/11499
Summary
Keywords
artificial intelligence, deep learning, foundation models, genomics, large language models, machine learning, multi-omics
Citation
Rodrigues WF, Parise MTD, Parise D, dos Santos LM, Ribeiro PS, Ristow P, Cardoso MS, Azevedo VAC, Soares SC, Minardi RCM and Goés-Neto A (2026) Artificial intelligence for genomic science: a scoping review of concepts, architectures, applications, and open challenges. Front. Bioinform. 6:1829576. doi: 10.3389/fbinf.2026.1829576
Received
13 March 2026
Revised
20 May 2026
Accepted
03 June 2026
Published
03 July 2026
Volume
6 - 2026
Edited by
Shahper Nazeer Khan, University of Manitoba, Canada
Reviewed by
Geet Madhukar, University of New Hampshire, United States
Adi Wijaya, Universitas Indonesia Maju, Indonesia
Updates

Check for updates
Copyright
© 2026 Rodrigues, Parise, Parise, dos Santos, Ribeiro, Ristow, Cardoso, Azevedo, Soares, Minardi and Goés-Neto.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Aristóteles Goés-Neto, arigoesneto@icb.ufmg.br; Wellington Francisco Rodrigues, wfr2023@ufmg.br
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.