EDITORIAL article

Front. Plant Sci., 21 July 2026

Sec. Functional and Applied Plant Genomics

Volume 17 - 2026 | https://doi.org/10.3389/fpls.2026.1869724

Editorial: Utilizing machine learning with phenotypic and genotypic data to enhance effective breeding in agricultural and horticultural crops

  • 1. U. S. Department of Agriculture, U. S. National Arboretum, Beltsville, MD, United States

  • 2. Facultad de Ciencias Agrarias—Departamento de Ciencias Forestales, Universidad Nacional de Colombia—Sede Medellín, Medellín, Colombia

Introduction

Environmental variability, agrobiodiversity loss, and market volatility are placing unprecedented pressures on agricultural systems (Hultgren et al., 2025), increasing the vulnerability of agricultural crops and horticultural trees, and shifting market preferences (Benitez-Alfonso et al., 2023). Rising temperatures have contributed to prolonged droughts (López-Hernández et al., 2023) and heat (López-Hernández et al., 2025c) waves, along with shifts in precipitation patterns, increased soil salinity (Guo et al., 2023), heightened pest and pathogen pressures (Guevara-Escudero et al., 2021), all of which severely threaten crop productivity and stability. Shifts in seasonal patterns and extreme temperatures affect not only yield but also food quality (Abberton et al., 2016). In turn, consumer narrow preferences, and market volatility and erratic prices of major staple and commodity crops, add uncertainty to an already highly vulnerable scenario (Smale and Jamora, 2020). In response, plant breeding efforts are focusing on developing environmentally resilient varieties capable of withstanding these multifaceted challenges (Campbell et al., 2025).

In this context, the development of varieties resilient to varying environmental conditions, along with market-versatile crop types (Peláez et al., 2022), is recognized as essential to mitigate these challenges while ensuring global sustainable food security (Benitez-Alfonso et al., 2023). Historically, breeding for complex traits, such as abiotic and biotic stress tolerance, yield and quality, and water- and nutrient-use efficiency, has relied on phenotypic recurrent selection, which more recently has been complemented by marker-assisted selection (MAS) (Pinson et al., 2022; Bedoya-Londoño et al., 2025) derived from simple sequence repeat (SSR) markers (Warnke et al., 2023), and quantitative trait loci (QTL) mapping (Ahmar et al., 2021; Fernandez-Baca et al., 2021a; Barnaby et al., 2022; Cortés, 2024). With the advancement in computing capabilities and high-throughput genotyping, particularly next-generation sequencing (NGS), genome-powered solutions have come into play by enabling a more precise identification of quantitative trait nucleotides (QTNs) associated with complex traits (Vega-Muñoz et al., 2025), as well as their prediction through genomic selection (GS) across environments and complex genetic backgrounds (Arenas et al., 2025).

Despite substantial advances in sequencing technologies, analytical pipelines, and computer power, the systematic phenotypic evaluation of large populations under field conditions remains a major limitation (Mir et al., 2019). Traditional phenotyping methods are laborious, time- and budget-consuming, and often subjected to observer bias and environmental inconsistencies. Furthermore, phenotype and adaptive potential forecasting remain a major bottleneck due to its poor precision, variability, and complex interactions (Li et al., 2025). ML and AI now offer transformative potential in overcoming these challenges from the plant genetics perspective (Ma et al., 2014b) as well as the crop breeding discipline itself (Varshney, 2021). ML and AI can capture complex, nonlinear relationships between genotypes, phenotypes, and environmental factors (Libbrecht and Noble, 2015; Schrider and Kern, 2018), leveraging in this way computational power to analyze large datasets (Tong and Nikoloski, 2021). This integration not only promises to accelerate the breeding pipeline but also is likely to enhance the precision of trait selection, particularly for complex—often plastic—polygenic traits (Arenas and Cortés, 2026), including the adaptive potential (Cortés, 2025), and traits influenced by soil microbial associations (Fernandez-Baca et al., 2021b, 2021).

Given the rapid advances in molecular crop breeding, this Research Topic on “Utilizing Machine Learning with Phenotypic and Genotypic Data to enhance Effective Breeding in Agricultural and Horticultural Crops” highlights how ML and AI can scale phenotyping, genotyping and their integration across modern breeding applications. Specifically, the Research Topic aims to compile innovative case studies where advanced phenotyping and genotypic tools are integrated through ML, AI, and deep learning (DL) to enhance breeding efficiency. The 15-contributing works (Table 1), spanning a diverse array of study systems, illustrate how these analytical technologies can enhance phenotyping, trait prediction, gene discovery, genomic selection, and multi-omics integration (Figure 1). These contributions collectively illustrate how ML functions not only as data-analysis tool but also catalyst of modern crop improvement. Specifically, ML/AI help overcome persistent bottlenecks in scalable, objective phenotyping and in forecasting complex genotype–environment interactions, complementing GS and GWAS pipelines and supporting the shift toward data-integrated breeding.

Table 1

SpeciesKey trait(s)Main goalML/AI method(s)Key result(s)Perspective(s)Ref.
ML-assisted high-throughput phenotyping and imaging technologies
Buckwheat (Fagopyrum esculentum)Seed traits (shape, size, color)Study inheritance of seed traits via image-based phenotypingRGB imaging + ImageJ + Statistical analysisMaternal inheritance of shape/color, overdominance in sizeIntegrating AI-based segmentation and multimodal imagingOh et al.
SoybeanYieldPredict soybean yield from RGB imagesGBDT best among 5 ML algorithmsR2 = 0.82 with multi-angle, multi-type featuresML-driven phenotyping for scalable yield predictionLi et al.
BarleyPathogen interactionHigh-throughput microscopy for pathogen interactionBluVision Micro with RF classifiers & CNNs30× faster than HyphArea, GWAS linked to resistance lociAccelerate disease resistance screening with ML toolsLück et al.
Bentgrass (Agrostis stolonifera × A. capillaris)Tiller numberAutomated instance-level tiller quantification in hybrid populationsYOLOv8 (one-stage detector) vs. Faster R-CNN & edge-based segmentationYOLOv8 achieved R² = 0.97, highest accuracy and fastest inference, robust to dense occlusionYOLOv8 enables scalable instance-level phenotyping for downstream genetic analysis & breedingFerm et al.
Machine learning models for phenotypic forecasting
LycheeFruit traits & cultivar IDPhenotyping and cultivar classificationCNN-based segmentation models + SVM, RF, LDASegmentation Dice >0.90, LDA accuracy ~79%Automated fruit analysis and cultivar ID with multimodal imagingXue et al.; Kim et al., (2024)
MaizeDays to pollen, yieldTrait prediction across environmentsCompositional Autoencoder (CAE)7×-10× better accuracy than PCA/PLSRLatent feature disentanglement Enables precise trait prediction Powadi et al.
AI and ML in trait mapping and candidate gene identification
SorghumCallus induction, regenerationIdentify QTLs for in vitro culture traitsMulti-locus GWAS via ML-based methods34 QTLs + 47 candidate genes identifiedML-GWAS for complex traits and in vitro breedingXu et al.
WheatPlant heightGWAS for plant height in wheatMLM with admixture & kinship correction44 SNPs identified, 7 linked to key genesML-GWAS to dissect polygenic traits in pre-breedingDing et al.
ML-enabled genomic selection
Rice, sunflower, wheat, maizeMultiple agronomic traitsEnhance genomic selection with deep learningParallel CNN with multi-kernel design+0.031 accuracy over DNNGP in 24 traitsFuture work on data balancing and multi-trait GSXie et al.
AlmondShelling fractionPredict shelling fraction with XAIRF + SHAP for SNP interpretationRF R² = 0.511, key SNP linked to seed developmentML-XAI integration for better GS interpretabilityNovielli et al.
CoffeeYield, fruit number, pest and disease resistanceBoost GS with ensemble learningSEL (GBLUP, RF, MARS, QRF + Meta-learners)Up to 200% gain in trait predictionSEL captures complex trait architecture in GSNascimento et al.
Multi-omics integration and predictive breeding
SoybeanProtein contentIdentify bHLH genes linked to protein synthesisComparative genomics + expression profilingSeveral bHLH genes linked to protein regulationGene editing for seed quality improvement
WheatFlower developmentStudy MADS-box gene evolution and functionML-based clustering + expression profilingIdentified 15 TaE genes, functional validationTargeted editing for flower traits and yieldBai et al.
Brassica napusFlower color (anthocyanin content)Decipher anthocyanin-based petal color regulationWGCNA + transcriptomics + metabolomicsMYB75-F3H module linked to color variationWGCNA enhances trait discovery in breedingCui et al.
Impatiens uliginosaFlower colorExplore factors driving flower colorMulti-omics + physiological assaysAnthocyanins + CHS expression drive colorCandidate targets for breeding via MAS/gene editingZhao et al.

Summary of the 15 studies compiled within this Research Topic, “Utilizing Machine Learning with Phenotypic and Genotypic Data to enhance Effective Breeding in Agricultural and Horticultural Crops”.

The studies highlight how machine learning (ML) and artificial intelligence (AI) are being applied across high-throughput phenotyping, phenotypic forecasting, trait mapping, genomic selection, and multi-omics integration to support modern plant-breeding efforts.

Figure 1

Advances in ML-assisted high-throughput phenotyping

High-throughput phenotyping (HTP) systems, powered by advancements in sensor technologies and imaging analysis, have revolutionized the way phenotypic data is collected and analyzed (Juliana et al., 2019; Mir et al., 2019; Barnaby et al., 2020; Kim et al., 2024). Imaging platforms, including RGB cameras, multispectral and hyperspectral imaging, thermal cameras, and 3D laser scanners, all allow for the non-invasive assessment of plant traits at various growth stages (Volpato et al., 2021; Gimode et al., 2023; Kim et al., 2024). Machine learning algorithms, particularly convolutional neural networks (CNNs), have shown remarkable success in analyzing complex image data for phenotypic traits (Xue et al.). These methods have been used for disease detection, biomass estimation, canopy structure analysis, and stress symptom identification in various crops (Cortés and Barnaby, 2023; Li et al.; Oh et al.). For instance, CNN-based models have achieved high accuracy in detecting foliar diseases in barley (Lück et al., 2025), predicting cultivar phenotypes in lychee (Xue et al.), and assessing agronomic traits in rice, sunflower, wheat and Maize (Xie et al.). In turn, deep learning models have enabled the extraction of subtle features from images that may not be discernible through traditional image processing techniques. This capability further enhances the reliability and repeatability of phenotypic measurements, which is critical for large-scale breeding programs.

To exemplify the power of ML for image processing, the study by Oh et al. in this Research Topic used high-throughput image-based ML phenotyping to explore the genetic inheritance of seed traits (i.e., shape, size, and coat color) in common buckwheat (Fagopyrum esculentum). Researchers captured RGB images of over 100 F1 seeds from diverse germplasm, and systematically extracted quantitative features (i.e., area, length, width, aspect ratio, and grayscale) using ImageJ software, enabling precise and scalable trait measurement. Statistical analysis revealed that seed coat color and shape exhibit maternal inheritance, while size parameters often showed overdominance, especially in crosses with contrasting parental seed sizes. The study also highlighted potential barriers for intercrossing linked to flower type and parent combinations. Through this work, the authors suggest integrating ML and advanced imaging techniques, such as AI-based segmentation and multimodal data analysis, to enhance phenotyping efficiency and accuracy, supporting breeding programs aiming for superior buckwheat varieties.

Similarly, the study by Li et al. also explored high-throughput RGB with ML to non-destructively estimate individual soybean yield at the R6 stage. Authors successfully extracted morphological, color, and textural features from over 18,000 top- and side-view images of 240 diverse soybean genotypes. To accomplish so, they applied feature selection at Pearson correlation ≥0.5 and tested five ML algorithms (i.e., CatBoost, LightGBM, Random Forest, GBDT, and MLP). The team found that Gradient Boosted Decision Trees (GBDT) was the best performer, achieving an of 0.82, especially when combining multi-angle, multi-type features. Therefore, this highly scalable ML-driven phenotyping approach offers an accurate prediction of yield in soybean. Authors predict further improvements in the prediction ability of GBDT by adding sensor types, expanding image acquisition across growth stages, and refining feature selection.

Machine learning is also capable of assisting phenotyping at microscopic scales. Specifically, writing for this Research Topic, Lück et al. presented BluVision Micro, a modular, open-source phenotyping platform that leverages machine learning and automated microscopy for detailed, high-throughput analysis of microscopic plant-pathogen interactions. The platform integrates both custom-made feature models (e.g., Random Forest classifiers) and convolutional neural networks (CNNs) to detect and quantify fungal microcolonies on barley leaves with high speed, up to 30× faster than the predecessor HyphArea. The tool also provides interpretable outputs via CNN heatmaps. Authors tested the pipeline on 196 diverse barley genotypes inoculated with powdery mildew. The software recorded statistics like colony count and area that were integrated into a GWAS model, which identified several associations enriched in defense-related genes. This tool offers a scalable, accurate disease-resistance screening, ultimately unveiling novel resistance loci.

To further highlight the versatility of ML-enabled phenotyping, the study by Ferm et al. in this Research Topic developed a high-throughput, deep learning–based workflow to quantify tiller number in interspecific bentgrass hybrids—a trait traditionally difficult to measure because of the thin, filamentous morphology of tillers and their frequent occlusion. Using controlled imaging of clipped tillers placed on standardized high-contrast backgrounds, the authors compared a classical edge-based segmentation pipeline with two convolutional neural network object detectors, Faster R-CNN and YOLOv8. Although the two-stage Faster R-CNN model showed higher localization precision, the one-stage YOLOv8 detector provided markedly superior total-count accuracy, achieving an overall R² of 0.97 while maintaining robust performance even in densely crowded samples exceeding 400 tillers per image. The study also demonstrated YOLOv8’s computational efficiency, completing training in ~4 hours compared with ~26 hours for Faster R-CNN, and revealed that recall-oriented one-stage detection is better suited for fine, overlapping plant structures where minimizing missed instances is more critical than boundary tightness. Altogether, this workflow provides the first scalable, instance-level solution for bentgrass tiller quantification and establishes a transferable phenotyping framework capable of generating high-quality trait data for downstream genetic and breeding applications. Complementary to this study, a low-cost automated greenhouse imaging platform integrating ML-based processing was developed to evaluate drought tolerance in bentgrass hybrids (Kim et al., 2024), demonstrating how accessible, controlled-environment phenotyping systems can generate high-frequency, reproducible trait data.

Machine learning models for phenotypic forecasting

Machine learning encompasses algorithms for regression, classification, clustering, and prediction, making it adaptable to many plant breeding applications (Tong and Nikoloski, 2021). Among these, (i) regression algorithms speed up the detection of proxy traits and age-age correlations with higher precision (Li et al.), (ii) unsupervised classification and clustering algorithms enable forecasting trait segregation in germplasm collections with genepool and racial substructure (Schrider and Kern, 2018; Park et al., 2020; Xue et al.; Denning-James et al., 2025; López-Hernández et al., 2025a), and (iii) supervised training of target breeding labels with environmental features simplify multi-locality field trials (Azodi et al., 2019; Powadi et al.) by directing them toward specific gaps within the phenotypic-environment continuum (Costa-Neto et al., 2020; Costa-Neto and Fritsche-Neto, 2021; Resende et al., 2021).

Commonly used algorithms include decision trees, random forests, support vector machines (SVM), gradient boosting machines (GBM), and deep neural networks (DNN) (Libbrecht and Noble, 2015). More recently, ensemble learning methods, which combine multiple algorithms to improve predictive performance, have gained popularity in genomic prediction studies (Abdollahi-Arpanahi et al., 2020; Banerjee et al., 2020). Ensemble techniques such as stacking and bagging enhance model robustness by reducing variance and bias (Nascimento et al.). In breeding programs, these methods have been applied to predict complex traits like yield (Li et al.; Nascimento et al.), drought tolerance, and disease resistance (Lück et al.) with improved accuracy over traditional statistical models. At the frontier of the field, explainable AI (XAI) approaches are increasingly being integrated into ML models to enhance their transparency and interpretability (Novielli et al.). XAI tools help breeders understand which features, such as phenotypic measurements, environmental variables, or specific genetic marker, are driving model fitting. This insight is invaluable for refining feature selection criteria and making informed breeding decisions.

Concerning trait clustering approaches, the study by Xue et al. developed an advanced deep learning pipeline for non-destructive phenotyping and cultivar classification in lychee using photon-counting micro−CT imaging. Researchers built and compared seven CNN-based segmentation models, including UNet variants (UNet/UNet++/UNet3+), AttentionUNet, SegNet, TransUNet, and DeepLabV3+, to automatically segment internal fruit structures (i.e., kernel, pulp, epicarp, endocarp) with exceptional accuracy (Precision ≈ 0.90–0.99). Authors further extracted morphological features to distinguish lychee species via classifiers like SVM, Random Forest, and LDA. The LDA model, using leave-one-out cross-validation, achieved ~78–79% classification accuracy, demonstrating the efficacy of combining deep learning-based feature extraction and classical ML for both phenotyping and cultivar discrimination. This approach enables automated fruit trait analysis and cultivar identification, with further improvements expected from transfer learning and multimodal imaging.

Regarding ML-assisted trait forecasting across environments, Powadi et al. introduce a cutting-edge compositional autoencoder (CAE) to disentangle genotype- and environment-specific latent features from high-dimensional maize hyperspectral data, representing a significant advancement in ML-driven plant trait prediction. Autoencoders are a type of artificial neural networks used to learn latent features (i.e., compressed representations of data) by training the network to reconstruct its own input. This being said, authors structured the CAE to partition latent representations into genotype, macro−environment, and micro−environment components. They tested the workflow on data from 578 maize inbreeds over multiple field environments and demonstrated that CAE enhances predictive accuracy by approximately sevenfold for “Days to Pollen” and tenfold for “Yield” compared to traditional principal components analyzes (PCA) and partial-least square regression (PLSR) autoencoders. This ML pipeline enables precise trait prediction across contrasting environments and opens new avenues for integrating sensor data, multimodal inputs, and time-series measurements to accelerate breeding (Sadohara et al., 2021).

Application of AI and ML in trait mapping and candidate gene identification

Recent studies have demonstrated the power of AI and ML in trait mapping and candidate gene identification (Payseur et al., 2016). For example, ML-based multi-locus genome-wide association studies (GWAS) could successfully identify significant QTLs related to callus induction and regeneration in crops. In particular, writing for this Research Topic, Xu et al. examined 236 sorghum mini−core lines using mature seeds to quantify four key in vitro tissue culture traits (i.e., callus induction, embryogenic callus formation, browning, and plant regeneration). Authors conducted multi−locus GWAS analyzes with over 6 million SNP markers, by relying on machine learning-based GWAS methods that, unlike traditional single-locus GWAS, are designed to capture complex genetic architectures by considering multiple loci simultaneously. These models, implemented in the mrMLM (i.e., multi-locus random-SNP-effect mixed linear model) package in R, include multiple multi-locus approaches such as mrMLM, FASTmrMLM, FASTmrEMMA, pLARmEB, pKWmEB, and ISIS EM-BLASSO. They identified five genotypes (i.e., IS5667, IS24503, IS8348, IS4698, and IS5295) exhibiting high regeneration potential, mapped a total of 34 QTLs (eight for callus induction, one each for browning and embryogenesis, and 24 for differentiation), and pinpointed 47 candidate genes within these regions. Fourteen of these genes have known orthologs implicated in cellular reprogramming and embryogenesis (including WIND1, which enhances callus and shoot formation). These integrative approaches also identified candidate genes, including transcription factors and kinases involved in cellular maintenance and embryogenesis. Such findings improve the prospects for in vitro cultivation, and downstream applications like gene editing (Doudna and Charpentier, 2014).

Also at the frontier of ML-assisted GWAS mapping, Ding et al. used an ML-based approach through mixed linear models (MLM) with admixture (Q) and kinship (K) correction as proxies to control for confounding factors like population stratification and relatedness. Authors analyzed over 55,000 SNP markers across 239 global wheat varieties grown in four environments. The GWAS models identified 44 significant SNPs on 12 chromosomes associated with plant height, with some explaining up to 11% of phenotypic variance. Crucially, seven loci were linked to candidate genes involved in photosynthesis, ion transport, transcription regulation, and protein degradation pathways. This work offers insights into the polygenic nature of wheat plant height and highlights the power of machine-learning driven GWAS for marker discovery and wheat pre-breeding. These findings reinforce the potential of ML in unraveling the genetic basis of complex traits, and accelerating the development of improved environmentally resilient and sustainable varieties (Grinberg et al., 2020).

Integration of phenotypic and genotypic data for ML-enabled genomic selection

The integration of high-throughput phenotypic data with genotypic information represents a significant advancement in modern breeding strategies, especially when selecting polygenic environmentally shaped complex traits (Qiu et al., 2016). Genomic selection (GS), which predicts the breeding value of individuals based on genome-wide markers (López-Hernández et al., 2025a), has been enhanced through the incorporation of phenomic data and more robust genotyping platforms (Tong and Nikoloski, 2021).

A milestone in the rapidly evolving field of ML-assisted GS is the work by Xie et al., who introduced PNNGS, a novel parallel convolutional neural network designed to substantially improve genomic selection accuracy. Unlike serial deep-learning models, PNNGS processes high-density SNP data through multiple convolutional branches that operate in parallel using different kernel sizes (e.g., 1×1, 3×3, 5×5, 7×7). PNNGS was tested against RR-BLUP, random forests, SVR, and a serial deep neural network (DNNGP) on 24 phenotypic labels across rice, sunflower, wheat, and maize. PNNGS consistently outperformed the other models, showing an average increase of 0.031 in prediction accuracy over DNNGP. Major advantages of PNNGS include parallel deep-learning architectures, use of residual learning, custom loss functions, and cluster-aware sampling. Still, authors acknowledge that future work should expand data-balancing strategies and test PNNGS across additional species and traits.

This Research Topic also demonstrates how advanced ML models and explainable AI (XAI) techniques improve the interpretability of genomic-enabled trait prediction. Specifically, the study by Novielli et al. tested this approach using a key agronomic fruit trait referred in almond as the “shelling fraction”—i.e., the ratio of edible kernel to whole fruit. They began by using high-dimensional SNP data from 98 almond cultivars, applying feature selection and training multiple ML regression models (i.e., Random Forest, Gradient Boosting, and AdaBoost, together with gBLUP and rrBLUP) to predict the shelling fraction. The Random Forest model outperformed others with an R2of 0.511 ± 0.025. Authors applied XAI via SHAP (SHapley Additive exPlanations) to better interpret model outputs and identify the most significant SNP markers. Interestingly, the top-ranked SNP mapped to a gene implicated in seed development. The work illustrates how integrating ML prediction with XAI enhances the accuracy of genotype-phenotype predictive models while offering a mechanistic understating of the underlying nature, boosting GS utilization in crop breeding.

Not being these major accomplishments in ML-assisted GS exciting enough, Nascimento et al. further released in this Research Topic an innovative stacking ensemble learning (SEL) framework designed to enhance GS by integrating through a meta-learner (a.k.a., blender) predictions from multiple models (i.e., base learners). Specifically, authors used the coffee tree (Coffea arabica) as proof of concept to demonstrate substantial improvements in the combined prediction of key traits, including yield, fruit number, leaf miner infestation, and cercosporiosis incidence. They used 5,970–21,211 SNPs from 195 coffee individuals to train GBLUP, MARS, Quantile Random Forests, and Random Forest, and merge their predictions via meta-learners such as Ridge Regression and two-kernel GBLUP within the SEL architecture. The SEL models outperformed all individual models, showing predictive ability gains of 87% for yield, 38% for fruit number, 200% for leaf miner infestation, and 15% for disease incidence compared to standard GBLUP. This approach captures complex genetic and omnigenic architectures (Boyle et al., 2017) by integrating multiple ML methods within a layered ensemble, rather than relying on comparisons among single models.

Multi-trait genomic prediction models could also leverage phenotypic data collected through HTP platforms to improve the accuracy of genomic-estimated breeding values (Guo et al., 2020; Lenz et al., 2020). So far, ML algorithms have facilitated the integration of these diverse datasets by handling high-dimensional data and identifying complex interactions between genetic markers and phenotypic traits. This approach allows for the discovery of novel marker-trait associations capable to handle second-order interactions like pleiotropic and epistatic effects (Haleem et al., 2022). More recently, integrating spectral reflectance indices with genomic data has additionally improved the prediction of yield and protein content under varying environmental conditions (Wu et al., 2025).

Multi-omics integration and predictive breeding in the era of ML

Beyond genomics and phenomics, integrating multi-omics datasets (Cortés and Du, 2023), like transcriptomics (López-Hernández and Cortés, 2022; 2026), proteomics (Castillejo et al., 2023), metabolomics (Henao-Rojas et al., 2022; Tienda-Parrilla et al., 2022; Cui et al.), epigenomics (Mehrmohamadi et al., 2021; Du et al., 2025), and more recently metagenomics (Zhang et al., 2021; Nwachukwu and Babalola, 2022), pangenomics (Hu et al., 2025) and enviromics (Cooper and Messina, 2021; Costa-Neto et al., 2021; Resende et al., 2021; Bedoya-Canas et al., 2024; López-Hernández et al., 2025b), provides a more complete understanding of complex trait architecture (Cortés et al., 2023). Machine learning models (Ma et al., 2014b; Myburg et al., 2019), particularly DL architectures, are well suited to integrating these complex, high-dimensional, and heterogenous datasets (Barrera-Redondo et al., 2020).

At the interface of genomics and transcriptomics, the contribution to this Research Topic by Wu et al. reports a comprehensive genome−wide identification and expression profiling of the basic helix–loop–helix (bHLH) transcription factor family in soybean, elucidating their potential role in regulating grain protein synthesis. A total of 188 Gm bHLH genes were identified, and detailed analyzes of their phylogenetic relationships, gene structures, conserved motifs, chromosomal localizations, and promoter cis−elements were done. Expression profiling during seed development revealed several Gm bHLH genes that were differentially regulated and showed strong correlations with specific stages of protein accumulation. Notably, a subset of these genes showed strong associations with known pathways involved in nitrogen metabolism and storage protein regulation. The study highlights promising candidate bHLH genes for future functional validation, offering a candidate for genetic manipulation of protein content and nutritional quality in soybean seeds.

Following a similar analytical roadmap, the study of Bai et al. in this Research Topic reports a genome-wide identification and characterization of the E-class MADS-box gene family in wheat my integrating comparative genomics, expression profiling, and interaction assays to unravel their evolutionary trajectory and roles in flower development. Authors identified 15 TaE genes in bread wheat and 9 orthologs in ancestral relatives, classifying them into three phylogenetic subgroups and demonstrating conserved structural motifs and nuclear localization. The study leveraged ML-based clustering for phylogenetic tree construction and collinearity/synteny analysis to trace how polyploidization and gene duplication shaped this family. Expression data highlighted spike- and floret-specific activity, notably of TaSEP5-A, which functional role and interactions with other TaEs were respectively verified through Arabidopsis transformation and yeast two-hybrid assays. Being flower development a key fitness trait, authors envision utilizing gene editing to manipulate flowering time and structure to enhance wheat yield and adaptability across varying environmental conditions.

At the exciting interface of transcriptomic and metabolomic multi-omic data integration, the study by Cui et al. investigates the molecular basis of flower color variation in Brassica napus spanning an impressive total of eight petal colors. Authors captured 275 anthocyanin-related genes via BLAST and performed RNA−seq, revealing significant upregulation of structural genes like F3H and UGT, and transcription factors including MYB75, GL3, and TTG1. The team further applied a weighted gene co-expression network analysis (WGCNA), which is a ML-based network clustering algorithm. This technique enabled authors to pinpoint core omnigenic (Boyle et al., 2017) regulators (i.e., CHS, F3H, MYB75/12/111) associated with anthocyanin accumulation in colored petals. This combined ML-driven transcriptome-metabolome approach suggests that MYB75-mediated activation of F3H plays a central role in pigment accumulation. This work not only deciphers transcriptional regulation underlying flower color variation, but also offers candidate targets for breeding ornamental varieties. ML tools such as WGCNA can accelerate trait discovery by identifying patterns across large omic datasets. Integrating transcriptomic and metabolomic data with phenotypic information could further improve the prediction of environmentally responsive traits in breeding programs. For instance, multi-omics integration has been applied successfully in legumes, perennial crops, and cereals to identify biomarkers for stress tolerance and quality traits (Villordo-Pineda et al., 2015; Espichan et al., 2022; Loupit et al., 2022).

Zhao et al. further explored multi-omic regulation of flower color, examining how cellular, physiological, biochemical, and genetic factors underly the distinct pigmentation patterns (i.e., deep red, red, pink, and white) in Impatiens uliginosa. The team combined a comprehensive set of methodologies, including colorimetric analyzes, microscopy of epidermal cells, vacuolar pH measurement, pigment quantification (i.e., anthocyanins and flavonoids), chemical profiling of key pigments (i.e., malvidin−3−galactoside, cyanidin−3−O−glucoside, delphinidin), and gene expression assays of chalcone synthase (IuCHS) via qRT−PCR. Authors found that (i) anthocyanins dominate pigment composition with highest concentrations in red and pink flowers, (ii) cellular pH values correlate with color intensity, and (iii) IuCHS expression peaks in pink flowers, reaching its lowest level in white ones. Biochemical differences, especially in anthocyanin levels and CHS gene activity, are central to color variation, even though epidermal cell structures are similar among the different flower types. Overall, this multi-omics approach clarifies the physiological and molecular mechanisms that can guide breeding of new Impatiens varieties through MAS and gene editing targeting CHS and anthocyanin pathways.

Despite these promising research avenues in the integration and understanding of regulator networks across multiple omics layers (Du et al., 2025), challenges remain in harmonizing data collected from different omics platforms and ensuring the scalability of models across diverse populations and environments (Bedoya-Canas et al., 2024; Juma et al., 2025). Nevertheless, advances in systems biology and computational data science continue to drive rapid progress in this field (Myburg et al., 2019).

Perspectives

Despite the advances highlighted in this Research Topic, several challenges must be addressed to release the full potential of AI/ML in crop breeding. First, data quality and standardization are often overlooked (Waldvogel et al., 2020), yet are essential for ensuring consist and accurate phenotypic and genotypic datasets, which are crucial for model training, testing and validation. Second, ML models must be validated across multiple environments and populations within a breeding program (Cortés et al., 2024) to ensure robustness and generalizability (Zhang et al., 2019; Thistlethwaite et al., 2020). Third, despite major increases in computational capacity over recent decades, integrating large-scale multi-omic datasets requires even greater computational infrastructure and more efficient data-compression methods to enable scalable data storage and processing (Barnaby et al., 2013; Barnaby et al., 2015; Kim et al., 2016; Barnaby et al., 2019a; Barnaby et al., 2019b; Wang et al., 2022; Barnaby et al., 2025). In addition, the rapid evolution of the field raises ethical considerations that have not yet fully addressed, including questions of data and model ownership, and equitable access to AI-driven breeding technologies, both of which are critical for fair global agricultural development (Cortés, 2025).

Addressing these challenges will require progress not only in technical fronts (Ma et al., 2014a; Qiu et al., 2016; Sadohara et al., 2021; Sousa et al., 2021) but also in societal and policy dimensions (Louwaars, 2019; Peláez et al., 2022; Kholova et al., 2024). Specifically, more scalable and cost-effective phenotyping platforms will help match the pace at which genomics has progressed (Khaipho-Burch et al., 2023). Incorporating real-time environmental data and sensor networks will further enhance the predictive capacity of breeding models. Similarly, refining ML algorithms for greater interpretability—not just usability—will help shift the emphasis from predictive performance alone to also understanding the underlying mechanisms. Finally, promoting fair and collaborative data-sharing frameworks (e.g., Mccouch et al., 2016; Spindel and Mccouch, 2016) will help balance advances in computational capabilities with on-farm phenotyping needs and market realities in the global South (Santantonio et al., 2020), which is essential for developing environmentally resilient and sustainable crops where they are most needed (Mccouch, 2004; Mccouch et al., 2013). Ultimate progress depends on standardized data collection and reporting, robust cross-environment validation, scalable computing and data infrastructure, and equitable access to AI tools (ownership to openness). Placing greater emphasis on interpretability (XAI) rather than predictive power alone will already enhance biological utility, while collaborative, open, and interoperable datasets and pipelines (Ramírez-Gil et al., 2026) will accelerate model generalization and application—particularly in under-resourced breeding programs.

Conclusions

The integration of machine learning with high-throughput phenotypic and genotypic data is reshaping plant genetics and crop breeding (Schrider and Kern, 2018). It is increasingly recognized that global food production is both affected by and contributes to long-term environmental shifts (Scherer et al., 2020). These technologies help address the need for varieties capable of withstanding diverse environmental stresses while maintaining high yield. By enabling precise, data-driven selection across a wide range of environments, populations, and omic layers, ML and AI improve the efficiency of identifying superior genotypes. As AI and ML tools become more accessible and training datasets expand, their use in crop improvement programs is expected to grow (Varshney, 2021), supporting sustainable agriculture and global food security (Benitez-Alfonso et al., 2023). Collaboration among breeders, geneticists, eco-physiologist, data scientists, stakeholders, and policymakers will be essential to fully harness these technologies (Mccouch et al., 2020; Branch, 2026). Bridging the gap between genomics (Cañas-Gutierrez et al., 2025), phenomics, computational biology, societal development, and farmer-level technology adoption (Kholova et al., 2024) will also be essential for releasing the full potential of machine learning to enhance plant breeding amid changing environmental conditions and evolving market needs. Because agriculture is both vulnerable to and influential on environmental conditions, integrated ML/AI pipelines combining phenomics, genomics, and multi-omics—paired with practical adoption by breeders and farmers—offer a concrete pathway toward sustainable productivity and resilience. Cross-disciplinary collaboration remains essential for achieving meaningful field-level impact.

Statements

Author contributions

JB: Writing – review & editing, Writing – original draft. AC: Writing – review & editing, Writing – original draft.

Funding

The author(s) declared that financial support was received for this work and/or its publication. AC appreciates funding from Vetenskapsrådet (grants 2022-04411 and 2016-00418), the British Council (grant 527023146), and Kungliga Vetenskapsakademien (grant BS20170036) while working on this Research Topic.

Acknowledgments

JB and AC thank all authors, reviewers, and editors, who worked hard to accomplish a balanced compilation of innovative works as part of this Research Topic on “Utilizing Machine Learning with Phenotypic and Genotypic Data to enhance Effective Breeding in Agricultural and Horticultural Crops”. JYB and AJC also appreciate the continue encouragement and assistance provided by Content Specialist Roxanne Palino, as well as by the extended editorial team at Frontiers in Plant Science. Being this a Research Topic that explores and discusses the broad applications of ML technologies in science, authors also appreciate the AI assistance provided by Grammarly and Quillbot in text fine-tuning and style refinement.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

The authors JB, AC declared that they were an editorial board member of Frontiers, at the time of submission. This had no impact on the peer review process and the final decision.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

  • 1

    AbbertonM.BatleyJ.BentleyA.BryantJ.CaiH.CockramJ.et al. (2016). Global agricultural intensification during climate change: A role for genomics. Plant Biotechnol. J.14, 10951098. doi: 10.1111/pbi.12467

  • 2

    Abdollahi-ArpanahiR.GianolaD.PeñagaricanoF. (2020). Deep learning versus parametric and ensemble methods for genomic prediction of complex phenotypes. Genet. Sel. Evol.52, 12. doi: 10.1186/s12711-020-00531-z

  • 3

    AhmarS.BallestaP.AliM.Mora-PobleteF. (2021). Achievements and challenges of genomics-assisted breeding in forest trees: From marker-assisted selection to genome editing. Int. J. Mol. Sci.22, 10583. doi: 10.3390/ijms221910583

  • 4

    ArenasS.CortésA. J. (2026). Unveiling the genomic architecture of phenotypic plasticity using multiple gwas approaches under contrasting conditions of water availability: A model for barley. Int. J. Mol. Sci.27 (2), 652. doi: 10.3390/ijms27020652

  • 5

    ArenasS.CortésA. J.Jaramillo-CorreaJ. P. (2025). “ Toward a genomic-enabled selection in natural tree populations for long-term management and conservation,” in Genomics Based Approaches for Tropical Tree Improvement and Conservation (Singapore: Springer Nature), 243278.

  • 6

    AzodiC. B.BolgerE.MccarrenA.RoantreeM.De Los CamposG.ShiuS.-H. (2019). Benchmarking parametric and machine learning models for genomic prediction of complex traits. G39, 36913702. doi: 10.1534/g3.119.400498

  • 7

    BanerjeeR.MarathiB.Singh1M. (2020). Efficient genomic selection using ensemble learning and ensemble feature reduction. J. Crop Sci. Biotechnol.23, 311323. doi: 10.1007/s12892-020-00039-4

  • 8

    BarnabyJ. Y.FermD.KimY. H.WarnkeS. E. (2025). Phenological and metabolic differences between two contrasting drought tolerance groups in an interspecific bentgrass breeding population. ISTRJ15, 11821184. doi: 10.1002/its2.70096

  • 9

    BarnabyJ. Y.FleisherD. H.ReddyV. R.SicherR. S. (2015). Combined effects of CO2 enrichment, diurnal light levels and water stress on foliar metabolites of potato plants grown in naturally sunlit controlled environment chambers. Physiol. Plant153, 243252. doi: 10.1057/9780230251038_4

  • 10

    BarnabyJ. Y.FleisherD. H.SicherR. S.ReddyV. R. (2019a). Combined effects of drought and CO2 enrichment on foliar metabolites of potato (Solanum tubersum L.) cultivars. J. Plant Interact.14, 110118. doi: 10.1080/17429145.2018.1562110

  • 11

    BarnabyJ. Y.HugginsT. D.LeeH. S.McClungA. M.PinsonS. R. M.OhM. R.et al. (2020). Vis/NIR hyperspectral imaging distinguishes sub-population, production environment, and physicochemical grain properties in rice. Sci. Rep.10, 9284. doi: 10.1038/s41598-020-65999-7

  • 12

    BarnabyJ. Y.KimM. S.BauchanG.BunceJ. A.ReddyV. R.SicherR. S. (2013). Drought responses of foliar metabolites in three maize hybrids differing in water stress tolerance. PloS One8, e77145. doi: 10.1371/journal.pone.0077145

  • 13

    BarnabyJ. Y.McClungA. M.EdwardsJ. E.PinsonS. R. M. (2022). Identification of quantitative trait loci for tillering, root, and shoot biomass at the maximum tillering stage in rice. Sci. Rep.12, 13304. doi: 10.1038/s41598-022-17109-y

  • 14

    BarnabyJ. Y.RohilaJ. S.HenryC. G.SicherR. C.ReddyV. R.McClungA. M. (2019b). Physiological and metabolic responses of rice to reduced soil moisture: Relationship of water stress tolerance and grain production. Int. J. Mol. Sci.20, 1846. doi: 10.3390/ijms20081846

  • 15

    Barrera-RedondoJ.PineroD.EguiarteL. E. (2020). Genomic, transcriptomic and epigenomic tools to study the domestication of plants and animals: A field guide for beginners. Front. Genet.11, 742. doi: 10.3389/fgene.2020.00742

  • 16

    Bedoya-CanasL. E.López-HernándezF.CortésA. J. (2024). Climate change adaptation of high-elevation Polylepis forests. Forests. 15 (5), 811. doi: 10.3390/f15050811

  • 17

    Bedoya-LondoñoS.Cañas-GutiérrezG. P.CortésA. J. (2025). “ Breeding without breeding: Enabling indirect selection schemes for tropical tree improvement,” in Genomics Based Approaches for Tropical Tree Improvement and Conservation (Singapore: Springer Nature), 1942.

  • 18

    Benitez-AlfonsoY.SoanesB. K.ZimbaS.SinanajB.GermanL.SharmaV.et al. (2023). Enhancing climate change resilience in agricultural crops. Curr. Biol.33, R1246R1261. doi: 10.1016/j.cub.2023.10.028

  • 19

    BoyleE. A.LiY. I.PritchardJ. K. (2017). An expanded view of complex traits: From polygenic to omnigenic. Cell169, 11771186. doi: 10.1016/j.cell.2017.05.038

  • 20

    BranchH. A. (2026). The sleeping giant needs coffee: Overlooked areas for integrating plant ecophysiology and evolutionary biology. Am. J. Bot.113, e70164. doi: 10.1002/ajb2.70164

  • 21

    CampbellQ.Castañeda-ÁlvarezN.DomingoR.Bishop-Von WettbergE.RunckB.NandkangréH.et al. (2025). Prioritizing parents from global genebanks to breed climate-resilient crops. Nat. Clim. Change15, 673681. doi: 10.1038/s41558-025-02333-x

  • 22

    Cañas-GutierrezG. P.López-HernandezF.CortésA. J. (2025). Whole genome resequencing of 205 avocado trees unveils the genomic patterns of racial divergence in the Americas. Int. J. Mol. Sci.26. doi: 10.3390/ijms262110353

  • 23

    CastillejoM. A.PascualJ.Jorrín-NovoJ. V.BalbuenaT. S. (2023). Proteomics research in forest trees: A 2012–2022 update. Front. Plant Sci.14. doi: 10.3389/fpls.2023.1130665

  • 24

    CooperM.MessinaC. D. (2021). Can we harness “enviromics” to accelerate crop improvement by integrating breeding and agronomy? Front. Plant Sci.12, 735143. doi: 10.3389/fpls.2021.735143

  • 25

    CortésA. J. (2024). Abiotic stress tolerance boosted by genetic diversity in plants. Int. J. Mol. Sci. 25 (10), 5367. doi: 10.3390/ijms25105367

  • 26

    CortésA. J. (2025). Unlocking genebanks for climate adaptation. Nat. Clim. Change15, 590592. doi: 10.1038/s41558-025-02336-8

  • 27

    CortésA. J.BarnabyJ. Y. (2023). Harnessing Genebanks: High-Throughput Phenotyping and Genotyping of Crop Wild Relatives and Landraces (Lausanne: Frontiers Media).

  • 28

    CortésA. J.CastillejoM. Á.YocktengR. (2023). ‘Omics’ approaches for crop improvement. Agronomy13, 1401. doi: 10.3390/agronomy13051401

  • 29

    CortésA. J.DuH. (2023). Molecular genetics enhances plant breeding. Int. J. Mol. Sci.24, 9977. doi: 10.3390/ijms24129977

  • 30

    CortésA. J.López-HernándezF.BlairM. W. (2024). “ Crop modeling for future climate change adaptation,” in Digital Agriculture: A Solution for Sustainable Food and Nutritional Security. Eds. PriyadarshanP. M.JainS. M.SuprasannaP.Al-KhayriJ. M., 625639.

  • 31

    Costa-NetoG.CrossaJ.Fritsche-NetoR. (2021). Enviromic assembly increases accuracy and reduces costs of the genomic prediction for yield plasticity in maize. Front. Plant Sci.126 (1), 92106. doi: 10.3389/fpls.2021.717552

  • 32

    Costa-NetoG.Fritsche-NetoR. (2021). Enviromics: Bridging different sources of data, building one framework. Crop Breed. Appl. Biotechnol.21, e393521S393512. doi: 10.1590/1984-70332021v21sa25

  • 33

    Costa-NetoG.Fritsche-NetoR.CrossaJ. (2020). Nonlinear kernels, dominance, and envirotyping data increase the accuracy of genome-based prediction in multi-environment trials. Heredity. doi: 10.1038/s41437-020-00353-1

  • 34

    Denning-JamesK. E.ChaterC.CortésA. J.BlairM. W.PeláezD.HallA.et al. (2025). Genome-wide association mapping dissects the selective breeding of determinacy and photoperiod sensitivity in common bean (Phaseolus vulgaris L.). G3 (Bethesda)15, 6. doi: 10.1093/g3journal/jkaf090

  • 35

    DoudnaJ. A.CharpentierE. (2014). Genome editing. The new frontier of genome engineering with crispr-cas9. Science346, 1258096. doi: 10.1126/science.1258096

  • 36

    DuH.LiangZ.CortésA. J. (2025). Editorial: Evolution of crop genomes and epigenomes, volume ii. Front. Plant Sci.16, 1623554. doi: 10.3389/fpls.2025.1623554

  • 37

    EspichanF.RojasR.QuispeF.CabanacG.MartiG. (2022). Metabolomic characterization of 5 native Peruvian chili peppers (Capsicum spp.) as a tool for species discrimination. Food Chem.386, 132704. doi: 10.1016/j.foodchem.2022.132704

  • 38

    Fernandez-BacaC. P.McClungA. M.EdwardsJ.CodlingE. E.ReddyV. R.BarnabyJ. Y. (2021a). Grain inorganic arsenic content in rice managed through targeted introgressions and irrigation management. Front. Plant Sci.11, 228. doi: 10.3389/fpls.2020.612054

  • 39

    Fernandez-BacaC. P.RiversA. R.KimW. J.IwataR.McClungA. M.RobertsD. P.et al. (2021c). Changes in rhizosphere soil microbial communities across plant stages of rice genotypes. Soil Biol. Biochem.156, 108233. doi: 10.1016/j.soilbio.2021.108233

  • 40

    Fernandez-BacaC. P.RiversA. R.MaulJ. E.KimW. J.McClungA. M.RobertsD. P.et al. (2021b). Rice plant–soil microbiome interactions driven by differential root and shoot biomass. Diversity13, 125. doi: 10.3390/d13030125

  • 41

    GimodeD.ChuY.HolbrookC. C.FoncekaD.PorterW.DobrevaI.et al. (2023). High-throughput canopy and belowground phenotyping of a set of peanut cssls detects lines with increased pod weight and foliar disease tolerance. Agronomy13, 1223. doi: 10.3390/agronomy13051223

  • 42

    GrinbergN. F.OrhoborO. I.KingR. D. (2020). An evaluation of machine-learning for predicting phenotype: Studies in yeast, rice, and wheat. Mach. Learn.109, 251277. doi: 10.1007/s10994-019-05848-5

  • 43

    Guevara-EscuderoM.OsorioA. N.CortésA. J. (2021). Integrative pre-breeding for biotic resistance in forest trees. Plants10, 2022. doi: 10.3390/plants10102022

  • 44

    GuoJ.KhanJ.PradhanS.ShahiD.KhanN.AvciM.et al. (2020). Multi-trait genomic prediction of yield-related traits in us soft wheat under variable water regimes. Genes11, 1270. doi: 10.3390/genes11111270

  • 45

    GuoH.NieC.-Y.LiZ.KangJ.WangX.-L.CuiY.-N. (2023). Physiological and transcriptional analyses provide insight into maintaining ion homeostasis of sweet sorghum under salt stress. Int. J. Mol. Sci.24 (13), 11045. doi: 10.3390/ijms241311045

  • 46

    HaleemA.KleesS.SchmittA. O.GültasM. (2022). Deciphering pleiotropic signatures of regulatory snps in Zea mays L. using multi-omics data and machine learning algorithms. Int. J. Mol. Sci.23, 5121. doi: 10.3390/ijms23095121

  • 47

    Henao-RojasJ. C.OsorioE.IsazaS.Madronero-SolarteI. A.SierraK.Zapata-VahosI. C.et al. (2022). Towards bioprospection of commercial materials of Mentha spicata L. using a combined strategy of metabolomics and biological activity analyzes. Molecules22. doi: 10.3390/molecules27113559

  • 48

    HuH.ZhaoJ.ThomasW. J. W.BatleyJ.EdwardsD. (2025). The role of pangenomics in orphan crop improvement. Nat. Commun.16, 118. doi: 10.1038/s41467-024-55260-4

  • 49

    HultgrenA.CarletonT.DelgadoM.GergelD. R.GreenstoneM.HouserT.et al. (2025). Impacts of climate change on global agriculture accounting for adaptation. Nature642, 644652. doi: 10.1038/s41586-025-09085-w

  • 50

    JulianaP.Montesinos-LópezO. A.CrossaJ.MondalS.González PérezL.PolandJ.et al. (2019). Integrating genomic‐enabled prediction and high‐throughput phenotyping in breeding for climate‐resilient bread wheat. Theor. Appl. Genet.132, 177194. doi: 10.1007/s00122-018-3206-3

  • 51

    JumaI.ValenciaJ. B.CortésA. J. (2025). Predicting suitable regions for avocado (Persea americana Mill.) tree cultivation in Tanzania. Horticulturae12 (1), 24. doi: 10.3390/horticulturae12010024

  • 52

    Khaipho-BurchM.CooperM.CrossaJ.LeonN. D.HollandJ.LewisR.et al. (2023). Scale up trials to validate modified crops’ benefits. Nature621, 470473. doi: 10.1038/d41586-023-02895-w

  • 53

    KholovaJ.UrbanM. O.BavorovaM.CeccarelliS.CosmasL.DesczkaS.et al. (2024). Promoting new crop cultivars in low-income countries requires a transdisciplinary approach. Nat. Plants10, 16101613. doi: 10.1038/s41477-024-01831-8

  • 54

    KimY.BarnabyJ. Y.WarnkeS. E. (2024). Development of a low-cost automated greenhouse imaging system with machine learning-based processing for evaluating drought tolerance in bentgrass. Comput. Electron. Agric.224, 108896. doi: 10.1016/j.compag.2024.108896

  • 55

    KimJ. Y.YangJ.YangR. H.SicherR. S.ChangC.TuckerM. (2016). Transcriptome analysis of soybean leaf abscission identifies transcriptional regulators of organ polarity and cell fate. Front. Plant Sci.7, 125. doi: 10.3389/fpls.2016.00125

  • 56

    LenzP. R. N.NadeauS.MottetM. J.PerronM.IsabelN.BeaulieuJ.et al. (2020). Multi-trait genomic selection for weevil resistance, growth, and wood quality in Norway spruce. Evol. Appl.13, 7694. doi: 10.1111/eva.12823

  • 57

    LiF.GatesD. J.BucklerE. S.HuffordM. B.JanzenG. M.Rellan-AlvarezR.et al. (2025). Environmental data provide marginal benefit for predicting climate adaptation. PloS Genet.21, e1011714. doi: 10.1371/journal.pgen.1011714

  • 58

    LibbrechtM. W.NobleW. S. (2015). Machine learning applications in genetics and genomics. Nat. Rev. Genet.16, 321332. doi: 10.1038/nrg3920

  • 59

    López-HernándezF.Burbano-ErazoE.León-PachecoR. I.Cordero-CorderoC. C.Villanueva-MejíaD. F.Tofiño-RiveraA. P.et al. (2023). Multi-environment genome-wide association studies of yield traits in common bean (Phaseolus vulgaris L.) × tepary bean (P. acutifolius A. Gray) interspecific advanced lines in humid and dry Colombian Caribbean subregions. Agronomy13, 1396. doi: 10.3390/agronomy13051396

  • 60

    López-HernándezF.CortésA. J. (2022). Whole transcriptome sequencing unveils the genomic determinants of putative somaclonal variation in mint (Mentha L.). Int. J. Mol. Sci.23, 5291. doi: 10.3390/ijms23105291

  • 61

    López-HernándezF.CortésA. J. (2026). “ Transcriptomic signatures of somaclonal variation,” in Plant Transcriptomics and Epitranscriptomics (Singapore: Springer Nature), 135153.

  • 62

    López-HernándezF.RodaF.CortésA. J. (2025a). “ Machine learning and pangenomics: Revolutionizing plant breeding for a sustainable future,” in Plant Breeding 2050. Eds. PriyadarshanP. M.OrtizR. (Singapore: Springer), 509523.

  • 63

    López-HernándezF.Rosero-AlpalaM. G.RoseroA.CortésA. J. (2025b). Projected shifts in Colombian sweet potato germplasm under climate change. Horticulturae11 (9), 1080. doi: 10.3390/horticulturae11091080

  • 64

    López-HernándezF.Villanueva-MejíaD. F.Tofiño-RiveraA. P.CortésA. J. (2025c). Genomic prediction of adaptation in common bean (Phaseolus vulgaris L.) × tepary bean (P. acutifolius A. Gray) hybrids. Int. J. Mol. Sci.26, 7370. doi: 10.3390/ijms26157370

  • 65

    LoupitG.FonayetJ. V.PrigentS.ProdhommeD.SpilmontA. S.HilbertG.et al. (2022). Identifying early metabolite markers of successful graft union formation in grapevine. Hortic. Res.9, uhab070. doi: 10.1093/hr/uhab070

  • 66

    LouwaarsN. (2019). Open source seed, a revolution in breeding or yet another attack on the breeder's exemption? Front. Plant Sci.10, 1127. doi: 10.3389/fpls.2019.01127

  • 67

    LückS.BourrasS.DouchkovD. (2025). Deep phenotyping platform for microscopic plant-pathogen interactions. Front. Plant Sci.16, 1462694. doi: 10.3389/fpls.2025.1462694

  • 68

    MaC.XinM.FeldmannK. A.WangX. (2014a). Machine learning–based differential network analysis: A study of stress-responsive transcriptomes in Arabidopsis. Plant Cell26, 520537. doi: 10.1105/tpc.113.121913

  • 69

    MaC.ZhangH. H.WangX. (2014b). Machine learning for big data analytics in plants. Trends Plant Sci.19, 798808. doi: 10.1016/j.tplants.2014.08.004

  • 70

    MccouchS. (2004). Diversifying selection in plant breeding. PloS Biol.2, 15071512. doi: 10.1371/journal.pbio.0020347

  • 71

    MccouchS.BauteG. J.BradeenJ.BramelP.BrettingP. K.BucklerE.et al. (2013). Feeding the future. Nature499, 2324. doi: 10.1038/499023a

  • 72

    MccouchS.NavabiK.AbbertonM.AnglinN. L.BarbieriR. L.BaumM.et al. (2020). Mobilizing crop biodiversity. Mol. Plant13, 13411344. doi: 10.1016/j.molp.2020.08.011

  • 73

    MccouchS. R.WrightM. H.TungC. W.MaronL. G.McnallyK. L.FitzgeraldM.et al. (2016). Open access resources for genome-wide association mapping in rice. Nat. Commun.7, 10532. doi: 10.1038/ncomms10532

  • 74

    MehrmohamadiM.SepehriM. H.NazerN.NorouziM. R. (2021). A comparative overview of epigenomic profiling methods. Front. Cell Dev. Biol.9. doi: 10.3389/fcell.2021.714687

  • 75

    MirR. R.ReynoldsM.PintoF.KhanM. A.BhatM. A. (2019). High-throughput phenotyping for crop improvement in the genomics era. Plant Sci.282, 6072. doi: 10.1016/j.plantsci.2019.01.007

  • 76

    MyburgA. A.HusseyS. G.WangJ. P.StreetN. R.MizrachiE. (2019). Systems and synthetic biology of forest trees: A bioengineering paradigm for woody biomass feedstocks. Front. Plant Sci.10, 775. doi: 10.3389/fpls.2019.00775

  • 77

    NwachukwuB. C.BabalolaO. O. (2022). Metagenomics: A tool for exploring key microbiome with the potentials for improving sustainable agriculture. Front. Sustain. Food Syst.6. doi: 10.3389/fsufs.2022.886987

  • 78

    ParkD. S.WillisC. G.XiZ.KarteszJ. T.DavisC. C.WorthingtonS. (2020). Machine learning predicts large scale declines in native plant phylogenetic diversity. New Phytol. 227 (5), 15441556. doi: 10.1111/nph.16621

  • 79

    PayseurB. A.SchriderD. R.KernA. D. (2016). S/Hic: Robust identification of soft and hard sweeps using machine learning. PloS Genet.12, e1005928. doi: 10.1371/journal.pgen.1005928

  • 80

    PeláezD.AguilarP. A.MercadoM.López-HernándezF.GuzmánM.Burbano-ErazoE.et al. (2022). Genotype selection, and seed uniformity and multiplication to ensure common bean (Phaseolus vulgaris L.) var. Liborino. Agronomy12, 2285. doi: 10.3390/agronomy12102285

  • 81

    PinsonS. R. M.HeuscheleD. J.EdwardsJ. E.JacksonA. K.SharmaS.BarnabyJ. Y. (2022). Relationships among arsenic-related traits in rice revealed by genome-wide association. Front. Genet.12, 787767. doi: 10.3389/fgene.2021.787767

  • 82

    QiuZ.ChengQ.SongJ.TangY.MaC. (2026). Application of machine learning-based classification to genomic selection and performance improvement. 9771, 412421. doi: 10.1007/978-3-319-42291-6_41

  • 83

    Ramírez-GilJ. G.López-HernandezF.Conejo-RodriguezD. F.Henao-RojasJ. C.Quiroga-BenavidesK. E.CortésA. J.et al. (2026). Germversity: A free and user-friendly interface to enhance the visualization and analysis of genebank data. PloS One21, e0340826. doi: 10.1371/journal.pone.0340826

  • 84

    ResendeR. T.PiephoH. P.RosaG. J. M.Silva‐JuniorO. B.SilvaF. F. E.ResendeM. D. V. D.et al. (2021). Enviromics in breeding: Applications and perspectives on envirotypic‐assisted selection. Theor. Appl. Genet.134, 95112. doi: 10.1007/s00122-020-03684-z

  • 85

    SadoharaR.LongY.IzquierdoP.UrreaC. A.MorrisD.CichyK. (2021). Seed coat color genetics and genotype × environment effects in yellow beans via machine‐learning and genome‐wide association. Plant Genome15, e20173. doi: 10.1002/tpg2.20173

  • 86

    SantantonioN.AtandaS. A.BeyeneY.VarshneyR. K.OlsenM.JonesE.et al. (2020). Strategies for effective use of genomic information in crop breeding programs serving Africa and South Asia. Front. Plant Sci.11. doi: 10.3389/fpls.2020.00353

  • 87

    SchererL.SvenningJ. C.HuangJ.SeymourC. L.SandelB.MuellerN.et al. (2020). Global priorities of environmental issues to combat food insecurity and biodiversity loss. Sci. Total Environ.730, 139096. doi: 10.1016/j.scitotenv.2020.139096

  • 88

    SchriderD. R.KernA. D. (2018). Supervised machine learning for population genetics: A new paradigm. Trends Genet.34, 301312. doi: 10.1016/j.tig.2017.12.005

  • 89

    SmaleM.JamoraN. (2020). Valuing genebanks. Food Secur.12, 905918. doi: 10.1007/s12571-020-01034-x

  • 90

    SousaI. C. D.NascimentoM.SilvaG. N.NascimentoA. C. C.CruzC. D.SilvaF. F. E.et al. (2021). Genomic prediction of leaf rust resistance to arabica coffee using machine learning algorithms. Sci. Agric.78, 4. doi: 10.1590/1678-992x-2020-0021

  • 91

    SpindelJ. E.MccouchS. R. (2016). When more is better: How data sharing would accelerate genomic selection of crop plants. New Phytol.212, 814826. doi: 10.1111/nph.14174

  • 92

    ThistlethwaiteF. R.Gamal El-DienO.RatcliffeB.KlapsteJ.PorthI.ChenC.et al. (2020). Linkage disequilibrium vs. pedigree: Genomic selection prediction accuracy in conifer species. PloS One15, e0232201. doi: 10.1371/journal.pone.0232201

  • 93

    Tienda-ParrillaM.López-HidalgoC.Guerrero-SanchezV. M.Infantes-GonzálezÁ.Valderrama-FernándezR.CastillejoM-Á.et al. (2022). Untargeted ms-based metabolomics analysis of the responses to drought stress in Quercus ilex L. leaf seedlings and the identification of putative compounds related to tolerance. Forests13, 551. doi: 10.3390/f13040551

  • 94

    TongH.NikoloskiZ. (2021). Machine learning approaches for crop improvement: Leveraging phenotypic and genotypic big data. J. Plant Physiol.257, 153354. doi: 10.1016/j.jplph.2020.153354

  • 95

    VarshneyR. K. (2021). The plant genome special issue: Advances in genomic selection and application of machine learning in genomic prediction for crop improvement. Plant Genome14, e20178. doi: 10.1007/978-981-99-4673-0_9

  • 96

    Vega-MuñozM. A.López-HernandezF.CortésA. J.RodaF.CastanoE.MontoyaG.et al. (2025). Pangenomic and phenotypic characterization of Colombian Capsicum germplasm reveals the genetic basis of fruit quality traits. Int. J. Mol. Sci.26 (17), 8205. doi: 10.3390/ijms26178205

  • 97

    Villordo-PinedaE.Gonzalez-ChaviraM. M.Giraldo-CarbajoP.Acosta-GallegosJ. A.Caballero-PerezJ. (2015). Identification of novel drought-tolerant-associated snps in common bean (Phaseolus vulgaris). Front. Plant Sci.6, 546. doi: 10.3389/fpls.2015.00546

  • 98

    VolpatoL.PintoF.González-PérezL.ThompsonI. G.BorémA.ReynoldsM.et al. (2021). High throughput field phenotyping for plant height using uav-based rgb imagery in wheat breeding lines: Feasibility and validation. Front. Plant Sci.12. doi: 10.3389/fpls.2021.591587

  • 99

    WaldvogelA. M.SchreiberD.PfenningerM.FeldmeyerB. (2020). Climate change genomics calls for standardised data reporting. Front. Ecol. Evol.8, 242. doi: 10.3389/fevo.2020.00242

  • 100

    WangP.SchumacherA. M.ShiuS.-H. (2022). Computational prediction of plant metabolic pathways. Curr. Opin. Plant Biol.66, 102171. doi: 10.1016/j.pbi.2021.102171

  • 101

    WarnkeS. E.BarnabyJ. Y. (2023). Genetic diversity of colonial bentgrass Agrostis capillaris based on simpe sequence repeat markers and high-resolution melt analysis with haplotype scoring. Crop Sci.63, 16281633. doi: 10.1002/csc2.20943

  • 102

    WuN.JiangT.FengY.YuanM. (2025). Genome-wide identification and expression analysis of soybean Bhlh transcription factor and its molecular mechanism on grain protein synthesis. Front. Plant Sci.16, 1481565. doi: 10.3389/fpls.2025.1481565

  • 103

    ZhangL.ChenF.ZengZ.XuM.SunF.YangL.et al. (2021). Advances in metagenomics and its application in environmental microorganisms. Front. Microbiol.12. doi: 10.3389/fmicb.2021.766364

  • 104

    ZhangH.YinL.WangM.YuanX.LiuX. (2019). Factors affecting the accuracy of genomic selection for agricultural economic traits in maize, cattle, and pig populations. Front. Genet.10, 189. doi: 10.3389/fgene.2019.00189

Summary

Keywords

artificial intelligence (AI), big data, crop breeding, deep learning (DL), genotyping, germplasm, high-throughput phenotyping (HTP), pre-breeding

Citation

Barnaby JY and Cortés AJ (2026) Editorial: Utilizing machine learning with phenotypic and genotypic data to enhance effective breeding in agricultural and horticultural crops. Front. Plant Sci. 17:1869724. doi: 10.3389/fpls.2026.1869724

Received

30 April 2026

Revised

12 June 2026

Accepted

19 June 2026

Published

21 July 2026

Volume

17 - 2026

Edited and reviewed by

Yuri Shavrukov, Flinders University, Australia

Updates

Copyright

*Correspondence: Jinyoung Y. Barnaby, ; Andrés J. Cortés,

†These authors have contributed equally to this work and share first authorship

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics