Abstract
Total sugar is one of the chemical indicators related to tobacco quality. However, due to its composition of various oligosaccharides, disaccharides, and polysaccharides, total sugar lacks a clear genetic target, making conventional breeding approaches for its improvement particularly challenging. Therefore, conducting genomic selection (GS) studies on total sugar content holds significant importance for the development of high-sugar tobacco varieties. In this study, 2,604 tobacco germplasm accessions with broad genetic diversity, sourced from the National tobacco Germplasm Repository, were used as experimental materials. High-throughput sequencing technologies were employed to perform comprehensive genetic evaluation and construct genomic selection models for total sugar content. Sixteen mainstream genomic prediction models were assessed through five-fold cross-validation to develop a high-accuracy prediction framework. Among these models, the Gradient Boosting Machine (GBM) achieved the highest prediction accuracy for total sugar content (0.85), followed by the rrBLUP model (0.81). Comparative analysis of computational speed and resource consumption revealed that GBM maintained rapid computation and low resource usage even in large sample sizes, demonstrating strong stability and superior performance. Considering all factors, GBM was preliminarily identified as the optimal model for predicting total sugar content in tobacco. The application of high-accuracy genomic prediction is expected to overcome the challenges of phenotypic evaluation in breeding programs and significantly enhance the efficiency of selecting high-sugar tobacco varieties.
Introduction
Tobacco (Nicotiana tabacum L.) is an important economic crop in China and is widely cultivated across both northern and southern provinces. At present, tobacco is grown and produced in 24 provinces and autonomous regions, forming six major tobacco-producing areas: North China, Northeast China, East China, Central and South China, Southwest China, and Northwest China (Yu et al., 2026). Among these regions, Yunnan is the largest tobacco-producing area in China. Tobacco is also a major component of Yunnan’s agricultural industry and plays an important role in the province’s economic development. In 2019, the total tobacco leaf production in Yunnan reached 835,400 tons, accounting for 38.79% of the national total tobacco production of 2,153,400 tons. Despite its high economic value, with the continuous increase in consumer demand for tobacco quality, the genetic improvement of total sugar content has gradually become an important focus in tobacco breeding (Tong et al., 2021; Ullah et al., 2025; Tsaliki et al., 2023).
Total sugar is one of the chemical traits associated with tobacco quality and has significant effects on sensory characteristics such as combustibility, smoking taste, aroma, and smoke softness (; ). An appropriate total sugar content helps improve flavor balance and smoking quality. In addition, during cigarette processing, the pyrolysis products of sugar compounds can participate in the formation of aroma-related substances, thereby enhancing overall product quality. Therefore, achieving proper regulation of sugar content has become one of the key targets in tobacco breeding. However, total sugar represents the sum of various oligosaccharides, disaccharides, and polysaccharides. Its metabolism is regulated by a complex network, and the target substances involved remain unclear. In addition, the genetic architecture of total sugar is complex, and major-effect loci have not been clearly identified. As a result, conventional phenotypic selection and marker-assisted selection (MAS) face clear limitations in the genetic improvement of total sugar (). Therefore, more efficient and systematic breeding strategies are needed to advance genetic improvement for this trait.
Genomic selection (GS), also referred to as genomic prediction (GP), is an extension of marker-assisted selection to the whole-genome level (; ). It mainly uses a large number of genome-wide molecular markers together with phenotypic data from a reference population to build prediction models, estimate the breeding value contributed by markers across the genome, and then predict the breeding values of offspring individuals using marker data alone for selection (; ). Implementation of GS requires two key populations: a training population (TP) and a breeding population (BP), both genotypic and phenotypic data are collected to train a regression model that estimates the effects of genetic markers across the whole genome. In contrast, each sample in the BP only needs to be genotyped, and its breeding value for the target trait can then be predicted using the trained model (). Individuals in the BP are subsequently ranked according to their estimated breeding values, and those with superior values are selected. GS was first applied in animal breeding and was later introduced into crop breeding, including major crops such as maize (Yan et al., 2021a; ), rice (; Wang et al., 2025a) and wheat, where it has greatly improved the efficiency of genetic improvement. Taken together, GS enables accurate prediction of total sugar content using genotypic data without the need for phenotypic evaluation, and therefore provides an important approach for breeding tobacco varieties with high sugar content.
In recent years, genomic selection has gradually been introduced into tobacco breeding, although its application remains largely exploratory. Previous studies have shown that tobacco quality-related traits, including total sugar, are controlled by a complex genetic architecture. For example, Zhang et al. identified SNP markers associated with several leaf chemical traits, including total sugar, based on a GWAS analysis of diverse tobacco germplasm, providing evidence for the genetic basis of these traits (). Tong et al. further analyzed QTL distributions for multiple agronomic and chemical traits using a high-density genetic map and, for the first time, evaluated the performance of several GS models in tobacco (Tong et al., 2021). Their results supported the feasibility of GS for tobacco genetic improvement and highlighted its potential value in molecular breeding.
In this study, 2,604 tobacco germplasm accessions with broad genetic diversity were used to conduct a systematic genetic evaluation of total sugar content. The prediction accuracy of 16 genomic selection models for total sugar was assessed using five-fold cross-validation, and a genomic selection model for total sugar in tobacco was established. This study provides a theoretical and methodological basis for improving total sugar through GS.
Materials and methods
Experimental materials and field design
A total of 2,604 tobacco germplasm accessions with broad genetic diversity were selected for this study, including backbone varieties approved or registered over the past 30 years, genetic materials with important breeding value from industry genome programs and major biological breeding projects, as well as widely cultivated varieties. A field experiment was conducted in 2022 in Longshan County, Hunan Province, China. The field trial was arranged in a randomized complete block design with two replicates. Each plot covered 22.0 m2, with a row spacing of 1.1 m, a plant spacing of 0.6 m, and a row length of 10.0 m. Two rows were planted in each plot, with one plant per hill and more than 20 plants per plot in total. Seeds were sown around February 27. When plants reached technological maturity in mid-September, leaves were harvested in three rounds. Field management, including irrigation and fertilization, followed local conventional cultivation practices.
Phenotype identification
Middle leaves were harvested from each tobacco accession at the appropriate maturity stage. After harvest, the leaves were subjected to heat treatment and drying to inactivate enzymatic activity and obtain dry samples for chemical analysis, thereby minimizing the effects of postharvest physiological changes on sugar determination. The dried leaf samples were then ground and passed through a 40-mesh sieve to obtain uniform powder. Sample preparation conditions were strictly controlled to ensure comparability among samples. Total sugar content was determined at the Quality and Safety Research Center, Tobacco Research Institute, Chinese Academy of Agricultural Sciences, using the continuous flow method according to the Chinese tobacco industry standard YC/T 159–2019 for the determination of water-soluble sugars in tobacco and tobacco products. Three technical replicates were performed for each sample, and the mean value was used as the phenotypic value. Total sugar content was expressed as grams of sugar per 100 g of dry sample (g/100 g). Experimental conditions were strictly controlled throughout the analysis to ensure consistency among samples. The resulting phenotypic data were used for subsequent genetic analysis and genomic selection.
The narrow-sense heritability of total sugar content was estimated as:where represents additive genetic variance, represents residual variance.
Genotype data processing
Genotyping of the 2,604 germplasm accessions was performed using genotyping-by-sequencing (GBS) (). Genomic DNA was first extracted using the CTAB method. After digestion with NlaIII and MseI, the resulting fragments were purified with AMPure XP beads and used for library construction. High-throughput sequencing was then carried out on the Illumina NovaSeq platform. Raw sequencing data were quality-filtered using fastp () to remove low-quality reads, and the clean reads were aligned to the ZY300 reference genome using BWA (). Variants were identified based on SAMtools mpileup (). High-quality SNPs were initially screened using the following criteria: average sequencing depth >2, minor allele frequency (MAF) > 0.03, and missing rate <0.5 for both individuals and loci (; ). To improve genotype completeness, missing genotypes were imputed using Beagle (), resulting in a final set of 98,667 high-quality SNP markers. Validation with whole-genome resequencing data showed that the average SNP calling accuracy reached 0.96, indicating high reliability.
Principal component analysis
PCA was conducted using the prcomp function with scaled genotype data. The first two principal components (PC1 and PC2) were retained for visualization and downstream clustering. K-means clustering was applied to the PC1 and PC2 scores. The optimal number of clusters (k) was determined using the within-cluster sum of squares (WSS) elbow method, implemented via factoextra R package (). Based on the inflection point observed in the WSS curve, k = 3 was selected as the optimal cluster number.
GWAS analysis
GWAS for total sugar content was performed using a mixed linear model implemented in GEMMA. The model incorporated SNP effects as fixed effects and controlled for population structure and genetic relatedness using principal components and a genomic relationship matrix. The genome-wide significance threshold was determined by Bonferroni correction based on the number of filtered SNPs. Manhattan and QQ plots were generated to display association signals and assess the effectiveness of population structure correction.
Construction of the genomic selection model
Based on the filtered set of 98,667 SNP markers, genomic selection analysis was performed for total sugar content in tobacco, using total sugar content as the phenotypic trait. A total of 16 widely used genomic selection (GS) models were applied to the filtered genotypic and phenotypic data for prediction analysis. These models can be broadly classified into three categories.
The first category included linear models, namely, ridge regression (RR), least absolute shrinkage and selection operator (LASSO) (), and elastic net (EN) (Zou and Hastie, 2005). These models are particularly useful when the underlying genetic architecture is expected to be sparse, as they can identify relevant markers while excluding unimportant ones. The second category included mixed linear models. These comprised additive models such as genomic best linear unbiased prediction (GBLUP) () and ridge regression best linear unbiased prediction (rrBLUP) (), non-additive kernel-based models such as reproducing kernel Hilbert space (RKHS) () and multi-kernel reproducing kernel Hilbert space (MKRKHS) (Yu et al., 2024), and Bayesian additive models including Bayes A (BA), Bayes B (BB), Bayes C (BC) (Webb et al., 2010), Bayesian ridge regression (BRR) (), and Bayesian LASSO (BL) (). These models are based on different assumptions about the genetic architecture of the trait. Some are more suitable for traits controlled mainly by many small-effect genes, some for traits influenced by major-effect genes, and others for traits jointly controlled by both major- and small-effect genes. The third category consisted of nonlinear machine learning models, including deep neural network genomic prediction (DNNGP) (Wang K. et al., 2023), gradient boosting machine (GBM) (Yan et al., 2021b), random forest (RF) (), and support vector machine (SVM) (). A key feature of these models is their ability to capture complex non-additive effects, making them potentially effective for traits influenced by genetic interactions.
Together, these 16 models from the three categories covered the major statistical and machine learning methods commonly used in current GS studies. They were selected to account for different possible genetic architectures of total sugar content and to provide a comprehensive evaluation for establishing a genomic selection framework for this trait in tobacco (; ). To improve the reproducibility of this study, all models were evaluated under the same five-fold cross-validation framework, and prediction accuracy was measured using the Pearson correlation coefficient between predicted and observed phenotypic values (Zhou et al., 2021; Wang H. et al., 2023). For the linear and Bayesian models, default parameter settings were used. For the machine learning and deep learning models, a grid search strategy was used for hyperparameter optimization. The main software packages and hyperparameter ranges used for each model are provided in Supplementary Table S1.
Cross-validation and model evaluation
To evaluate the predictive performance of the 16 models, five-fold cross-validation (k = 5) was used. The population was randomly divided into five equal subsets (Wang et al., 2025b; Wang et al., 2026). In each round, 20% of the individuals were randomly assigned to the TP, which contained genotypic data only, while the remaining 80% were used as the TP, which included both genotypic and phenotypic data. The GS models were trained on the TP and then used to predict the phenotypic values of individuals in the TP (Zhang et al., 2019). Model performance was evaluated using the Pearson correlation coefficient (r), calculated as the correlation between predicted breeding values and observed breeding values. In addition, differences among models in computational efficiency, including running time and computational resource use under different sample sizes, were compared to identify the optimal GS model.
Computational efficiency
Computational efficiency and resource consumption are important factors affecting the application of genomic selection models in large-scale breeding. Therefore, this study used computer-based simulations to compare the computational speed and resource requirements of different GS models when handling large datasets. All simulations were performed on a laptop equipped with a 12th Gen Intel processor (12 cores, 2.50 GHz), 16 GB RAM, and a 64-bit Windows 11 operating system. All analyses were conducted in the R environment (version 4.4.1).
Result
Genetic diversity assessment of the 2,604 accessions in the training population
In this study, genotyping-by-sequencing was performed on 2,604 tobacco accessions from the National Tobacco Germplasm Bank. The materials included flue-cured tobacco (1,063 accessions), sun-cured tobacco (12 accessions), cigar tobacco (50 accessions), burley tobacco (133 accessions), oriental tobacco (68 accessions), air-cured tobacco (1,198 accessions), Maryland tobacco (11 accessions), and medicinal tobacco (69 accessions). Principal component analysis (PCA) () showed that the population could be divided into three subgroups based on genetic similarity (Figure 1). The purple group consisted mainly of introduced tobacco varieties (294 accessions), the dark green group represented Chinese-bred varieties (1,106 accessions), and the yellow group included Chinese landraces and local varieties (1,204 accessions). These results indicate that the genetic structure of this population was mainly influenced by geographic origin and germplasm type, such as local varieties, bred varieties, and introduced varieties.
FIGURE 1
Phenotypic variation and heritability of total sugar content
Phenotypic analysis of total sugar content in the 2,604 accessions showed that the values ranged from 0.07 to 39.68 mg/g, with a mean of 11.87 mg/g (Figure 2). Approximately 79% of the samples were distributed within the range of 0.07–20 mg/g. Among the major cultivated varieties in China, the total sugar contents of K326 and Yunyan 87 were 17.72 mg/g and 32.82 mg/g, respectively. These results indicate substantial phenotypic variation in total sugar content among Chinese tobacco varieties, introduced materials, and local germplasm resources. Heritability analysis showed that the narrow-sense heritability of total sugar content was h2 = 0.36 ± 0.03. The likelihood ratio test further indicated that the genetic effect was highly significant (LRT = 1,655.76, df = 1, P < 10−100) (; Yang et al., 2011; Visscher et al., 2008). This suggests that total sugar content is a trait with moderate heritability, and that a considerable proportion of the phenotypic variation can be explained by genetic factors, indicating good potential for genetic improvement.
FIGURE 2
GWAS reveals the polygenic architecture of total sugar content
To further investigate the genetic basis of total sugar content in tobacco, a genome-wide association study (GWAS) was conducted using the phenotypic data from 2,604 accessions together with the genotypic data obtained by genotyping-by-sequencing. As shown in Figure 3, using a significance threshold of −log10 (P) ≥ 5, one significant association signal was detected at 77, 398, 396 bp on chromosome 4. However, this locus explained only a small proportion of the genetic variation, whereas most of the phenotypic variation appeared to be contributed by a large number of loci with smaller effects that did not reach genome-wide significance. Overall, these results suggest that total sugar content in tobacco has a polygenic genetic architecture characterized by many small-effect loci. In other words, this trait is likely regulated by the combined effects of multiple minor loci, possibly through additive effects and gene interactions, rather than being controlled by a single major-effect gene. This finding provides further evidence for understanding the genetic basis of total sugar content.
FIGURE 3
Prediction performance of genomic selection models
In this study, the predictive performance of 16 genomic selection (GS) models, including Bayes A, Bayes B, Bayes C, BL, BRR, GBLUP, rrBLUP, LASSO, GBM, EN, RR, RKHS, SVM, MKRKHS, RF, and DNNGP, was evaluated for total sugar content. Prediction accuracy was calculated as the average Pearson correlation coefficient across ten repeated five-fold cross-validations. As shown in Figure 4, except for RKHS, most models achieved prediction accuracies of around 0.65. Among them, the GBM model showed the best performance, with a prediction accuracy of 0.85, followed by the rrBLUP model, which reached 0.81. Overall, GBM was identified as the optimal model and can be considered the preferred approach for genomic prediction of total sugar content in tobacco.
FIGURE 4
Computational efficiency of genomic selection models
Using random sampling, we generated multiple simulated populations with sample sizes of 500, 1,000, 1,500, and 2000. To ensure that the simulated genotypic data reflected the genetic structure of tobacco populations in real breeding programs, each simulated dataset was generated by random sampling with replacement from the 2,604 real tobacco accessions. All 16 models were applied to each simulated population, and five-fold cross-validation was used to evaluate running time (seconds), memory usage, and CPU utilization. As shown in Figure 5, the GBM model consistently showed high computational efficiency and low memory consumption across all sample sizes. In particular, when the sample size reached 2000, the GBM model used only 1.83 GB of memory.
FIGURE 5
Field validation of the prediction model
To validate the prediction model for total sugar content in tobacco leaves, the GBM model was used to predict total sugar content for 5,500 accessions from the Chinese Tobacco Germplasm Collection (Zan et al., 2025). Based on the prediction results, four accessions with low predicted total sugar content (TS1–TS4) and four accessions with high predicted total sugar content (TS5–TS8) were selected for field planting. The total sugar content of these eight accessions was then measured and compared with the predicted values. The Pearson correlation coefficient between the predicted and observed total sugar contents was 0.833, demonstrating the effectiveness of the GBM model for predicting total sugar content in tobacco leaves (Figure 6).
FIGURE 6
Discussion
In this study, we conducted a comprehensive phenotypic and genetic analysis of total sugar content in tobacco and found substantial phenotypic variation among Chinese tobacco varieties, introduced materials, and local germplasm resources. Total sugar showed moderate heritability, and high-sugar genetic materials were identified within the germplasm collection, indicating considerable potential for improving total sugar content in major cultivated varieties such as Yunyan 87 and K326. Although GWAS identified one significant locus with a relatively large effect, total sugar in tobacco is still likely to be mainly controlled by many small-effect loci. The limited number of significant loci may be related to the small effects of individual loci and the conservative Bonferroni threshold, which can reduce the power to detect minor-effect variants for highly polygenic traits. By evaluating 16 genomic selection models in terms of prediction accuracy, computational resource use, and running time, we established a genomic selection framework for total sugar content, with the GBM model showing the best overall performance. The good performance of GBM may be related to its ability to capture complex and nonlinear relationships between SNP markers and phenotypic variation, which is suitable for traits with a polygenic genetic architecture.
Although this study provides a useful foundation for the genetic improvement of total sugar content in tobacco, several limitations remain. First, the field experiment was conducted in only one environment, and no multi-environment phenotypic correction was applied before model construction. Therefore, the current results mainly reflect model performance under the sampled environment. Second, although repeated five-fold cross-validation was performed, the cross-validation was based on random partitioning within the same germplasm population. Closely related accessions were not explicitly separated across folds, so the high prediction accuracy of GBM should be interpreted with caution. Third, the field validation included only eight accessions selected from the two extreme predicted groups, which may have increased the observed correlation between predicted and measured values. Therefore, this validation should be regarded as preliminary evidence rather than conclusive proof of model robustness. Future studies should use larger independent validation populations, multi-environment field trials, and population-structure- or kinship-based cross-validation to further evaluate the stability and practical value of the proposed model.
Statements
Data availability statement
The original contributions presented in the study are publicly available. This data can be found in the NCBI repository with the accession number PRJNA936601.
Author contributions
YY: Conceptualization, Data curation, Resources, Software, Writing – original draft. QX: Data curation, Software, Writing – original draft. JF: Data curation, Formal Analysis, Software, Writing – original draft. LG: Data curation, Formal Analysis, Software, Writing – original draft. ZeH: Investigation, Software, Writing – original draft. HS: Investigation, Methodology, Writing – original draft. HW: Data curation, Formal Analysis, Investigation, Writing – original draft. GS: Investigation, Methodology, Project administration, Writing – original draft. YZ: Funding acquisition, Investigation, Methodology, Project administration, Writing – original draft. ZoH: Investigation, Methodology, Project administration, Writing – original draft, Writing – review and editing.
Funding
The author(s) declared that financial support was received for this work and/or its publication. This research was funded by the National Science Foundation of China (32200503), Taishan Young Scholar Program and Distinguished Overseas Young Talents Program from Shandong province (2024HWYQ-079), Key Science and Technology Project from China National Tobacco Corporation (110202101040 JY-17), Jiangsu Tobacco Industrial Co., Ltd. (H200407), Hubei Province (2025KY3CGJCYN2A011), Agricultural Science and Technology Innovation Program (ASTIP-TRIC01) from Chinese Academy Agriculture Sciences, and the key Science and Technology Program of Shandong Tobacco Corporation (202417). China National Tobacco Corporation and Jiangsu Tobacco Industrial Co., Ltd. were not involved in the study design, collection, analysis, interpretation of data, the writing of this article, or the decision to submit it for publication.
Conflict of interest
Authors YY, QX, JF, and ZH were employed by China Tobacco Jiangsu Industrial Co., Ltd. Author ZH was employed by Shandong Tobacco Corporation.
The remaining author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Supplementary material
The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fgene.2026.1886099/full#supplementary-material
References
1
Abdollahi-ArpanahiR.GianolaD.PenagaricanoF. (2020). Deep learning versus parametric and ensemble methods for genomic prediction of complex phenotypes. Genet. Sel. Evol.52, 12. 10.1186/s12711-020-00531-z
2
AlemuA.ÅstrandJ.Montesinos-LópezO. A.Isidro y SánchezJ.Fernández-GónzalezJ.TadesseW.et al (2024). Genomic selection in plant breeding: key factors shaping two decades of progress. Mol. Plant17, 552–578. 10.1016/j.molp.2024.03.007
3
BerlinetA.Thomas-AgnanC. (2011). Reproducing Kernel Hilbert Spaces in Probability and Statistics. Springer Science and Business Media.
4
BrowningB. L.BrowningS. R. (2016). Genotype imputation with millions of reference samples. Am. J. Hum. Genet.98, 116–126. 10.1016/j.ajhg.2015.11.020
5
ChaoD.WangH.WanF.YanS.FangW.YangY. (2025). MtCro: multi-task deep learning framework improves multi-trait genomic prediction of crops. Plant Methods21, 12. 10.1186/s13007-024-01321-0
6
ChenS.ZhouY.ChenY.GuJ. (2018). Fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics34, i884–i890. 10.1093/bioinformatics/bty560
7
ClarkS. A.Van Der WerfJ. (2013). in Genome-Wide Association Studies and Genomic Prediction 321-330 (Springer).
8
CollardB. C.MackillD. J. (2008). Marker-assisted selection: an approach for precision plant breeding in the twenty-first century. Philos. Trans. R. Soc. Lond B Biol. Sci.363, 557–572. 10.1098/rstb.2007.2170
9
ElshireR. J.GlaubitzJ. C.SunQ.PolandJ. A.KawamotoK.BucklerE. S.et al (2011). A robust, simple genotyping-by-sequencing (GBS) approach for high diversity species. PLoS One6, e19379. 10.1371/journal.pone.0019379
10
EndelmanJ. B. (2011). Ridge regression and other kernels for genomic selection with R package rrBLUP. Plant Genome4, 250–255. 10.3835/plantgenome2011.08.0024
11
HeK.YuT.GaoS.ChenS.LiL.ZhangX.et al (2025). Leveraging automated machine learning for environmental data-driven genetic analysis and genomic prediction in maize hybrids. Adv. Sci. (Weinh)12, e2412423. 10.1002/advs.202412423
12
HeslotN.YangH.-P.SorrellsM. E.JanninkJ.-L. (2012). Genomic selection in plant breeding: a comparison of models. Crop Sci.52, 146–160. 10.2135/cropsci2011.06.0297
13
KangH. M.ZaitlenN. A.WadeC. M.KirbyA.HeckermanD.DalyM. J.et al (2008). Efficient control of population structure in model organism association mapping. Genetics178, 1709–1723. 10.1534/genetics.107.080101
14
KassambaraA.MundtF. (2016). Factoextra: Extract and Visualize the Results of Multivariate Data Analyses. Vienna, Austria: The R Foundation for Statistical Computing.
15
KumarR.DasS. P.ChoudhuryB. U.KumarA.PrakashN. R.VermaR.et al (2024). Advances in genomic tools for plant breeding: harnessing DNA molecular markers, genomic selection, and genome editing. Biol. Res.57, 80. 10.1186/s40659-024-00562-6
16
LiH.DurbinR. (2009). Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics25, 1754–1760. 10.1093/bioinformatics/btp324
17
LiH.HandsakerB.WysokerA.FennellT.RuanJ.HomerN.et al (2009). The Sequence Alignment/Map format and SAMtools. Bioinformatics25, 2078–2079. 10.1093/bioinformatics/btp352
18
LuF.LipkaA. E.GlaubitzJ.ElshireR.CherneyJ. H.CaslerM. D.et al (2013). Switchgrass genomic diversity, ploidy, and evolution: novel insights from a network-based SNP discovery protocol. PLos Genet.9, e1003215. 10.1371/journal.pgen.1003215
19
MaX.WangH.WuS.HanB.CuiD.LiuJ.et al (2024). DeepCCR: large-scale genomics-based deep learning method for improving rice breeding. Plant Biotechnol. J.22, 2691–2693. 10.1111/pbi.14384
20
MaX.WangH.YanS.ZhouC.ZhouK.ZhangQ.et al (2025). Large-scale genomic and phenomic analyses of modern cultivars empower future rice breeding design. Mol. Plant18, 651–668. 10.1016/j.molp.2025.03.007
21
MeuwissenT. H.HayesB. J.GoddardM. E. (2001). Prediction of total genetic value using genome-wide dense marker maps. Genetics157, 1819–1829. 10.1093/genetics/157.4.1819
22
Montesinos-LópezO. A.Kismiantini and Montesinos-LópezA. (2023). Designing optimal training sets for genomic prediction using adversarial validation with probit regression. Plant Breed.142, 594–606. 10.1111/pbr.13124
23
ParkT.CasellaG. (2008). The bayesian lasso. J. Am. Stat. Assoc.103, 681–686. 10.1198/016214508000000337
24
PattersonN.PriceA. L.ReichD. (2006). Population structure and eigenanalysis. PLos Genet.2, e190. 10.1371/journal.pgen.0020190
25
PolandJ.EndelmanJ.DawsonJ.RutkoskiJ.WuS.ManesY.et al (2012). Genomic selection in wheat breeding using genotyping‐by‐sequencing. Plant Genome5, plantgenome2012.06.0006. 10.3835/plantgenome2012.06.0006
26
RanstamJ.CookJ. A. (2018). LASSO regression. J. Br. Surg.105, 1348. 10.1002/bjs.10895
27
RigattiS. J. (2017). Random forest. J. Insurance Medicine47, 31–39. 10.17849/insm-47-01-31-39.1
28
ShiQ.Abdel-AtyM.LeeJ. (2016). A Bayesian ridge regression analysis of congestion's impact on urban expressway safety. Accid. Analysis and Prev.88, 124–137. 10.1016/j.aap.2015.12.001
29
SuthaharanS. (2016). in Machine Learning Models and Algorithms for Big Data Classification: Thinking with Examples for Effective Learning (Springer), 207–235.
30
TongZ.FangD.ChenX.JiaoF.ZhangY.LiY.et al (2020). Genome-wide association study of leaf chemistry traits in tobacco. Breed. Sci.70, 253–264. 10.1270/jsbbs.19067
31
TongZ.XiuZ.MingY.FangD.ChenX.HuY.et al (2021). Quantitative trait locus mapping and genomic selection of tobacco (Nicotiana tabacum L.) based on high-density genetic map. Plant Biotechnol. Rep.15, 845–854. 10.1007/s11816-021-00713-1
32
TsalikiE.MoysiadisT.ToumpasE.KalivasA.PanorasI.GrigoriadisI. (2023). Evaluation of Greek tobacco varieties (Nicotiana tabacum L.) grown in different regions οf Greece. Agriculture13, 1394. 10.3390/agriculture13071394
33
UllahA.TongZ.KamranM.LinF.ZhuT.ShahzadM.et al (2025). Integration of QTL mapping and GWAS reveals the complicated genetic architecture of chemical composition traits in tobacco leaves. Front. Plant Sci.16, 1616591. 10.3389/fpls.2025.1616591
34
VisscherP. M.HillW. G.WrayN. R. (2008). Heritability in the genomics era--concepts and misconceptions. Nat. Rev. Genet.9, 255–266. 10.1038/nrg2322
35
WangK.AbidM. A.RasheedA.CrossaJ.HearneS.LiH. (2023a). DNNGP, a deep neural network-based method for genomic prediction using multi-omics data in plants. Mol. Plant16, 279–293. 10.1016/j.molp.2022.11.004
36
WangH.LinY. N.YanS.HongJ. P.TanJ. R.ChenY. Q.et al (2023b). NRTPredictor: identifying rice root cell state in single-cell RNA-seq via ensemble learning. Plant Methods19, 119. 10.1186/s13007-023-01092-0
37
WangH.YanS.WangW.ChenY.HongJ.HeQ.et al (2025a). Cropformer: an interpretable deep learning framework for crop genomic prediction. Plant Commun.6, 101223. 10.1016/j.xplc.2024.101223
38
WangH.FuX.LiuL.WangY.HongJ.PanB.et al (2025b). PhytoCluster: a generative deep learning model for clustering plant single-cell RNA-seq data. aBIOTECH6, 189–201. 10.1007/s42994-025-00196-6
39
WangH.YanS.MaX.SiH.LuQ.ChenY.et al (2026). PhytoCell: an ensemble learning framework for identifying cell states in plant scRNA-seq data. Crop J.10.1016/j.cj.2026.02.021
40
WebbG. I.KeoghE.MiikkulainenR. (2010). Naïve bayes. Encycl. Machine Learning15, 713–714. 10.1007/978-0-387-30164-8_576
41
YanJ.XuY.ChengQ.JiangS.WangQ.XiaoY.et al (2021a). LightGBM: accelerated genomically designed crop breeding ensemble learning. Genome Biol.22, 271. 10.1186/s13059-021-0242-y
42
YanJ.XuY.ChengQ.JiangS.WangQ.XiaoY.et al (2021b). LightGBM: accelerated genomically designed crop breeding through ensemble learning. Genome Biol.22, 271. 10.1186/s13059-021-02492-y
43
YangJ.LeeS. H.GoddardM. E.VisscherP. M. (2011). GCTA: a tool for genome-wide complex trait analysis. Am. J. Hum. Genet.88, 76–82. 10.1016/j.ajhg.2010.11.011
44
YuL.DaiY.ZhuM.GuoL.JiY.SiH.et al (2024). ShinyGS—a graphical toolkit with a serial of genetic and machine learning models for genomic selection: application, benchmarking, and recommendations. Front. Plant Sci.15, 1480902. 10.3389/fpls.2024.1480902
45
YuL.GuoL.LiuL.RenM.ChengL.LiangL.et al (2026). Construction of genomic prediction models for leaf protein content in Nicotiana tabacum. Ind. Crops Prod.243, 123090. 10.1016/j.indcrop.2026.123090
46
ZanY.ChenS.RenM.LiuG.LiuY.HanY.et al (2025). The genome and GeneBank genomics of allotetraploid Nicotiana tabacum provide insights into genome evolution and complex trait regulation. Nat. Genet.57, 986–996. 10.1038/s41588-025-02126-0
47
ZhangH.YinL.WangM.YuanX.LiuX. (2019). Factors affecting the accuracy of genomic selection for agricultural economic traits in maize, cattle, and pig populations. Front. Genet.10, 189. 10.3389/fgene.2019.00189
48
ZhouJ.BoS.WangH.ZhengL.LiangP.ZuoY. (2021). Identification of disease-related 2-Oxoglutarate/Fe (II)-dependent oxygenase based on reduced amino acid cluster strategy. Front. Cell. Dev. Biol.9707938, 10.3389/fcell.2021.707938
49
ZouH.HastieT. (2005). Regularization and variable selection via the elastic net. J. R. Stat. Soc. Ser. B Stat. Methodol.67, 301–320. 10.1111/j.1467-9868.2005.00503.x
Summary
Keywords
genome-wide association analysis, genomic selection, machine learning, tobacco, total sugar
Citation
Yang Y, Xu Q, Fu J, Guo L, Huang Z, Si H, Wang H, Shen G, Zan Y and Hu Z (2026) Construction and optimization of a genomic selection model for total sugar content in tobacco. Front. Genet. 17:1886099. doi: 10.3389/fgene.2026.1886099
Received
20 May 2026
Revised
10 June 2026
Accepted
07 July 2026
Published
23 July 2026
Volume
17 - 2026
Edited by
Mingxing Cheng, Wuhan University, China
Reviewed by
Yan Tomason, West Virginia State University, United States
Prakash Kumar, Indian Council of Agricultural Research (ICAR), India
Updates
Copyright
© 2026 Yang, Xu, Fu, Guo, Huang, Si, Wang, Shen, Zan and Hu.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Guanwang Shen, gwshen@swu.edu.cn; Yanjun Zan, zanyanjun@caas.cn; Zongyu Hu, huzy707@sina.com
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.