DATA REPORT article

Front. Plant Sci., 28 November 2022

Sec. Functional and Applied Plant Genomics

Volume 13 - 2022 | https://doi.org/10.3389/fpls.2022.1056891

TeaGVD: A comprehensive database of genomic variations for uncovering the genetic architecture of metabolic traits in tea plants

  • 1. National Center for Tea Improvement, Tea Research Institute of the Chinese Academy of Agricultural Science, Hangzhou, China

  • 2. Tea Research Institute, Lishui Academy of Agricultural and Forestry Sciences, Lishui, China

  • 3. Research Institute of Climate Change and Agriculture, National Institute of Horticultural and Herbal Science, Jeju, South Korea

  • 4. Department of Horticulture, Faculty of Agriculture, Ataturk University, Erzurum, Turkey

Introduction

Tea plant (Camellia sinensis (L.) O. Kuntze) is one of the most important nonalcoholic beverage crops. As a result of distinctive metabolites beneficial to health, tea plant is also widely used to uncover the molecular mechanisms underlying the synthesis of specific metabolites, such as catechins and caffeine (; ; Zhu et al., 2021). The rapid development of high-throughput sequencing technologies has led to an exponential increase in the volume of biological sequence data of tea plants over the past decade, providing valuable insights into the diversity and evolution of tea germplasms and the mechanism of important metabolites and agronomic traits in tea plant. Tea genome is large (~3 Gb) and complex, harboring a large number of repetitive sequences and high heterozygosity due to its self-incompatibility. Recently, the completion and availability of genome assemblies of tea plant have accelerated the investigations of evolutionary dynamics of whole-genome duplication, tandem duplication, and long terminal repeat retrotransposons that resulted in the diversification of tea germplasms (; ; ; Zhang et al., 2020a; Zhang et al., 2020c; ; Zhang et al., 2021). Meanwhile, large-scale resequencing and RNA-seq projects of tea germplasms have been performed and enabled novel insights into the diversity, evolution and domestication in tea germplasms (; ; Yu et al., 2020; Zhang et al., 2020c; Zhang et al., 2021). Genome-wide linkage study and genome-wide association study (GWAS) have revealed numerous sites and genes controlling relevant agronomical traits of tea plant, such as leaf traits (; ) and metabolites (Zhang et al., 2020b), which provide an important foundation for further decoding the molecular mechanism of traits in tea plant.

However, the lack of a standardized data processing and visualizing platform hinders the availability of such data. The construction of a user-friendly web-based platform for big data deposition, integration, accession and visualization has become a crucial issue for maximizing these valuable sequence data. Recently, several specialized web-based databases have been developed for the storage and utilization of biological sequence data in tea plant, such as TPIA (), TeaPGDB (), TeaCoN (Zhang et al., 2020b), and TeaAS (). However, these databases did not comprehensively integrate a large-scale genomic variation of various tea genetic resources and genotype-to-phenotype associations (G2Ps) for understanding the complex traits in tea plants, hindering the availability of big omics data. Here, we collected and identified more than 70 million genomic variations and 17,974 high-quality G2Ps for 464 tea metabolites. A comprehensive and user-friendly database of genomic variations for tea plants, TeaGVD (http://www.teaplant.top/teagvd), has been developed for storage, retrieval, visualization and utilization of these data, which will facilitate understanding of the genetic architecture of metabolic and agronomic traits, molecular assistant breeding, and molecular design breeding in tea plants.

Materials and methods

Data sources

Currently, the raw reads of whole-genome sequencing (WGS), GBS data and RNA-seq data from eight datasets of tea germplasms comprising 1,229 accessions were collected (Table S1). All the species and varieties in Camellia L. Sect. Thea (L.) Dyer were covered, including C. sinensis (L.) O. Kuntze var. sinensis, var. assamica (Masters) Kitamura, var. pubilimba Chang, C. taliensis (W. W. Smith) Melchior, C. tachangensis F. C. Zhang, C. crassicolumna Chang, and C. gymnogyna Chang (). Four datasets of WGS germplasms representing genetic diversity and improvement of tea plants were downloaded from NCBI with BioProject accession numbers PRJNA646044, PRJNA597714, PRJNA665594, and PRJNA716079 (; ; ; Zhang et al., 2021). GBS data were downloaded from the Genome Sequence Archive in National Genomics Data Center, China National Center for Bioinformation/Beijing Institute of Genomics, Chinese Academy of Sciences with CRA001438 (). Other datasets were RNA-seq data and downloaded from PRJNA595795 and PRJNA562973 with 217 and 136 tea accessions, respectively (Yu et al., 2020; Zhang et al., 2020c). In addition, GA and eight catechin compounds in three leaf samples of 176 tea accessions (Zhang et al., 2020c) and 437 annotated metabolites detected by UPLC-QTOF MS of 136 tea accessions (Yu et al., 2020) were integrated into the database. Because a high-quality chromosome-level genome assembly is basis for identification of genomic variations and genome-wide association analysis, the reference genome (C. sinensis var. sinensis ‘Shuchazao’), functional annotation and gene expression were downloaded from the Tea Plant Information Archive (http://tpdb.shengxin.ren/) (). Two previously published draft genomes of C. sinensis var. sinensis ‘Shuchazao’ and C. sinensis var. assamica ‘Yunkang 10’ have widely applied in genetic and functional studies in tea plants (Xia et al., 2017; ). For users’ convenience, a total of 31,780 orthologous gene sets were identified for the three tea genome assemblies by using BLASTP () based on the bidirectional best hit (BBH) method (Table S2).

Data processing

To identify the genomic variation of tea germplasms accurately, the raw reads were trimmed by Sickle (https://github.com/najoshi/sickle) with default parameters to remove low-quality sequences. In WGS germplasms, the trimmed reads were aligned to the tea pant reference genome using Burrows Wheeler Aligner (BWA) () and PCR duplicates were filtered by Sambamba () with parameters “–overflow-list-size 1000000 –hash-table-size 1000000”. After filtering low-quality alignments, SNP and InDel were identified by SAMtools () and FreeBayes (). In GBS germplasms, the trimmed reads were aligned to the tea pant reference genome using BWA () and SNP and InDel were identified by HaplotypeCaller of GATK with parameters “–minimum-mapping-quality 30 -ERC GVCF –dont-use-soft-clipped-bases” (). In RNA-seq germplasms, the trimmed RNA-seq reads were mapped to the reference genome using HISAT2 with default parameters (). PCR duplicates were removed by Picard (https://broadinstitute.github.io/picard). SNP and InDel calling was performed by HaplotypeCaller of GATK (). These SNPs and InDels were further filtered by VCFtools with parameters “–max-missing 0.5 –minQ 30 –maf 0.05” (). The identified genomic variations were annotated by SnpEff (), ANNOVAR () and VEP () based on the gene annotation file of the tea plant genome with default parameters.

To explore the genetic diversity of tea germplasms, the SNP density, nucleotide diversity (θπ), and Tajima’s D statistics of 461 WGS germplasms were calculated by VCFtools (). In addition, GWAS was performed with EMMAX () and GAPIT () with GLM, MLM, CMLM and FarmCPU model to find genetic variations or genes associated with a particular metabolic trait. The threshold of significant candidate loci (lead SNPs) was determined by GEC software (). The LD Score regression intercept and heritability were estimated by LDSC software (https://github.com/bulik/ldsc).

Implementation

The interactive web interface of TeaGVD was built based on Flask, a lightweight Python Web framework (https://palletsprojects.com/p/flask/), and it integrated all pre-processed data. The frontend pages were developed and visualized by HTML5, CSS5, jQuery, Bootstrap (https://getbootstrap.com/), ECharts (https://echarts.apache.org/), and Bokeh (https://bokeh.org/). The BLAST tool was implemented using SequenceServer (). In the PCR primer design tool, Primer3 () was used to pick PCR primers based on the reference genome with customization.

Database contents

To take advantages of omics data in tea plant, the sequencing data of 1,229 accessions of tea germplasms were collected and analyzed using a standardized pipeline. In total, more than 70 million genomic variations (SNPs and InDels) were identified from the sequencing data (Table 1). The missing rate and level of heterozygosity were 20.27% and 16.73%, respectively. Among these, 6,193,642, 30,938, and 944,449 genomic variations were present in or around gene regions (e.g., exon, intron, upstream, and downstream), accounting for 8.74%, 17.66% and 77.15% of these in WGS, GBS, and RNA-seq data, respectively. In addition, 17,974 high-quality G2Ps for 464 tea metabolites have been identified by GWAS. To facilitate the exploration of these data, we developed a comprehensive and user-friendly database of genomic variations in tea plants (TeaGVD) that was built and organized into three functional modules for various data types and applications, including Genotype, Phenotype, and Tools modules (Figure 1A). These modules provide user-friendly web interfaces to retrieve and visualize genomic variations and their related information. In the Genotype module, users can retrieve available SNP/InDel information by multiple search strategies with filter parameters. Moreover, TeaGVD can figure out the polymorphic SNPs/InDels between two or more germplasms rapidly by comparison of varieties, which is convenient to develop molecular markers. In the Phenotype module, TeaGVD shows the detailed trait values, value distribution, and GWAS results for each available metabolite. Users also can further explore candidate genes and functional markers associated with the metabolite of interest by the Candidate Region and Lead SNP Genotype submodules, respectively. To better utilize these data, the BLAST, Extract Sequence, Primer Design, and Population Genetic Analysis (SNP density, nucleotide diversity, and Tajima’s D statistics) tools were established in the Tools module.

Table 1

WGS GermplasmsGBS GermplasmsRNA-seq GermplasmsMetabolites
SNPInDelSNPInDelSNPInDelG2Ps
Chr15,222,352145,70015,10563890,7505,0141,760
Chr24,969,532138,45312,52058590,3605,0131,558
Chr34,482,574123,31310,23043684,3453,8601,427
Chr44,646,325126,95412,72155776,3674,1241,464
Chr54,673,005121,24110,39245064,8763,0901,634
Chr64,086,500117,34710,73851375,7034,0501,754
Chr74,473,987118,74610,99947375,8633,838896
Chr84,022,025100,78010,85744749,4032,284757
Chr93,869,937105,27210,17946270,1293,616966
Chr104,006,236106,2479,07339458,0592,810688
Chr112,921,98782,1278,40339459,7773,4511,209
Chr123,814,285101,3329,22038948,8652,520674
Chr133,139,35687,9708,15837857,9643,112762
Chr142,985,43285,2528,13438154,9142,9251,310
Chr152,811,58677,3176,73628045,3772,4111,115
UN8,859,214246,97314,293648162,0597,287
Total68,984,3331,885,024167,7587,4251,164,81159,40517,974

Statistics of genomic variations and genotype-to-phenotype associations for metabolites in tea plants.

Figure 1

Use cases

These data and tools will facilitate understanding of the genetic architecture of metabolic traits and molecular breeding in tea plants. We take EC-GC dimer isomer 4 under NEG mode as an example. Histogram plot of value distribution and table of detail value for each tea germplasm are shown by selecting the corresponding trait in Trait Search (Figure 1B). GWAS results present the Multiple GWAS comparison, GWAS Manhattan plot, QQ plot, LDSC analysis, and significant candidate loci (lead SNP) associated with EC-GC dimer isomer 4, which can be dynamically visualized by clicking on given SNP/InDel links to various detailed information pages of variation (Figure 1C). On the basis of the GWAS results, we specified genomic coordinate (Chr1:190796254-191238806) in Candidate Region and identified 12 genes in the genomic region. The gene distribution, functional annotation, and expression of these genes are displayed in the web interface (Figure 1D). The given gene links direct users to gene detailed information interface, which includes a visualized variation map around the gene, basic gene information, gene annotation (GO, KEGG, and Pfam), and gene expression of eight tissues (Figure 1E). Among these, CSS0005646 (also known as CsMYB111) has been reported to be associated with anthocyanin, catechin, and flavanol biosynthesis (Li et al., 2022). In addition, TEAV1S01r00071715 significantly associated with EC-GC dimer isomer 4 (P-value < 2.16e-12) was identified by GWAS. Comparisons of different genotypes in TEAV1S01r00071715 showed that the content of EC-GC dimer isomer 4 of genotypes AA and AG was significantly higher than that of genotype GG (P-value < 0.01, two-sided Wilcoxon test; Figure 1G) by lead SNP genotyping. We also found that genotype AA was only present in Yunnan, China, which was the center of origin for tea plants (Figure 1F).

Funding

This research was supported by the National Key Research and Development Program of China (2021YFD1200203), the Zhejiang Provincial Natural Science Foundation of China (LQ20C160010), the Zhejiang Science and Technology Major Program on Agricultural New Variety Breeding-Tea Plant (2021C02067) and the Fundamental Research Fund for Tea Research Institute of the Chinese Academy of Agricultural Sciences (1610212022009).

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Statements

Data availability statement

The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/Supplementary material and http://www.teaplant.top/teagvd.

Author contributions

LC, M-ZY, J-DC, and W-ZH conceived and designed the study. J-DC, SC, Q-YC, J-QM, J-QJ and C-LM. performed the data analysis and web design. J-DC, W-ZH, LC, M-ZY, D-GM and SE prepared the manuscript. All authors contributed to the article and approved the submitted version.

Conflict of interest

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fpls.2022.1056891/full#supplementary-material

References

  • 1

    AltschulS.MaddenT.SchäfferA.ZhangJ.ZhangZ.MillerW.et al. (1997). Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res.25 (17), 33893402. doi: 10.1093/nar/25.17.3389

  • 2

    AnY.ChenL.TaoL.LiuS.WeiC. (2021). QTL mapping for leaf area of tea plants (Camellia sinensis) based on a high-quality genetic map constructed by whole genome resequencing. Front. Plant Sci.12, 705285. doi: 10.3389/fpls.2021.705285

  • 3

    ChenL.YuF.TongQ. (2000). Discussions on phylogenetic classification and evolution of section thea. J. Tea Sci.20, 8994. doi: 10.13305/j.cnki.jts.2000.02.003

  • 4

    ChenJ.ZhengC.MaJ.JiangC.ErcisliS.YaoM.et al. (2020). The chromosome-scale genome reveals the evolution and diversification after the recent tetraploidization event in tea plant. Horticul. res.7, 63. doi: 10.1038/s41438-020-0288-2

  • 5

    CingolaniP.PlattsA.Wang leL.CoonM.NguyenT.WangL.et al. (2012). A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of drosophila melanogaster strain w1118; iso-2; iso-3. Fly6, 8092. doi: 10.4161/fly.19695

  • 6

    DanecekP.AutonA.AbecasisG.AlbersC. A.BanksE.DePristoM. A.et al. (2011). The variant call format and VCFtools. Bioinformatics27, 21562158. doi: 10.1093/bioinformatics/btr330

  • 7

    GarrisonE.MarthG. (2012). Haplotype-based variant detection from short-read sequencing. arXiv Prep. arXiv1207, 3907. doi: 10.48550/arXiv.1207.3907

  • 8

    JiangC.MaJ.ApostolidesZ.ChenL. (2019). Metabolomics for a millenniums-old crop: tea plant (Camellia sinensis). J. Agric. Food Chem.67, 64456457. doi: 10.1021/acs.jafc.9b01356

  • 9

    JinJ.MaJ.YaoM.MaC.ChenL. (2017). Functional natural allelic variants of flavonoid 3',5'-hydroxylase gene governing catechin traits in tea plant and its relatives. Planta245, 523538. doi: 10.1007/s00425-016-2620-5

  • 10

    KangH.SulJ.ServiceS.ZaitlenN.KongS.FreimerN.et al. (2010). Variance component model to account for sample structure in genome-wide association studies. Nat. Genet.42, 348354. doi: 10.1038/ng.548

  • 11

    KimD.PaggiJ. M.ParkC.BennettC.SalzbergS. L. (2019). Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat. Biotechnol.37, 907915. doi: 10.1038/s41587-019-0201-4

  • 12

    KoressaarT.RemmM. (2007). Enhancements and modifications of primer design program Primer3. Bioinformatics23, 12891291. doi: 10.1093/bioinformatics/btm091

  • 13

    LeiX.WangY.ZhouY.ChenY.ChenH.ZouZ.et al. (2021). TeaPGDB: Tea plant genome database. Bev. Plant Res.1, 5. doi: 10.48130/BPR-2021-0005

  • 14

    LiH.DurbinR. (2009). Fast and accurate short read alignment with burrows-wheeler transform. Bioinformatics25, 17541760. doi: 10.1093/bioinformatics/btp324

  • 15

    LiH.HandsakerB.WysokerA.FennellT.RuanJ.HomerN.et al. (2009). The sequence alignment/map format and SAMtools. Bioinformatics25, 20782079. doi: 10.1093/bioinformatics/btp352

  • 16

    LiM.YeungJ.ChernyS. S.ShamP. C. (2012). Evaluating the effective numbers of independent tests and significant p-value thresholds in commercial genotyping arrays and public imputation reference datasets. Hum. Genet.131, 747756. doi: 10.1007/s00439-011-1118-2

  • 17

    LuL.ChenH.WangX.ZhaoY.YaoX.XiongB.et al. (2021). Genome-level diversification of eight ancient tea populations in the guizhou and yunnan regions identifies candidate genes for core agronomic traits. Horticul. Res.8, 190. doi: 10.1038/s41438-021-00617-9

  • 18

    McKennaA.HannaM.BanksE.SivachenkoA.CibulskisK.KernytskyA.et al. (2010). The genome analysis toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res.20, 12971303. doi: 10.1101/gr.107524.110

  • 19

    McLarenW.GilL.HuntS. E.RiatH. S.RitchieG. R.ThormannA.et al. (2016). The ensembl variant effect predictor. Genome Biol.17, 122. doi: 10.1186/s13059-016-0974-4

  • 20

    MiX.YueY.TangM.AnY.XieH.QiaoD.et al. (2021). TeaAS: a comprehensive database for alternative splicing in tea plants (Camellia sinensis). BMC Plant Biol.21, 280. doi: 10.1186/s12870-021-03065-8

  • 21

    NiuS.KoiwaH.SongQ.QiaoD.ChenJ.ZhaoD.et al. (2020). Development of core-collections for guizhou tea genetic resources and GWAS of leaf size using SNP developed by genotyping-by-sequencing. PeerJ8, e8572. doi: 10.7717/peerj.8572

  • 22

    PriyamA.WoodcroftB. J.RaiV.MoghulI.MunagalaA.TerF.et al. (2019). Sequenceserver: a modern graphical user interface for custom BLAST databases. Mol. Biol. Evol.36, 29222924. doi: 10.1093/molbev/msz185

  • 23

    TarasovA.VilellaA. J.CuppenE.NijmanI. J.PrinsP. (2015). Sambamba: fast processing of NGS alignment formats. Bioinformatics31, 20322034. doi: 10.1093/bioinformatics/btv098

  • 24

    WangX.FengH.ChangY.MaC.WangL.HaoX.et al. (2020). Population sequencing enhances understanding of tea plant evolution. Nat. Commun.11, 4447. doi: 10.1038/s41467-020-18228-8

  • 25

    WangK.LiM.HakonarsonH. (2010). ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res.38, e164. doi: 10.1093/nar/gkq603

  • 26

    WangP.YuJ.JinS.ChenS.YueC.WangW.et al. (2021). Genetic basis of high aroma and stress tolerance in the oolong tea cultivar genome. Horticul. Res.8, 107. doi: 10.1038/s41438-021-00542-x

  • 27

    WangJ.ZhangZ. (2021). GAPIT version 3: Boosting power and accuracy for genomic association and prediction. Genom. Proteomics Bioinf.19, 629640. doi: 10.1016/j.gpb.2021.08.005

  • 28

    WeiC.YangH.WangS.ZhaoJ.LiuC.GaoL.et al. (2018). Draft genome sequence of Camellia sinensis var. sinensis provides insights into the evolution of the tea genome and tea quality. Proc. Natl. Acad. Sci. United States America115, E4151E4158. doi: 10.1073/pnas.1719622115

  • 29

    XiaE.LiF.TongW.LiP.WuQ.ZhaoH.et al. (2019). Tea plant information archive: a comprehensive genomics and bioinformatics platform for tea plant. Plant Biotechnol. J.17, 19381953. doi: 10.1111/pbi.13111

  • 30

    XiaE.TongW.HouY.AnY.ChenL.WuQ.et al. (2020). The reference genome of tea plant and resequencing of 81 diverse accessions provide insights into its genome evolution and adaptation. Mol. Plant13, 10131026. doi: 10.1016/j.molp.2020.04.010

  • 31

    XiaE.ZhangH.ShengJ.LiK.ZhangQ.KimC.et al. (2017). The tea tree genome provides insights into tea flavor and independent evolution of caffeine biosynthesis. Mol. Plant10, 866877. doi: 10.1016/j.molp.2017.04.002

  • 32

    YuX.XiaoJ.ChenS.YuY.MaJ.LinY.et al. (2020). Metabolite signatures of diverse Camellia sinensis tea populations. Nat. Commun.11, 5586. doi: 10.1038/s41467-020-19441-1

  • 33

    ZhangX.ChenS.ShiL.GongD.ZhangS.ZhaoQ.et al. (2021). Haplotype-resolved genome assembly provides insights into evolutionary history of the tea plant Camellia sinensis. Nat. Genet.53, 12501259. doi: 10.1038/s41588-021-00895-y

  • 34

    ZhangQ.LiW.LiK.NanH.ShiC.ZhangY.et al. (2020a). The chromosome-level reference genome of tea tree unveils recent bursts of non-autonomous LTR retrotransposons in driving genome size evolution. Mol. Plant13, 935938. doi: 10.1016/j.molp.2020.04.009

  • 35

    ZhangR.MaY.HuX.ChenY.HeX.WangP.et al. (2020b). TeaCoN: a database of gene co-expression network for tea plant (Camellia sinensis). BMC Genomics21, 461. doi: 10.1186/s12864-020-06839-w

  • 36

    ZhangW.ZhangY.QiuH.GuoY.WanH.ZhangX.et al. (2020c). Genome assembly of wild tea tree DASZ reveals pedigree and selection history of tea varieties. Nat. Commun.11, 3719. doi: 10.1038/s41467-020-17498-6

  • 37

    ZhuB.GuoJ.DongC.LiF.QiaoS.LinS.et al. (2021). CsAlaDC and CsTSI work coordinately to determine theanine biosynthesis in tea plants (Camellia sinensis L.) and confer high levels of theanine accumulation in a non-tea plant. Plant Biotechnol. J.19 (12), 23952397. doi: 10.1111/pbi.13722

Summary

Keywords

tea plant, genomic variation, database, genotype-to-phenotype associations, metabolite

Citation

Chen J-D, He W-Z, Chen S, Chen Q-Y, Ma J-Q, Jin J-Q, Ma C-L, Moon D-G, Ercisli S, Yao M-Z and Chen L (2022) TeaGVD: A comprehensive database of genomic variations for uncovering the genetic architecture of metabolic traits in tea plants. Front. Plant Sci. 13:1056891. doi: 10.3389/fpls.2022.1056891

Received

29 September 2022

Accepted

08 November 2022

Published

28 November 2022

Volume

13 - 2022

Edited by

Fei Shen, Beijing Academy of Agricultural and Forestry Sciences, China

Reviewed by

Tangchun Zheng, Beijing Forestry University, China; Kai Fan, Fujian Agriculture and Forestry University, China

Updates

Copyright

*Correspondence: Liang Chen, ; Ming-Zhe Yao,

†These authors have contributed equally to this work

This article was submitted to Functional and Applied Plant Genomics, a section of the journal Frontiers in Plant Science

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics