Abstract
Various methods have been proposed for genomic prediction (GP) in livestock. These methods have mainly focused on statistical considerations and did not include genome annotation information. In this study, to improve the predictive performance of carcass traits in Chinese Simmental beef cattle, we incorporated the genome annotation information into GP. Single nucleotide polymorphisms (SNPs) were annotated to five genomic classes: intergenic, gene, exon, protein coding sequences, and 3′/5′ untranslated region. Haploblocks were constructed for all markers and these five genomic classes by defining a biologically functional unit, and haplotype effects were modeled in both numerical dosage and categorical coding strategies. The first-order epistatic effects among SNPs and haplotypes were modeled using a categorical epistasis model. For all makers, the extension from the SNP-based model to a haplotype-based model improved the accuracy by 5.4–9.8% for carcass weight (CW), live weight (LW), and striploin (SI). For the five genomic classes using the haplotype-based prediction model, the incorporation of gene class information into the model improved the accuracies by an average of 1.4, 2.1, and 1.3% for CW, LW, and SI, respectively, compared with their corresponding results for all markers. Including the first-order epistatic effects into the prediction models improved the accuracies in some traits and genomic classes. Therefore, for traits with moderate-to-high heritability, incorporating genome annotation information of gene class into haplotype-based prediction models could be considered as a promising tool for GP in Chinese Simmental beef cattle, and modeling epistasis in prediction can further increase the accuracy to some degree.
Introduction
Genomic prediction (GP), which uses whole-genome markers to predict genomic breeding value, has been widely used in breeding programs of plants (; ; ; ) and domestic animals (; ; ; ), disease risk prediction for humans (; ; ), and phenotype prediction of model organisms (; ). Accompanied by the fast development of genotyping and sequencing technologies, various methods with different underlying statistical assumptions have been proposed for GP, including penalized and Bayesian regression methods (; ; ; ; ; ; ; ). These methods have been applied in cattle populations to improve the prediction accuracy of direct genomic estimated breeding values (DGVs) to some degree (; ; ; ; ; ). However, these established prediction methods have mainly focused on statistical considerations and did not consider the abundantly available biological information. Incorporating biological knowledge, like annotation information () and gene expression (), into GP using an appropriate method may bridge the gap between mathematical models and the underlying biological processes; thus, this information has the potential to improve the prediction accuracy under certain circumstances ().
Given the availability of genome annotation information, some studies have tried to integrate this information into prediction models to improve the predictive accuracies (; ; ; ; ). Single nucleotide polymorphisms (SNPs) were divided into different genomic classes based on the genome annotation information, and GP was conducted for genomic classes using two strategies. The first strategy was to assess the prediction accuracy for each genomic class, and then the genomic class that give the best prediction accuracy was further used for GP (; ; ). Another strategy was to assign different prior distributions for the different genomic classes, and then all genomic classes were used for prediction (). These approaches for incorporating annotation information into GP slightly improved the prediction accuracy in some cases. For instance, found that SNPs in the transcribed class produce better predictive performance than other classes in dairy cattle, with a slight increase in prediction accuracy of 0.03 for milk yield, fat yield, and protein yield traits on average. However, others discovered that the prediction accuracy of genomic classes was trait-dependent in the commercial chicken population, and the predictive performance of the whole-genome region remained more accurate (). Generally, these studies have not achieved significant improvements over their corresponding predictions without annotation information. Most studies simply applied standard prediction models for genomic classes based on individual SNPs, with the basic underlying assumption is that at least one marker is in linkage disequilibrium (LD) with each quantitative trait locus (QTL) under high-density markers. The marker density of genomic classes declined after the partitioning, which caused fewer bi-allelic SNPs in LD with a QTL.
An alternative is treating haplotypes that are on tuples of SNPs as predictor variables in GP to compensate for the imperfect LD between SNPs and QTLs (; ). The main benefit of using haplotypes for GP is that a haplotype is expected to have a higher LD with a QTL than an individual marker (), and has better ability to identify mutations than a single SNP (). For a trait controlled by rare QTLs, the fitting haplotype could yield a higher accuracy, regardless of the minor allele frequency (MAF) of the QTL (). When a high-density SNPs chip was annotated into different genomic classes, at least two SNPs may be included in a genome feature; thus, multi-allelic haplotype-based prediction models are expected to capture the state of a QTL better than single-SNP-based prediction models for genomic classes (; ).
In this study, we used annotation information of the cattle genome to divide Illumina BovineHD BeadChip into five genomic classes, including intergenic regions (IGR), gene, exon, protein coding sequences (CDS), and 3′/5′ untranslated regions (UTR) classes. Then, haploblocks were created () and haplotype effects were modeled using both numerical dosage and categorical coding strategies () for each genomic class. Although an additive model may explain a major part of the genetic variance in different datasets (), this model does not explicitly capture any kind of interaction that may be present in biochemical pathways that connect gene expression with the ultimate target phenotype. Therefore, statistical models that incorporate interactions between loci are considered as potentially beneficial for GP (; ; ; ). Epistasis resulting from interactions between genes at different loci was recognized as an important component in dissecting genetic pathways and understanding the evolution of complex genetic systems (; ). Overall, the objectives of this study were (1) to compare the predictive accuracies of haplotype-based prediction models with SNP-based prediction models, (2) to characterize the predictive performance when genome annotation information was incorporated into haplotype-based prediction model, and (3) to investigate the contribution of epistasis for the accuracy of GP for carcass traits in Chinese Simmental beef cattle.
Materials and Methods
Data
Our dataset includes 1346 Simmental cattle born between 2008 and 2015 from Ulgai, Xilingol League, and Inner Mongolia, China. After weaning, cattle were moved to Jinweifuren Co., Ltd. (Beijing, China) for fattening under the same feeding and management conditions. A more detailed description of the management processes was reported in previous studies (Zhu et al., 2016, 2017). All individuals were slaughtered at an average age of 20 months, and carcass and meat quality traits were measured in accordance with the guidelines proposed by the Institutional of Meat Purchase Specifications. All animals used in the study were treated following the guidelines established by the Council of China Animal Welfare. Protocols of the experiments were approved by the Science Research Department of the Institute of Animal Sciences, Chinese Academy of Agricultural Sciences (CAAS) (Beijing, China). The approval ID/permit numbers are SYXK (Beijing) 2008-007 and SYXK (Beijing) 2008-008. In our study, carcass weight (CW), live weight (LW), and striploin (SI) were analyzed, and their statistical description was summarized in Table 1.
TABLE 1
| Traits1 | The number of phenotype | Mean (SD) | Maximum | Minimum | h2(SE) |
| CW | 1346 | 270.67 ± 45.20 | 486.00 | 162.60 | 0.42 ± 0.05 |
| LW | 1342 | 504.95 ± 70.22 | 776.00 | 318.00 | 0.38 ± 0.07 |
| SI | 1342 | 8.55 ± 1.99 | 15.90 | 3.21 | 0.40 ± 0.05 |
Statistical description and heritability estimation of three traits in Chinese Simmental beef cattle.
1Carcass weight (CW), live weight (LW), and striploin (SI).
Genotyping and Quality Control
The DNA for each animal was obtained from blood using routine procedures. Samples were genotyped with Illumina BovineHD BeadChip. This array contains 777,962 SNPs with an average probe spacing of 3.43 kb and a median spacing of 2.68 kb. Before statistical analysis, the original SNP dataset was filtered using PLINK (v1.90) (; ). Individuals and autosomal SNPs that failed in any of the following criteria were removed, SNPs call rate (>0.90) (MAF > 0.01), Hardy–Weinberg Equilibrium (p > 10–6) and individual call rate (>0.90). Missing genotypes were imputed using BEAGLE (v4.1) (). Consequently, 1331 individuals and 671,204 SNPs remained. SNPs were coded as the number of copies of the minor allele, i.e., 0, 1, and 2 for the first homozygote, the heterozygote, and the second homozygote, respectively. About population structure, like principal component analysis (PCA) and linkage disequilibrium (LD) were performed in previous studies, which have shown that this population could be separated into five clusters, and the LD (r2) dropped below 0.2 at distances of 34 kb, indicating that the implementation of GS in this population requires at least 77,941 markers (; ).
Heritability Estimation
Phenotypes were adjusted for the environmental fixed effects, including sex, year, and the covariates of body weight upon entering the fattening farm, and the number of fattening days. Subsequently, the adjusted phenotypes were used for further analysis. Variance components were estimated using the following univariate animal model in ASREML (v4.1) ():
where y is the vector of the adjusted phenotypes, 1nis an n × 1 vector with entries equal to 1; μ is the overall mean; is a vector of random additive genetic effect, where G is the additive genetic relationship matrix constructed using all SNPs and is the additive genetic variance, Z is incidence matrix associating a; and is a vector of random residuals, where I is the identity matrix and is the residual variance. The heritability of each trait was estimated using .
SNP Annotation
The latest bovine genome annotation (Bos_taurus.ARS-UCD1.2) was downloaded from Ensemble1. According to genome annotation information, the bovine genome was partitioned into five genomic classes: (1) intergenic regions (IGR), (2) gene, (3) exon, (4) protein coding sequences (CDS), and (5) 3′/5′ untranslated regions (UTR) classes. Gene class contained the exon class, and exon class represented a combination of CDS and UTR classes. Thus, overlapping existed among different genomic classes. Then, the SNPs of BovineHD Beadchip were annotated into the corresponding genomic class based on their physical position.
Haplotype Derivation and Encoding
For the gene, exon, CDS, and UTR classes, a genome feature refers to a single gene, exon, CDS, and UTR, respectively; for the IGR class, a genome feature refers to an interval between two adjacent genes. A group of SNPs that were annotated in a certain genome feature of the five genomic classes was called an SNP set. The phased consecutive SNPs were used for haploblock construction via the approach described by for each SNP set. The number of SNPs contained in each haploblock depends on the predefined number of types for haplotype allele configurations; here, we used 10 as the maximum number of types (). For SNP sets containing only one SNP, the 0-, 1-, or 2-encoded genotypes were retained for further analysis. Subsequently, haploblocks with at least two haplotype alleles were generated for each SNP set of different genomic classes.
Haplotype effects were then modeled using both numerical dosage (; ; ) and categorical () coding strategies. In the numerical dosage model, pseudo-markers were generated for haploblocks by counting the number of copies of the respective allele carried by a certain individual, where the intra-locus additive effects were assumed. The additivity assumption was not necessary in the categorical coding, where the pseudo-markers of haploblocks were coded according to the haplotype allele configurations (genotypes), and each haplotype allele had its own independent effect. Table 2 shows the coding of a haplotype formed by two consecutive SNPs. Thus, for the five genomic classes, the pseudo-marker matrixes with entries 0, 1, and 2 were reconstructed in both numerical dosage and categorical models (CMs). For all markers, haploblocks were constructed for each chromosome separately using the same approach described above, and the process started from the first marker and followed by their physical order, whereas the genome annotation information was not used to define a biologically functional unit.
TABLE 2
| Haplotype allele 1 | Haplotype allele 2 | Categorical coding of haploblock1 | Numerical coding of haploblock | |||
| AB | Ab | aB | ab | |||
| AB | AB | AB|AB | 2 | 0 | 0 | 0 |
| AB | Ab | AB|Ab | 1 | 1 | 0 | 0 |
| AB | aB | AB|aB | 1 | 0 | 1 | 0 |
| AB | ab | AB|ab | 1 | 0 | 0 | 1 |
| Ab | Ab | Ab|Ab | 0 | 2 | 0 | 0 |
| Ab | aB | Ab|aB | 0 | 1 | 1 | 0 |
| Ab | ab | Ab|ab | 0 | 1 | 0 | 1 |
| aB | aB | aB|aB | 0 | 0 | 2 | 0 |
| aB | ab | aB|ab | 0 | 0 | 1 | 1 |
| ab | ab | ab|ab | 0 | 0 | 0 | 2 |
Numerical and categorical coding of a haploblock formed by two consecutive single nucleotide polymorphisms (SNPs).
1separates the strands of DNA. Considering this haploblock (let {A, a} and {B, b} denote alleles harbored by the two SNPs, respectively), four possible types of gametes—AB, Ab, aB, and ab—could be generated and 10 types of genotypes are possibly formed in a large population (imprinting is not considered).
Prediction Models
The prediction model used in this study was basically the same as in Eq. (1), except for the different genomic relatedness matrices G, which were constructed based on respective prediction approaches (Table 3). In our study, the predictive accuracies of using all markers were considered as a benchmark.
TABLE 3
| Models | Description | Relatedness matrices | Use1 |
| GBLUP | Genomic best linear unbiased prediction | All markers | |
| GHBLUP | Haplotype based GBLUP | All markers | |
| GHBLUP|GA | Haplotype based GBLUP given genome annotation | Genomic classes | |
| CM | Categorical marker effect model | All markers | |
| CE | Categorical epistasis model | E = 0.5 × mS#(mS + 1n×n)/m2 | All markers |
| CHM | Haplotype based CM | All markers | |
| CHE | Haplotype based CE | All markers | |
| CHM|GA | CHM given genome annotation | Genomic classes | |
| CHE|GA | CHE given genome annotation | Genomic classes |
Genomic relatedness matrices for different genomic prediction models for all markers or haplotypes.
1Refers to the whole genome-wide SNP; genomic classes refer to IGR, gene, exon, CDS, and UTR class.
In numerical dosage models, GBLUP () was performed for all markers, and the genomic relatedness matrix was calculated as , where M denotes the (0, 1, and 2) encoded genotype matrix, pi is the MAF of marker i, m is the number of markers, and P is a matrix with columns equal to 2pi. The haplotype-based genomic best linear unbiased prediction (GHBLUP) was performed for all markers. The haplotype-based genomic relatedness matrix in GHBLUP was constructed as the dot product of the haplotype allele matrix (MH) and expressed as , where MH is the pseudo-markers matrix with entries 0, 1, and 2 representing the number of copies of each haplotype allele in a haploblock, and QH is the total number of haploblocks of whole genome.
For the five genomic classes, haplotype-based genomic best linear unbiased prediction given genome annotation (GHBLUP|GA) was implemented. Similarly, the haplotype-based genomic relatedness matrix in GHBLUP|GA was constructed as , where is the haplotype allele matrix with pseudo-markers encoded with (0, 1, and 2), and is the total number of haploblocks in the corresponding genomic class.
In CMs, the SNP-based CM () was applied for all markers, and the genomic relatedness matrix in CM is expressed as S with entries , in which φjik was scored 1 if individual j and i shared the same genotype on marker k; otherwise, φjikwas scored 0, and m was the number of markers. The haplotype-based CM (CHM) was applied for all markers as well, in which the number of haploblocks that were in the same state between pairs of individuals were counted. The genomic relatedness matrix in CHM is expressed as SHwith entries , where φjiq was scored 1 if individual i and j share the same haplotype allele configuration on haploblock q; otherwise, φjiq was scored 0; QH was the total number of haploblocks, which is the same with that in GH. Therefore, the entries of SH represented the proportion of haploblocks with an identical state between pairs of individuals. For the five genomic classes, the haplotype-based CM assigned the genome annotation CHM|GA was applied. Similarly, the genomic relatedness matrix was built by counting the number of haploblocks that were in an identical state between pairs of individuals () and expressed as with entries , whereφjiq is the same as in CHM, but is the total number of haploblocks in certain genomic class, which is the same with that in .
To capture the first-order epistasis among SNPs, the CM model can be extended to categorical epistasis (CE) model (). In the CE model, the genotype combinations of each pair of loci were treated as categorical variables, and the relatedness of two individuals was measured by counting the number of pairs of markers in the same state. The genomic relatedness matrix in the CE model was be deduced from S via the formulaE = 0.5 × mS#(mS + 1n×n)/m2, where # denotes the Hadamard product. The first-order epistasis between pairs of haploblocks was modeled by extending CHM to the haplotype-based categorical epistasis model (CHE) (), where the genotype combinations of each pair of haploblocks were treated as a new categorical variable, and the genomic relatedness matrix was calculated as . The corresponding epistatic model that included the first-order epistasis among haploblocks was developed for the five genomic classes and was denoted as CHE|GA (), where the genomic relatedness matrix was constructed as .
Assessment of Prediction Accuracy
The accuracy of GP was assessed using fivefold cross-validation (CV), which assigns animals randomly into five separate subsets with near-equal size. Each subset was used as the validation set only once, with phenotype masked, and the remaining four subsets were treated as a training set. In order to reduce random sampling effects, the CV layout described above was replicated twenty times, where a new randomization was implemented for each replicate so that the each of the subset contains different individuals. DGVs were calculated for each validation subset based on the genomic relatedness matrix. For each replicate, the prediction accuracies were assessed by the correlation between the DGVs and the pre-adjusted phenotypes in the validation set divided by square root of heritability. In addition, in order to assess the extent of bias on GP, linear regression coefficients [b (y, DGV)] of the pre-adjusted phenotypes (y) on the DGVs was calculated for individuals in the validation set. Unbiased models are expected to do not significantly different from 1, whereas values greater than 1 indicate a biased deflation prediction of DGVs and values smaller than 1 indicate a biased inflation prediction of DGVs.
Results
SNP Annotation and Heritability Estimation
We annotated 671,204 filtered SNPs into five genomic classes based on their physical positions. The annotation results and descriptive statistics of each genomic class are displayed in Table 4. Overall, 67.03 and 32.97% of the total SNPs were annotated into the IGR and gene classes, respectively. Only 1.46, 1.05, and 0.39% of the total SNPs were annotated into the exon, CDS, and UTR class, respectively. The average MAF among these five genomic classes was in the range of 0.25 to 0.26. The number of haploblocks of gene, exon, CDS, and UTR classes were 87,407, 45,748, 9287, 6799, and 2409, respectively. We counted the number of genome features that were annotated by SNPs for each genomic class (Table 4). For instance, 16,286 genes were annotated by SNPs in the gene class, representing 66.30% of the total genes in the bovine genome. Based on the GREML method, the heritability estimates of CW, LW, and SI, were 0.42, 0.38, and 0.40 respectively.
TABLE 4
| Genomic class | # of SNPs1 | MAF | Mean MAF (SD) | # of haploblocks | # of represented genome feature2 |
| IGR class | 449,918 (67.03%) | 0.009–0.5 | 0.26 (0.15) | 87,407 | |
| Gene class | 221,286 (32.97%) | 0.009–0.5 | 0.26 (0.15) | 45,748 | 16,286 (66.30%) |
| Exon class | 9814 (1.46%) | 0.010–0.5 | 0.25 (0.15) | 9287 | 9287 (4.08%) |
| CDS class | 7024 (1.05%) | 0.010–0.5 | 0.25 (0.14) | 6799 | 6799 (3.17%) |
| UTR class | 2614 (0.39%) | 0.010–0.5 | 0.25 (0.15) | 2409 | 2409 (7.26%) |
| All markers | 671,204 | 0.009–0.5 | 0.26 (0.15) | 115,005 |
Mapping results and statistical descriptions of each genomic classes.
1The number of SNPs annotated in five genomic classes, and their percentage of the whole genome-wide markers is indicated in parentheses. 2The number of genomic features represented by SNPs in the corresponding genomic class, and their percentage of the total genome features of the reference genome in parentheses. The bovine reference genome contains 24,559 genes, 227,610 exons, 214,584 CDS, and 33,137 UTR. # means “the number.”
Prediction Accuracy of Haplotype-Based Prediction Model
We first compared the prediction accuracies of all markers between haplotype-based prediction models (GHBLUP and CHM) and the SNP-based prediction models (GBLUP and CM). The results showed that the predictive performances of GHBLUP and CHM were more accurate than GBLUP and CM in CW, LW, and SI (Figure 1). In the numerical dosage models, the accuracy of GHBLUP was 5.4, 9.8, and 7.1% higher than GBLUP in CW, LW, and SI, respectively (Table 5). In the CMs, CHM improved the accuracies by 7.8, 9.5, and 9.4% in CW, LW, and SI, respectively, compared with the CM results. Generally, the numerical dosage models performed better than CMs for most traits. For all markers, GBLUP slightly outperformed CM with 3.0, 0.7, and 1.2% higher accuracy in CW, LW, and SI, respectively (Table 5). The predictive performance of GHBLUP was 1% more accurate than CHM only in LW.
FIGURE 1
TABLE 5
| Trait1 | Numerical dosage model | Categorical model | Categorical epistasis model | ||||
| CW | All maker | GBLUP | 0.336 (0.05) | CM | 0.316 (0.06) | CE | 0.327 (0.06) |
| All maker | GHBLUP | 0.390 (0.06) | CHM | 0.394 (0.06) | CHE | 0.408 (0.06) | |
| IGR class | GHBLUP|GA | 0.397 (0.06) | CHM|GA | 0.387 (0.06 | CHE|GA | 0.381 (0.06) | |
| Gene class | GHBLUP|GA | 0.403 (0.05) | CHM|GA | 0.410 (0.06) | CHE|GA | 0.403 (0.06) | |
| Exon class | GHBLUP|GA | 0.246 (0.06) | CHM|GA | 0.215 (0.05) | CHE|GA | 0.217 (0.05) | |
| CDS class | GHBLUP|GA | 0.225 (0.06) | CHM|GA | 0.197 (0.05) | CHE|GA | 0.200 (0.05) | |
| UTR class | GHBLUP|GA | 0.232 (0.06) | CHM|GA | 0.188 (0.05) | CHE|GA | 0.192 (0.05) | |
| LW | All maker | GBLUP | 0.384 (0.05) | CM | 0.377 (0.06) | CE | 0.381 (0.06) |
| All maker | GHBLUP | 0.482 (0.06) | CHM | 0.472 (0.06) | CHE | 0.479 (0.05) | |
| IGR class | GHBLUP|GA | 0.423 (0.06) | CHM|GA | 0.425 (0.06) | CHE|GA | 0.428 (0.06) | |
| Gene class | GHBLUP|GA | 0.502 (0.07) | CHM|GA | 0.494 (0.07) | CHE|GA | 0.508 (0.07) | |
| Exon class | GHBLUP|GA | 0.217 (0.06) | CHM|GA | 0.237 (0.06) | CHE|GA | 0.237 (0.06) | |
| CDS class | GHBLUP|GA | 0.176 (0.06) | CHM|GA | 0.206 (0.06) | CHE|GA | 0.203 (0.06) | |
| UTR class | GHBLUP|GA | 0.193 (0.06) | CHM|GA | 0.178 (0.05) | CHE|GA | 0.179 (0.05) | |
| SI | All maker | GBLUP | 0.414 (0.07) | CM | 0.402 (0.07)) | CE | 0.402 (0.07) |
| All maker | GHBLUP | 0.485 (0.06) | CHM | 0.496 (0.06) | CHE | 0.479 (0.06) | |
| IGR class | GHBLUP|GA | 0.487 (0.06) | CHM|GA | 0.500 (0.06) | CHE|GA | 0.485 (0.06) | |
| Gene class | GHBLUP|GA | 0.506 (0.06) | CHM|GA | 0.501 (0.06) | CHE|GA | 0.500 (0.06) | |
| Exon class | GHBLUP|GA | 0.430 (0.06) | CHM|GA | 0.388 (0.06) | CHE|GA | 0.394 (0.06) | |
| CDS class | GHBLUP|GA | 0.394 (0.06) | CHM|GA | 0.350 (0.05) | CHE|GA | 0.358 (0.05) | |
| UTR class | GHBLUP|GA | 0.357 (0.06) | CHM|GA | 0.310 (0.06) | CHE|GA | 0.315 (0.06) | |
The prediction accuracies (SD) of different genomic classes in three traits of Chinese Simmental beef cattle.
1Carcass weight (CW), live weight (LW), and striploin (SI); prediction accuracies are averaged over the fivefold cross-validation (CV) and then over the 20 replicates.
Prediction Accuracy of Haplotype-Based Prediction Model Given Genome Annotation
Under the haplotyped-based model, we further compared the prediction accuracies for the genomic classes with all markers to characterize the benefits of usage genome annotation information in GP. We found that the accuracy of using gene annotation to define haploblocks was consistently higher than that of all markers across all traits (Figure 1). In GHBLUP|GA, the prediction accuracy of gene class was 0.403, 0.502, and 0.506 for CW, LW, and SI, which were 1.3, 2.0, and 2.1% higher than using GHBLUP, respectively. In the CM, CHM|GA outperformed CHM in gene class, with accuracy improvements of 1.6, 2.2, and 0.5% in CW, LW, and SI, respectively. For IGR, exon, CDS, and UTR genomic classes, the accuracies using the two haplotype-based prediction models were not improved. In GHBLUP|GA, gene class had 0.6–7.9, 7.6–28.5, 11.2–32.6, and 14.9–30.9% higher accuracies than IGR, exon, CDS, and UTR classes for the three traits, respectively. Analogously, in CHM|GA, the accuracies of the three traits using gene class were 0.1–6.9, 11.3–25.7, 15.1–28.8, and 19.1–31.6% higher than that of IGR, exon, CDS, and UTR classes, respectively (Table 5). Comparing the prediction accuracy of numerical dosage with the CM, we found that GHBLUP|GA maintained more accurate predictive performance than CHM|GA in most genomic classes (Table 5).
Prediction Accuracy of Epistasis Model
Considering the prediction model including epistatic effects may increase the accuracy and reduce the bias of DGVs. The results showed that incorporation of first-order epistatic effects into prediction model can slightly improve the prediction accuracies for most traits and genomic classes (Figure 1). When including the epistatic effects amongst SNPs into the CE model for all markers, prediction accuracy increased by 1.1 and 0.4% in CW and LW, respectively (Table 5). Similarly, the extension of CHM to CHE for all markers improved the prediction accuracies by 1.4 and 0.7% in CW and LW, respectively. For the five genomic classes, compared with CHM|GA, CHE|GA also had higher prediction accuracies in the IGR class of LW (0.3%), gene class of LW (1.4%), exon class of CW (0.2%) and SI (0.6%), CDS class of CW (0.3%) and SI (0.8%), and UTR class of CW (0.4%) and SI (0.5%).
Regression Coefficient
Table 6 displayed the slope of the regression of the adjusted phenotype on DGVs. For numerical dosage models, the regression coefficients of all marker, IGR, and gene classes were not significantly different from 1 in all traits, indicating the predictions were not significantly biased. For CMs, the regression coefficients of gene, exon, CDS, and UTR classes were significantly different from 1 in CW and LW. However, the regression coefficients for the predictions using the CMs that included the first-order epistasis were significantly different from 1 in all markers and genomic classes, suggesting that these models increased the biasedness of GPs. Generally, among five genomic classes, the regression coefficients of IGR and gene classes were similar to those of all markers, and they contribute to less bias prediction than exon, CDS, and UTR classes. When compared haplotype-based prediction models without including epistasis to the corresponding SNP-based prediction models, we found that the formers’ regression coefficients were closer to one, with less biasedness prediction.
TABLE 6
| Trait1 | Numerical dosage model | Categorical model | Categorical epistasis model | ||||
| CW | All maker | GBLUP | 1.102 (0.08) | CM | 1.097 (0.05) | CE | 1.087 (0.05) |
| All maker | GHBLUP | 1.062 (0.06) | CHM | 1.079 (0.06) | CHE | 1.388 (0.08) | |
| IGR class | GHBLUP|GA | 1.064 (0.06) | CHM|GA | 1.080 (0.07) | CHE|GA | 1.318 (0.07) | |
| Gene class | GHBLUP|GA | 1.071 (0.06) | CHM|GA | 1.090 (0.06) | CHE|GA | 1.300 (0.07) | |
| Exon class | GHBLUP|GA | 1.131 (0.16) | CHM|GA | 1.143 (0.18) | CHE|GA | 1.135 (0.18) | |
| CDS class | GHBLUP|GA | 1.173 (0.18) | CHM|GA | 1.169 (0.23) | CHE|GA | 1.156 (0.21) | |
| UTR class | GHBLUP|GA | 1.165 (0.16) | CHM|GA | 1.232 (0.16) | CHE|GA | 1.218 (0.16) | |
| LW | All maker | GBLUP | 0.984 (0.10) | CM | 1.062 (0.09) | CE | 1.094 (0.09) |
| All maker | GHBLUP | 1.009 (0.07) | CHM | 1.023 (0.08) | CHE | 1.546 (0.10) | |
| IGR class | GHBLUP|GA | 1.051 (0.07) | CHM|GA | 1.073 (0.08) | CHE|GA | 1.311 (0.08) | |
| Gene class | GHBLUP|GA | 1.051 (0.04) | CHM|GA | 1.088 (0.04) | CHE|GA | 1.629 (0.04) | |
| Exon class | GHBLUP|GA | 1.187 (0.30) | CHM|GA | 1.159 (0.22) | CHE|GA | 1.165 (0.22) | |
| CDS class | GHBLUP|GA | 1.386 (0.31) | CHM|GA | 1.285 (0.25) | CHE|GA | 1.294 (0.25) | |
| UTR class | GHBLUP|GA | 1.197 (0.29) | CHM|GA | 1.282 (0.33) | CHE|GACHE|GA | 1.278 (0.32) | |
| SI | All maker | GBLUP | 1.079 (0.03) | CM | 1.076 (0.05) | CE | 1.083 (0.05) |
| All maker | GHBLUP | 1.038 (0.05) | CHM | 1.046 (0.04) | CHE | 1.414 (0.07) | |
| IGR class | GHBLUP|GA | 1.038 (0.05) | CHM|GA | 1.049 (0.04) | CHE|GA | 1.338 (0.07) | |
| Gene class | GHBLUP|GA | 1.050 (0.05) | CHM|GA | 1.050 (0.05) | CHE|GA | 1.643 (0.06) | |
| Exon class | GHBLUP|GA | 1.055 (0.03) | CHM|GA | 1.048 (0.05) | CHE|GA | 1.052 (0.05) | |
| CDS class | GHBLUP|GA | 1.058 (0.05) | CHM|GA | 1.046 (0.07) | CHE|GA | 1.049 (0.08) | |
| UTR class | GHBLUP|GA | 1.064 (0.07) | CHM|GA | 1.081 (0.10) | CHE|GA | 1.080 (0.10) | |
Regression coefficients (SD) of pre-adjusted phenotypes on DGVs for three traits of Chinese Simmental beef cattle.
1Carcass weight (CW), live weight (LW), and striploin (SI); for each trait (row), the values in bold face indicate the coefficient are significantly different from 1 (p < 0.05); regression coefficients are averaged over the fivefold cross-validation (CV) and then over the 20 replicates.
Discussion
Advances in high-throughput genotyping technology and the availability of genome annotation information have contributed to the improvement of the predictive performance of complex quantitative traits in livestock species (; ; ; ). To bridge the gap between mathematical models and underlying biological processes, we combined bovine genome annotation information with haplotype-based prediction models to improve the predictive accuracies in Chinese Simmental beef cattle. In this study, whole genome-wide SNPs of BovineHD Beadchip were annotated to five genomic classes. The predictive performance of five genomic classes and all markers was assessed using both numerical and CMs, and the contribution of first-order epistatic effects among SNPs and haploblocks were modeled using categorical coding strategy.
Predictive Performance of Haplotype-Based Prediction Model
Haplotypes have been used widely in human genetics research (; ; ); in animal breeding studies, haplotypes have been used for the GP of breeding values with the use of high density SNP chips (; ; ; ). In this study, haplotype-based prediction models (GHBLUP and CHM) were applied to the whole genome-wide markers, and the result of this scenario was treated as a benchmark. We found that the predictive performance of haplotype-based prediction models was superior to corresponding SNP-based prediction models in the three traits (Figure 1), with higher accuracy and less bias. This was consistent with previously reported results in simulated datasets (; ), dairy cattle (; ; ) and beef cattle (). This may be attributable to haplotypes better capturing LDs with causative mutation or QTLs than single SNPs.
In livestock, SNPs are commonly bi-allelic. When mutations occur, the allele frequencies may remain (almost) unaltered. However, mutations in different loci tend to cause major changes in the haplotype frequencies (). Thus, when haplotypes were analyzed, a QTL that was not in complete LD with any individual bi-allelic SNP marker may be in complete LD with a multi-marker haplotype. To use a haplotype as an indicator variable in GP, previous studies defined haploblocks by setting windows with a fixed number of SNPs to be placed together as a haploblock (; ; ), or by considering only the first locus out of 10 consecutive loci in genomic evaluation (; ). Although their prediction accuracies were improved in GP, the number of SNPs used to outline haploblocks was arbitrarily defined.
To efficiently use the genome properties to define haploblocks and reduce the number of variables for the GP models, several researchers used only haplotypes with a high frequency in the population () or based on LD threshold to define haploblocks (). For instance, used an average LD threshold (≥0.45) to construct haploblocks and found that prediction accuracies increased for the three traits compared with the commonly-used individual SNP. Similarly, we used the cattle genome annotation information to define a biologically functional unit and constructed a haploblock for each unit. This strategy may reflect underlying biological processes and avoid haploblocks being arbitrarily defined. Our study contributes to the improvement of prediction accuracy using a haplotype-based model, since the functional unit contains the combined effects of tightly linked cis-acting causal variants (; ), and the number of haplotypes having effects was significantly larger than that for SNP models (). indicated that the increase in accuracy bringing by haplotype-based prediction models may be explained by this model capitalizing on local epistatic effects among markers.
Predictive Performance Among Five Genomic Classes
In our study, we applied | GA approaches based on the concept of defining biologically functional units as predictor variables. The results showed that the accuracies and biasedness of prediction for gene and IGR classes were consistently better than those for the exon, CDS, and UTR classes, regardless of which | GA prediction models were used. Firstly, this finding may be attributed to the number of SNPs annotated in its corresponding genomic class, which decreased from the IGR to UTR classes. As previously suggested, the number of markers plays an important role in affecting the GP performance (; ). With decreasing number of markers, the physical distance increased between the markers and QTLs and reduced the LD between markers and QTLs, which would lead to poor predictive power (; ; ). found that when the causative mutation loci had a lower MAF, a decrease in marker density would result in an incomplete linkage between the SNP and causative mutation loci; thus, these markers only explained a limited genetic variance.
In our study, 67.03 and 32.97% of the total SNPs were located within the IGR class and gene class, respectively, whereas only 0.39% of total SNPs was annotated in the UTR class, which had the lowest predictive accuracy. Secondly, the average number of SNPs in a haploblock may affect the prediction accuracy of genomic classes as well. It is clear that if each haploblock consisted of only one marker, the haplotype-based prediction models were exactly identical to the corresponding SNP-based prediction models (). In the IGR and gene classes, 87,407 and 45,748 haploblocks were constructed (Table 4), respectively, and 96.82 and 94.21% of the total haploblocks consisted of more than one SNP, which resulted in 5.15 and 4.84 SNPs per haploblock on average, respectively. However, only 9287, 6799, and 2409 haploblocks were constructed in the exon, CDS, and UTR classes. The average number of SNPs per haploblock was 1.06, 1.03, and 1.08, respectively, which indicated haplotype-based prediction models for these genomic classes were similar to SNP-based prediction models. Finally, the number of biological functional units that was used to construct the statistical framework in the | GA approaches could also be a key factor in affecting the predictive accuracies, since the biological functional units may reflect the underlying biological process.
According to the bovine genome annotation information, the bovine reference genome contained 24,559 genes, 227,610 exons, 214,584 CDS, and 33,137 UTR. In this study, gene class represented 66.3% (16,286 out of 24,559 genes) of the total genes of the reference genome, whereas 4.08% (9287 out of 227,610 exons), 3.17% (6799 out of 214,584 CDS), and 7.26% (2409 out of 33,137 UTR) of the total exons, CDS, and UTR of reference genome were respectively represented by exon, CDS, and UTR classes. Consequently, the high proportion of biological-functional-unit-like genes may contribute to stronger predictive power. Taken together, these factors may explain the outstanding predictive performance displayed in gene class compared with the other classes.
Benefits of Using Genome Annotation Information in GP
When the genome annotation information was incorporated into the haplotype-based prediction models, we also observed a slight or moderate improvement in prediction accuracies for the three traits. This can be explained by the traits having different genetic architectures (). The number of QTLs and the distribution of their effects may influence the prediction accuracies of genomic classes. For three traits, the gene class improved the prediction accuracy in comparison with the result of all markers using the haplotype-based prediction model, which was consistent with reported results in mouse and drosophila populations (). This may reflect that genetic signals of the gene class are well tagged in these traits, despite more haploblocks being constructed in the scenario of all markers. The method of defining a biological unit through haplotypes might have increased the linkage of markers and QTLs, which not only allowed the effects of QTL to be better captured but also reduced the density of unrelated markers. Studies have reported that gene class has the most potential to be enriched for trait-associated variants and was more likely to explain a large proportion of the total additive variance (; ; ). However, and found that the gene class did not lead to an improvement in predictive ability, and the whole genome-wide SNP-based prediction model remained the most efficient method for GP in chicken. These studies only annotated SNPs to the corresponding genomic class and applied the routine GP process for genomic classes. In this case, the genome annotation information cannot be comprehensively used in the SNP-based model because the biologically functional units were not defined as predictor variables in the model.
The usage of genome annotation information of the IGR class also led to a slight improvement in prediction accuracy in CW and SI. Studies have suggested that the IGR class, such as non-coding conserved regions, miRNA, and regulatory regions, might harbor important genetic variants associated with complex traits in crops (; ) and humans (; ). For instance, a study suggested that more than 75% of identified SNPs are embedded in regulatory genome segments in common human diseases (). Therefore, the IGR class may contribute to a large phenotypic variation. Overall, combining the genome annotation information of the gene class with the haplotype-based prediction models can improve the prediction accuracies, and this can be considered as a promising tool of GP for economically important traits in Chinese Simmental beef cattle.
Effects of Numerical and Categorical Model on Prediction Accuracy
When comparing the predictive performance of the numerical model with the CM, we found that GBLUP slightly outperformed the SNP-based CM in three traits. compared the predictive performance of CM with GBLUP, and found only slight differences in predictive ability between CM and GBLUP among 13 traits in mouse. The CM does not use the assumption of constant allele substitution effects like GBLUP; instead, it models the independent effect of each genotype at a locus, which enables the modeling of dominance (). The advantages of CM depend on the population structure and the influence of the dominance effects on a particular trait. One reason to use CM instead of GBLUP might be the population having prevalent heterosis, since heterosis creates a deviation from the linear dosage model. When most loci are mainly present in only two of the three possible SNP genotypes, the CM cannot substantially outperform GBLUP (). found that GHBLUP outperformed CHM in eight traits, and CHM outperformed GHBLUP in three traits. Analogously, in our study, GHBLUP|GA displayed better predictive performance than CHM|GA in most of the genomic classes among three traits. However, a similar pattern was not observed by , who found that CHM|GA performed better than GHBLUP|GAin the gene class among most traits.
Contribution of First-Order Epistasis to Prediction Accuracy
Epistasis has long been recognized as a biologically influential component contributing to the genetic architecture of quantitative traits (). Several genomic selection approaches have been developed to model both additive and epistatic effects (; ; ; ). To minimize the inherently high computational costs of those methods, EGBLUP () and kernel Hilbert space regression accommodating epistasis within the GP models were proposed (). Generally, the influence of epistasis on GP ranges from positive to negative. In some studies, prediction accuracies increased (; ; ; ), whereas in others, modeling epistasis adversely affected prediction accuracies (). For instance, extended GBLUP to EGBLUP to estimate both additive and additive by additive epistatic genetic effects. They found that the epistatic variance accounted for 9.5% of the total phenotypic variance, and the predictive reliabilities of genomic predicted breeding values increased by 0.3%, which was consistent with the results reported by . These discrepancies can be explained by the complexities of the studied traits, which are controlled by many loci exhibiting small effects entailing a low QTL detection power.
In this study, the first-order epistatic effects were captured by the categorical epistasis model, which can eliminate the undesired coding-dependent properties of EGBLUP (; ). Although EGBLUP has been applied in other studies (), suggested that both EGBLUP and the Gaussian kernel in an RKHS approach respond differently to a change in marker coding: a translation of the coding impacts the predictive ability of EGBLUP, but not that of the Gaussian kernel. The difference of coding strategy in the CM with the traditional encoding (0, 1, 2) in EGBLUP meant that the additivity assumption was not necessary in the categorical coding and the encoding of SNPs or haploblocks corresponded to the allele configurations, which enables the modeling of dominance (). In CMs, for all markers, the first-order epistasis of pairs of SNPs were modeled by the CE model, and we found an increase in predictive accuracies from step CM to the CE model in all traits except SI. also found that CE was slightly better than CM in the simulated and mouse datasets.CHE modeling of the first-order epistasis between pairs of haploblocks also increased the predictive accuracies of all makers of CW and LW. Similarly, found an improvement in predictive ability from CM to CE, and from CHM to CHE. For genomic classes, we observed a slight increase in accuracy in the gene class of LW and the CDS class of SI from CHM|GA to CHE|GA. These findings suggest that the first-order epistatic effects captured by markers was likely to contribute to some of the phenotypic variations of the traits observed in this study.
Conclusion
In our study, genome annotation information was incorporated into the haplotype-based prediction model for GP of three carcass traits in Chinese Simmental beef cattle. To enable comparison, the SNP-based and haplotype-based prediction methods were applied for all markers, and their results were treated as a benchmark. We found that when the haplotype was treated as a predictor variable, the prediction accuracy improved in most traits. After combining the genome annotation information of the gene class with the haplotype-based prediction model, a further increase in accuracy was observed in most traits compared with the results of all markers obtained by haplotype-based prediction models without genome annotation. The first-order epistatic effects among SNPs and haplotypes slightly improved the prediction accuracy of all markers in LW and CW. In conclusion, incorporating genome annotation information of gene classes into GP models through haplotype-based models could be considered as a promising tool for the GP of carcass traits in Chinese Simmental beef cattle.
Statements
Data availability statement
Genotype data have been submitted to Dryad: doi: 10.5061/dryad.4qc06. Bovine genome annotation (Bos_taurus.ARS-UCD1.2) was downloaded from Ensemble (http://asia.ensembl.org/index.html).
Ethics statement
The animal study was reviewed and approved by Science Research Department of the Institute of Animal Sciences, Chinese Academy of Agricultural Sciences (CAAS) (Beijing, China).
Author contributions
LX simulated and analyzed the data and wrote the manuscript. ZW, LX, and YL collected the data. NG, YC, XG, HG, LYX, LZ, BZ, and JL discussed and improved the manuscript. All authors read and approved the final manuscript.
Funding
This work was supported by National Natural Science Foundation of China (31802049, 31372294, and 31201782), Chinese Academy of Agricultural Sciences of Technology Innovation Project (CAAS-XTCX2016010, CAAS-ZDXT2018006, and ASTIP-IAS03), Program of National Beef Cattle and Yak Industrial Technology System (CARS-37), Cattle Breeding Innovative Research Team of Chinese Academy of Agricultural Sciences (cxgc-ias-03, Y2016PT17, and 2019-YWF-YTS-11), Beijing Natural Science Foundation (6154032).
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Footnotes
References
1
Abdollahi-ArpanahiR.MorotaG.ValenteB. D.KranisA.RosaG. J.GianolaD. (2016). Differential contribution of genomic regions to marked genetic variation and prediction of quantitative traits in broiler chickens.Genet. Sel. Evol.48:10. 10.1186/s12711-12016-10187-z
2
AbrahamG.HavulinnaA. S.BhalalaO. G.ByarsS. G.De LiveraA. M.YetukuriL.et al (2016). Genomic prediction of coronary heart disease.Eur. Heart J.373267–3278. 10.1093/eurheartj/ehw450
3
AkeyJ. M.AbrahamG.Tye-DinJ. A.BhalalaO. G.KowalczykA.ZobelJ.et al (2014). Accurate and robust genomic prediction of celiac disease using statistical learning.PLoS Genet.10:e1004137. 10.1001371/journal.pgen.1004137
4
BennewitzJ.SolbergT.MeuwissenT. (2009). Genomic breeding value estimation using nonparametric additive regression models.Genet. Sel. Evol.41:20. 10.1186/1297-9686-1141-1120
5
BoichardD.GuillaumeF.BaurA.CroiseauP.RossignolM.-N.BoscherM. Y.et al (2012). Genomic selection in French dairy cattle.Anim. Prod. Sci.52115–120. 10.1186/s12711-019-0495-1
6
BolormaaS.PryceJ. E.KemperK.SavinK.HayesB. J.BarendseW.et al (2013). Accuracy of prediction of genomic breeding values for residual feed intake and carcass and meat quality traits in Bos taurus, Bos indicus, and composite beef cattle1.J. Anim. Sci.913088–3104. 10.2527/jas.2012-5827
7
BrowningB. L.BrowningS. R. (2016). Genotype imputation with millions of reference samples.Am. J. Hum. Genet.98116–126. 10.1016/j.ajhg.2015.11.020
8
CaiX.HuangA.XuS. (2011). Fast empirical Bayesian LASSO for multiple quantitative trait locus mapping.BMC Bioinformatics12:211. 10.1186/1471-2105-12-211
9
CalusM. P.MeuwissenT. H.de RoosA. P.VeerkampR. F. (2008). Accuracy of genomic selection using different methods to define haplotypes.Genetics178553–561. 10.1534/genetics.107.080838
10
ChangC. C.ChowC. C.TellierL. C. A. M.VattikutiS.PurcellS. M.LeeJ. J. (2015). Second-generation PLINK: rising to the challenge of larger and richer datasets.Gigascience4:7. 10.1186/s13742-015-0047-8
11
ChapmanJ. M.CooperJ. D.ToddJ. A.ClaytonD. G. (2003). Detecting disease associations due to linkage disequilibrium using haplotype tags: a class of tests and the determinants of statistical power.Hum. Heredity5618–31. 10.1159/000073729
12
CurtisD. (2007). Comparison of artificial neural network analysis with other multimarker methods for detecting genetic association.BMC Genet.8:49. 10.1186/1471-2156-8-49
13
CurtisD.NorthB.ShamP. (2001). Use of an artificial neural network to detect association between a disease and multiple marker genotypes.Ann. Hum. Genet.6595–107. 10.1046/j.1469-1809.2001.6510095.x
14
CuyabanoB. C. D.SuG.LundM. S. (2014). Genomic prediction of genetic merit using LD-based haplotypes in the Nordic Holstein population.BMC Genomics15:1171. 10.1186/1471-2164-1115-1171
15
CuyabanoB. C. D.SuG.LundM. S. (2015). Selection of haplotype variables from a high-density marker map for genomic prediction.Genet. Sel. Evol.47:61. 10.1186/s12711-015-0143-3
16
DaY. (2015). Multi-allelic haplotype model based on genetic partition for genomic prediction and variance component estimation using SNP markers.BMC Genet.16:144. 10.1186/s12863-015-0301-1
17
DaetwylerH. D.Pong-WongR.VillanuevaB.WoolliamsJ. A. (2010). The impact of genetic architecture on genome-wide evaluation methods.Genetics1851021–1031. 10.1534/genetics.110.116855
18
de los CamposG.HickeyJ. M.Pong-WongR.DaetwylerH. D.CalusM. P. L. (2013). Whole-genome regression and prediction methods applied to plant and animal breeding.Genetics193327–345. 10.1534/genetics.112.143313
19
DoD. N.JanssL. L.JensenJ.KadarmideenH. N. (2015). SNP annotation-based whole genomic prediction and selection: an application to feed efficiency and its component traits in pigs.J. Anim. Sci.932056–2063. 10.2527/jas.2014-8640
20
EdwardsS. M.SorensenI. F.SarupP.MackayT. F.SorensenP. (2016). Genomic prediction for quantitative traits is improved by mapping variants to gene ontology categories in Drosophila melanogaster.Genetics2031871–1883. 10.1534/genetics.116.187161
21
ErbeM.HayesB. J.MatukumalliL. K.GoswamiS.BowmanP. J.ReichC. M.et al (2012). Improving accuracy of genomic predictions within and between dairy cattle breeds with imputed high-density single nucleotide polymorphism panels.J. Dairy Sci.954114–4129. 10.3168/jds.2011-5019
22
Fernandes JúniorG. A.RosaG. J. M.ValenteB. D.CarvalheiroR.BaldiF.GarciaD. A.et al (2016). Genomic prediction of breeding values for carcass traits in Nellore cattle.Genet. Sel. Evol.48:7. 10.1186/s12711-016-0188-y
23
FinucaneH. K.Bulik-SullivanB.GusevA.TrynkaG.ReshefY.LohP.-R.et al (2015). Partitioning heritability by functional annotation using genome-wide association summary statistics.Nat. Genet.47:1228. 10.1038/ng.3404
24
GaoN.MartiniJ. W. R.ZhangZ.YuanX.ZhangH.SimianerH.et al (2017). Incorporating gene annotation into genomic prediction of complex phenotypes.Genetics207489–501. 10.1534/genetics.117.300198
25
GarnierS.TruongV.BrochetonJ.ZellerT.RovitalM.WildP. S.et al (2013). Genome-wide haplotype analysis of cis expression quantitative trait loci in monocytes.PLoS Genet.9:e1003240. 10.1371/journal.pgen.1003240
26
GianolaD. (2013). Priors in whole-genome regression: the bayesian alphabet returns.Genetics194573–596. 10.1534/genetics.113.151753
27
GianolaD.FernandoR. L.StellaA. (2006). Genomic-assisted prediction of genetic value with semiparametric procedures.Genetics1731761–1776. 10.1534/genetics.105.049510
28
GilmourA.GogelB.CullisB.WelhamS.ThompsonR. (2015). ASReml User Guide Release 4.1 Structural Specification.Hemel Hempstead: VSN International Ltd,
29
GusevA.LeeS. H.TrynkaG.FinucaneH.VilhjálmssonB. J.XuH.et al (2014). Partitioning heritability of regulatory and cell-type-specific variants across 11 common diseases.Am. J. Hum. Genet.95535–552. 10.1016/j.ajhg.2014.10.004
30
HabierD.FernandoR. L.KizilkayaK.GarrickD. J. (2011). Extension of the bayesian alphabet for genomic selection.BMC Bioinformatics12:186. 10.1186/1471-2105-1112-1186
31
HayesB. J.BowmanP. J.ChamberlainA. J.GoddardM. E. (2009). Invited review: genomic selection in dairy cattle: progress and challenges.J. Dairy Sci.92433–443. 10.3168/jds.2008-1646
32
HayesB. J.ChamberlainA. J.McPartlanH.MacleodI.SethuramanL.GoddardM. E. (2007). Accuracy of marker-assisted selection with single markers and marker haplotypes in cattle.Genet. Res.89215–220. 10.1017/S0016672307008865
33
HayesB. J.CoganN. O.PembletonL. W.GoddardM. E.WangJ.SpangenbergG. C.et al (2013). Prospects for genomic selection in forage plant species.Plant Breed.132133–143. 10.1371/journal.pone.0059668
34
HayesB. J.PryceJ.ChamberlainA. J.BowmanP. J.GoddardM. E. (2010). Genetic architecture of complex traits and accuracy of genomic prediction: coat colour, milk-fat percentage, and type in Holstein cattle as contrasting model traits.PLoS Genet.6:e1001139. 10.1001371/journal.pgen.1001139
35
HeD.WangZ.ParidaL. (2015). Data-driven encoding for quantitative genetic trait prediction.BMC Bioinformatics16:S10. 10.1186/1471-2105-16-S1-S10
36
HeS.SchulthessA. W.MirditaV.ZhaoY.KorzunV.BotheR.et al (2016). Genomic selection in a commercial winter wheat population.Theor. Appl. Genet.129641–651. 10.1007/s00122-015-2655-1
37
HeffnerE. L.SorrellsM. E.JanninkJ.-L. (2009). Genomic selection for crop improvement.Crop Sci.491–12. 10.3389/fpls.2013.00023
38
HessM.DruetT.HessA.GarrickD. (2017). Fixed-length haplotypes can improve genomic prediction accuracy in an admixed dairy cattle population.Genet. Sel. Evol.49:54. 10.1186/s12711-017-0329-y
39
HillW. G.GoddardM. E.VisscherP. M. (2008). Data and theory point to mainly additive genetic variance for complex traits.PLoS Genet.4:e1000008. 10.1371/journal.pgen.1000008
40
HindorffL. A.SethupathyP.JunkinsH. A.RamosE. M.MehtaJ. P.CollinsF. S.et al (2009). Potential etiologic and functional implications of genome-wide association loci for human diseases and traits.Proc. Natl. Acad. Sci. U.S.A.1069362–9367. 10.1073/pnas.0903103106
41
JiangY.ReifJ. C. (2015). Modeling epistasis in genomic selection.Genetics201759–768.
42
JiangY.SchmidtR. H.ReifJ. C. (2018). Haplotype-based genome-wide prediction models exploit local epistatic interactions among markers.G3 (Bethesda)81687–1699. 10.1534/g3.117.300548
43
KamanuF. K.MedvedevaY. A.SchaeferU.JankovicB. R.ArcherJ. A.BajicV. B. (2012). Mutations and binding sites of human transcription factors.Front. Genet.3:100. 10.1371/journal.pgen.1006207
44
KarimiZ.SargolzaeiM.RobinsonJ. A. B.SchenkelF. S. (2018). Assessing haplotype-based models for genomic evaluation in Holstein cattle.Can. J. Anim. Sci.98750–759.
45
KindtA. S. D.NavarroP.SempleC. A. M.HaleyC. S. (2013). The genomic signature of trait-associated variants.BMC Genomics14:108. 10.1186/1471-2164-1114-1108
46
KookeR.KruijerW.BoursR.BeckerF.KuhnA.van de GeestH.et al (2016). Genome-wide association mapping and genomic prediction elucidate the genetic architecture of morphological traits in Arabidopsis.Plant Physiol.1702187–2203. 10.1104/pp.15.00997
47
KoufariotisL.ChenY.-P. P.BolormaaS.HayesB. J. (2014). Regulatory and coding genome regions are enriched for trait associated variants in dairy and beef cattle.BMC Genomics15:436. 10.1186/1471-2164-1115-1436
48
LiZ.GaoN.MartiniJ. W. R.SimianerH. (2019). Integrating gene expression data into genomic prediction.Front. Genet.10:126. 10.3389/fgene.2019.00126
49
LorenzanaR. E.BernardoR. (2009). Accuracy of genotypic value predictions for marker-based selection in biparental plant populations.Theor. Appl. Genet.120151–161. 10.1007/s00122-009-1166-3
50
LuanT.WoolliamsJ. A.LienS.KentM.SvendsenM.MeuwissenT. H. E. (2009). The accuracy of genomic selection in Norwegian red cattle assessed by cross-validation.Genetics1831119–1126. 10.1534/genetics.109.107391
51
MackayT. F. (2014). Epistasis and quantitative traits: using model organisms to study gene–gene interactions.Nat. Rev. Genet.15:22. 10.1038/nrg3627
52
MacLeodI. M.BowmanP. J.Vander JagtC. J.Haile-MariamM.KemperK. E.ChamberlainA. J.et al (2016). Exploiting biological priors and sequence variants enhances QTL discovery and genomic prediction of complex traits.BMC Genomics17:144. 10.1186/s12864-016-2443-6
53
MartiniJ. W.GaoN.CardosoD. F.WimmerV.ErbeM.CantetR. J.et al (2017). Genomic prediction with epistasis models: on the marker-coding-dependent performance of the extended GBLUP and properties of the categorical epistasis model (CE).BMC Bioinformatics18:3. 10.1186/s12859-12016-11439-12851
54
MauranoM. T.HumbertR.RynesE.ThurmanR. E.HaugenE.WangH.et al (2012). Systematic localization of common disease-associated variation in regulatory DNA.Science3371190–1195. 10.1126/science.1222794
55
MehrbanH.LeeD. H.MoradiM. H.IlChoC.NaserkheilM.Ibáñez-EscricheN. (2017). Predictive performance of genomic selection methods for carcass traits in Hanwoo beef cattle: impacts of the genetic architecture.Genet. Sel. Evol.49:1. 10.1186/s12711-016-0283-0
56
MeuwissenT. H.HayesB. J.GoddardM. E. (2001). Prediction of total genetic value using genome-wide dense marker maps.Genetics1571819–1829.
57
MeuwissenT. H.OdegardJ.Andersen-RanbergI.GrindflekE. (2014). On the distance of genetic relationships and the accuracy of genomic prediction in pig breeding.Genet. Sel. Evol.46:49. 10.1186/1297-9686-1146-1149
58
MorotaG.Abdollahi-ArpanahiR.KranisA.GianolaD. (2014). Genome-enabled prediction of quantitative traits in chickens using genomic annotation.BMC Genomics15:109. 10.1186/1471-2164-15-109
59
MorotaG.GianolaD. (2014). Kernel-based whole-genome prediction of complex traits: a review.Front. Genet.5:363. 10.3389/fgene.2014.00363
60
MuchaA.WierzbickiH.KamiñskiS.OleñskiK.HeringD. (2019). High-frequency marker haplotypes in the genomic selection of dairy cattle.J. Appl. Genet.60179–186. 10.1007/s13353-019-00489-9
61
MuñozP. R.ResendeM. F.GezanS. A.ResendeM. D. V.de los CamposG.KirstM.et al (2014). Unraveling additive from nonadditive effects using genomic relationship matrices.Genetics1981759–1768. 10.1534/genetics.114.171322
62
NaniJ. P.RezendeF. M.PeñagaricanoF. (2019). Predicting male fertility in dairy cattle using markers with large effect and functional annotation data.BMC Genomics20:258. 10.1186/s12864-019-5644-y
63
NiuH.ZhuB.GuoP.ZhangW. G.XueJ. L.ChenY.et al (2016). Estimation of linkage disequilibrium levels and haplotype block structure in Chinese Simmental and Wagyu beef cattle using high-density genotypes.Livest. Sci.1901–9.
64
OberU.AyrolesJ. F.StoneE. A.RichardsS.ZhuD.GibbsR. A.et al (2012). Using whole-genome sequence data to predict quantitative trait phenotypes in Drosophila melanogaster.PLoS Genet.8:e1002685. 10.1001371/journal.pgen.1002685
65
PalucciV.SchaefferL. R.MigliorF.OsborneV. (2007). Non-additive genetic effects for fertility traits in Canadian holstein cattle (open access publication).Genet. Sel. Evol.39:181. 10.1186/1297-9686-39-2-181
66
PetterssonM.BesnierF.SiegelP. B.CarlborgÖ (2011). Replication and explorations of high-order epistasis using a large advanced intercross line pedigree.PLoS Genet.7:e1002180. 10.1371/journal.pgen.1002180
67
PhillipsP. C. (2008). Epistasis — the essential role of gene interactions in the structure and evolution of genetic systems.Nat. Rev. Genet.9:855. 10.1038/nrg2452
68
PurcellS.NealeB.Todd-BrownK.ThomasL.FerreiraM. A. R.BenderD.et al (2007). PLINK: a tool set for whole-genome association and population-based linkage analyses.Am. J. Hum. Genet.81559–575.
69
RiedelsheimerC.Czedik-EysenbergA.GriederC.LisecJ.TechnowF.SulpiceR.et al (2012). Genomic and metabolic prediction of complex heterotic traits in hybrid maize.Nat. Genet.44217–220.
70
SchaubM. A.BoyleA. P.KundajeA.BatzoglouS.SnyderM. (2012). Linking disease associations with regulatory information in the human genome.Genome Res.221748–1759.
71
SchrootenC.SchopenG.ParkerA.MedleyA.BeatsonP. (2013). Across-breed genomic evaluation based on bovine high density genotypes and phenotypes of bulls and cows.Proc. Assoc. Advmt. Anim. Breed. Genet.20138–141.
72
SonessonA. K.MeuwissenT. H. (2009). Testing strategies for genomic selection in aquaculture breeding programs.Genet. Sel. Evol.41:37. 10.1186/s12864-12017-13557-12861
73
SuG.ChristensenO. F.OstersenT.HenryonM.LundM. S. (2012). Estimating additive and non-additive genetic variances and predicting genetic merits using genome-wide dense single nucleotide polymorphism markers.PLoS One7:e45293. 10.1371/journal.pone.0045293
74
ToghianiS.HayE.SumreddeeP.GearyT. W.RekayaR.RobertsA. J. (2017). Genomic prediction of continuous and binary fertility traits of females in a composite beef cattle breed.J. Anim. Sci.954787–4795.
75
VanRadenP. M. (2008). Efficient methods to compute genomic predictions.J. Dairy Sci.914414–4423.
76
VazquezI.de los CamposG.KlimentidisY. C.RosaG. J.GianolaD.YiN.et al (2012). A comprehensive genetic approach for improving prediction of skin cancer risk in humans.Genetics1921493–1502.
77
VillumsenT. M.JanssL.LundM. S. (2009). The importance of haplotype length and heritability using genomic selection in dairy cattle.J. Anim. Breed. Genet.1263–13.
78
WangD.El-BasyoniI. S.BaenzigerP. S.CrossaJ.EskridgeK. M.DweikatI. (2012). Prediction of genetic values of quantitative traits with epistatic effects in plant breeding populations.Heredity109:313.
79
WhittakerJ. C.ThompsonR.DenhamM. C. (2000). Marker-assisted selection using ridge regression.Genet Res.75249–252.
80
WittenburgD.MelzerN.ReinschN. (2011). Including non-additive genetic effects in Bayesian methods for the prediction of genetic values based on genome-wide markers.BMC Genet.12:74. 10.1186/1471-2156-12-74
81
XiaJ.QiX.WuY.ZhuB.XuL. Y.GaoH. J.et al (2016). Genome-wide association study identifies loci and candidate genes for meat quality traits in Simmental beef cattle.Mamm. Genome27246–255.
82
XuS. (2007). An empirical Bayes method for estimating epistatic effects of quantitative trait loci.Biometrics63513–521.
83
YangJ.BenyaminB.McEvoyB. P.GordonS.HendersA. K.NyholtD. R.et al (2010). Common SNPs explain a large proportion of the heritability for human height.Nat. Genet.42565–569.
84
ZhangZ.DingX.LiuJ.ZhangQ.de KoningD. J. (2011). Accuracy of genomic prediction using low-density marker panels.J. Dairy Sci.943642–3650.
85
ZhongS.DekkersJ. C.FernandoR. L.JanninkJ.-L. (2009). Factors affecting accuracy from genomic selection in populations derived from multiple inbred lines: a barley case study.Genetics182355–364.
86
ZhuB.NiuH.ZhangW.WangZ.LiangY.GuanL.et al (2017). Genome wide association study and genomic prediction for fatty acid composition in Chinese Simmental beef cattle using high density SNP array.BMC Genomics18:464. 10.1186/s12864-017-3847-7
87
ZhuB.ZhuM.JiangJ.NiuH.WangY.WuY.et al (2016). The impact of variable degrees of freedom and scale parameters in Bayesian methods for genomic prediction in Chinese Simmental beef cattle.PLoS One11:e0154118. 10.1371/journal.pone.0154118
Summary
Keywords
genomic prediction, genome annotation, haplotype, Chinese Simmental beef cattle, prediction accuracy
Citation
Xu L, Gao N, Wang Z, Xu L, Liu Y, Chen Y, Xu L, Gao X, Zhang L, Gao H, Zhu B and Li J (2020) Incorporating Genome Annotation Into Genomic Prediction for Carcass Traits in Chinese Simmental Beef Cattle. Front. Genet. 11:481. doi: 10.3389/fgene.2020.00481
Received
12 December 2019
Accepted
17 April 2020
Published
15 May 2020
Volume
11 - 2020
Edited by
Guilherme J. M. Rosa, University of Wisconsin–Madison, United States
Reviewed by
Matthew L. Spangler, University of Nebraska–Lincoln, United States; Fernando Baldi, São Paulo State University, Brazil
Updates
Copyright
© 2020 Xu, Gao, Wang, Xu, Liu, Chen, Xu, Gao, Zhang, Gao, Zhu and Li.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Bo Zhu, zhubo@caas.cnJunya Li, lijunya@caas.cn
This article was submitted to Livestock Genomics, a section of the journal Frontiers in Genetics
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.