GENERAL COMMENTARY article

Front. Genet., 05 November 2013

Sec. Statistical Genetics and Methodology

Volume 4 - 2013 | https://doi.org/10.3389/fgene.2013.00225

The curse of the missing heritability

  • XS

    Xia Shen *

  • Division of Computational Genetics, Department of Clinical Sciences, Swedish University of Agricultural Sciences Uppsala, Sweden

Since “the case of the missing heritability” was highlighted 5 years ago (Maher, ), scientists have been investigating various possible explanations for this issue (Manolio et al., ; Slatkin, ; Eichler et al., ; Zuk et al., ). Recently, Bloom et al. () conducted a linkage analysis in a large yeast Saccharomyces cerevisiae cross with high statistical power to map functional quantitative trait loci (QTL) and found that nearly all the additive genetic contribution can be explained by the detected QTL. It is striking that the “old-fashioned” linkage analysis can resolve the missing heritability problem arisen in the high-throughput genome-wide association study (GWAS) era. Compared to human population studies, an intercross creates large linkage disequilibrium (LD) blocks that greatly enhance statistical power but also reduce QTL mapping resolution. Simple simulations (Figure 1) indicate that the real sources or architecture of missing heritability will remain undiscovered due to LD. Breaking down LD would provide better resolution but reduce the power. This commentary is raised to emphasize the trade-off between resolution and statistical power in mapping functional loci.

Figure 1

Linkage analysis or QTL interval mapping in an experimental design is a classic method in quantitative genetics to detect QTL, which allows inferring QTL effects in an un-typed chromosomal interval harbored by flanking genetic markers (Lynch and Walsh, ). In an F2 cross, the observed LD blocks are often very large, due to limited number of recombination events happened in the F1 individuals, though the recombination rate in yeast is relatively high. For example, among the detected QTL for yeast growth in E6 berbamine (Figure 3 in Bloom et al., ), the two QTL on chromosome 1 covered the two clear LD blocks (not shown) on the chromosome, and the QTL on chromosome 9 covered most of the chromosome. The finding that the detected QTL can explain almost all the narrow sense heritability (h2) is expected given that the kinship estimates using only the significant QTL are similar to the genomic kinship. Even a small number of randomly selected markers can resemble the genomic kinship and give similar heritability estimates (Figure 1B), because the number of LD blocks in the entire genome is limited. The prediction of trait values using detected QTL was good according to cross validation, because the specific F2 population share similar LD patterns, but such prediction would not perform as superior in another population with different LD pattern. Related empirical evidence can be seen in human height (Makowsky et al., ) and marker-assisted selection (Dekkers, ), where detected QTL were unsuccessful for out-sample prediction purposes.

If a future generation (e.g., F8) with small LD blocks is developed from the F2, the statistical power for mapping QTL will decrease. One reason is that a single-locus test for QTL within a large LD block is very likely boosted by multiple QTL within the LD block whose effects are much smaller. The single QTL effect can be simply a combined effect of multiple QTL, and its standard error is underestimated without considering the linkage with other QTL in the same LD region. Assume that there are two functional SNPs x1 and x2 in a chromosomal region with high LD, and the phenotype y is determined by y = x1β1 + x2 β2 + e (1), where β1 and β2 are the effects of the two SNPs; y, x1, and x2 are column vectors of data; e is a vector of residuals. Due to the high LD, x1x2 if x1 and x2 are on the same scale, so that yx11 + β2) + e. In a regression model on the single SNP x1, y = x1β + e (2), the estimated effect for β will be approximately β1 + β2, i.e., a combined effect of both variants. Comparing regression models (1) and (2), the standard error (s.e.) of the estimated β is an underestimate of the s.e. of β1. This is because the s.e. of β1 is inversely proportional to where r is the correlation coefficient between x1 and x2, which is close to 1 due to the high LD, therefore the s.e. of β1 becomes much larger than that of β. When the large LD blocks are broken down, such a combined effect will substantially decrease, leading to lack of statistical power for mapping multiple QTL in the original large LD blocks. One previous empirical example was found in chicken advanced intercross lines (AIL), where only five out of nine QTL detected in the F2 were confirmed by the AIL (Besnier et al., ).

Bloom et al.'s study clearly shows that nearly all the h2 in yeast is written in the DNA, which improves our understanding of missing heritability though some resolution is sacrificed. Researchers are searching for genetic architecture that answers not only where but also what and how the sources contribute to the heritability. However, the curse of missing heritability forces us to choose between resolution and power. For many complex traits, such as human height (Yang et al., ), their polygenic nature makes it extremely difficult to fine-map even the major contribution of the heritability. In future studies, it is important to check the prediction performance in a validation population, in order to show the real sources of missing heritability. Also, biological information and useful tools other than statistical methods need to be developed and utilized.

Statements

Acknowledgments

Xia Shen is funded by a Future Research Leaders grant from Swedish Foundation for Strategic Research (SSF) to Prof. Örjan Carlborg.

References

Summary

Keywords

missing heritability, quantitative trait loci, intercross, linkage analysis, genomic kinship

Citation

Shen X (2013) The curse of the missing heritability. Front. Genet. 4:225. doi: 10.3389/fgene.2013.00225

Received

19 July 2013

Accepted

16 October 2013

Published

05 November 2013

Volume

4 - 2013

Edited by

Frank Emmert-Streib, Queen's University Belfast, UK

Reviewed by

Gaurav Sablok, Istituto Agrario San Michele, Italy; Zhixiang Lu, University of California, Los Angeles, USA; Pavlos Pavlidis, Heidelberg Institute of Theoretical Studies, Germany

Copyright

*Correspondence:

This article was submitted to Statistical Genetics and Methodology, a section of the journal Frontiers in Genetics.

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics