<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="methods-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fgene.2020.00362</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Exact Distribution of Linkage Disequilibrium in the Presence of Mutation, Selection, or Minor Allele Frequency Filtering</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Qu</surname> <given-names>Jiayi</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Kachman</surname> <given-names>Stephen D.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Garrick</surname> <given-names>Dorian</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/37713/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Fernando</surname> <given-names>Rohan L.</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/22048/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Cheng</surname> <given-names>Hao</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/774102/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Animal Science, University of California, Davis</institution>, <addr-line>Davis, CA</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Statistics, University of Nebraska Lincoln</institution>, <addr-line>Lincoln, NE</addr-line>, <country>United States</country></aff>
<aff id="aff3"><sup>3</sup><institution>School of Agriculture, Massey University</institution>, <addr-line>Wellington</addr-line>, <country>New Zealand</country></aff>
<aff id="aff4"><sup>4</sup><institution>Department of Animal Science, Iowa State University</institution>, <addr-line>Ames, IA</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Jacob A. Tennessen, Harvard University, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Guanglin He, Sichuan University, China; Dan Skelly, Jackson Laboratory, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Hao Cheng <email>qtlcheng&#x00040;ucdavis.edu</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Evolutionary and Population Genetics, a section of the journal Frontiers in Genetics</p></fn></author-notes>
<pub-date pub-type="epub">
<day>21</day>
<month>04</month>
<year>2020</year>
</pub-date>
<pub-date pub-type="collection">
<year>2020</year>
</pub-date>
<volume>11</volume>
<elocation-id>362</elocation-id>
<history>
<date date-type="received">
<day>08</day>
<month>11</month>
<year>2019</year>
</date>
<date date-type="accepted">
<day>25</day>
<month>03</month>
<year>2020</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2020 Qu, Kachman, Garrick, Fernando and Cheng.</copyright-statement>
<copyright-year>2020</copyright-year>
<copyright-holder>Qu, Kachman, Garrick, Fernando and Cheng</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>Linkage disequilibrium (LD), often expressed in terms of the squared correlation (<italic>r</italic><sup>2</sup>) between allelic values at two loci, is an important concept in many branches of genetics and genomics. Genetic drift and recombination have opposite effects on LD, and thus <italic>r</italic><sup>2</sup> will keep changing until the effects of these two forces are counterbalanced. Several approximations have been used to determine the expected value of <italic>r</italic><sup>2</sup> at equilibrium in the presence or absence of mutation. In this paper, we propose a probability-based approach to compute the exact distribution of allele frequencies at two loci in a finite population at any generation <italic>t</italic> conditional on the distribution at generation <italic>t</italic> &#x02212; 1. As <italic>r</italic><sup>2</sup> is a function of this distribution of allele frequencies, this approach can be used to examine the distribution of <italic>r</italic><sup>2</sup> over generations as it approaches equilibrium. The exact distribution of LD from our method is used to describe, quantify, and compare LD at different equilibria, including equilibrium in the absence or presence of mutation, selection, and filtering by minor allele frequency. We also propose a deterministic formula for expected LD in the presence of mutation at equilibrium based on the exact distribution of LD.</p></abstract>
<kwd-group>
<kwd>linkage disequilibrium</kwd>
<kwd>effective population size</kwd>
<kwd>mutation rate</kwd>
<kwd>selection</kwd>
<kwd>minor allele frequency filtering</kwd>
</kwd-group>
<contract-num rid="cn001">2018-67015-27957</contract-num>
<contract-sponsor id="cn001">National Institute of Food and Agriculture<named-content content-type="fundref-id">10.13039/100005825</named-content></contract-sponsor>
<counts>
<fig-count count="4"/>
<table-count count="1"/>
<equation-count count="15"/>
<ref-count count="15"/>
<page-count count="10"/>
<word-count count="6354"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Linkage disequilibrium (LD), the non-random association of alleles at two or more loci, is an important concept in various areas of genetics, including evolutionary, quantitative, and statistical genetics. In evolutionary genetics, LD is used to detect the genomic locations of historical selection and to estimate the time of divergence and geographic subdivision between populations (Sved and Hill, <xref ref-type="bibr" rid="B12">2018</xref>). For example, LD was used to date the divergence of the European population from the African population using the HapMap data (Sved et al., <xref ref-type="bibr" rid="B13">2008</xref>). In the field of quantitative and statistical genetics, methods for genomic prediction and genome-wide association studies (GWAS) rely on the existence of LD between the molecular markers and the unobserved quantitative trait loci (QTL).</p>
<p>Two statistics that have been widely used to quantify the LD between allelic values at two loci are: their covariance (<italic>D</italic>); or their squared correlation (<italic>r</italic><sup>2</sup>) (Hill and Robertson, <xref ref-type="bibr" rid="B4">1968</xref>). For a diallelic system, they can be expressed as</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>and</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>respectively, where <italic>p</italic><sub><italic>A</italic><sub><italic>i</italic></sub></sub> represents the frequency of the <italic>ith</italic> allele at locus A, <italic>p</italic><sub><italic>B</italic><sub><italic>j</italic></sub></sub> represents the frequency of the <italic>jth</italic> allele at locus B, and <italic>p</italic><sub><italic>A</italic><sub><italic>i</italic></sub><italic>B</italic><sub><italic>j</italic></sub></sub> represents the frequency of haplotype <italic>A</italic><sub><italic>i</italic></sub><italic>B</italic><sub><italic>j</italic></sub> in the population. Although the covariance of different pairs of alleles depends on their corresponding haplotype frequencies and the respective allele frequencies, a single value of D is sufficient to characterize the disequilibrium between two loci in a diallelic system, i.e., <italic>D</italic> &#x0003D; <italic>D</italic><sub><italic>A</italic><sub>1</sub><italic>B</italic><sub>1</sub></sub> &#x0003D; &#x02212;<italic>D</italic><sub><italic>A</italic><sub>1</sub><italic>B</italic><sub>2</sub></sub> &#x0003D; &#x02212;<italic>D</italic><sub><italic>A</italic><sub>2</sub><italic>B</italic><sub>1</sub></sub> &#x0003D; <italic>D</italic><sub><italic>A</italic><sub>2</sub><italic>B</italic><sub>2</sub></sub>. The value of <italic>D</italic> depends on how the alleles are coded (i.e., alleles <italic>A</italic><sub>1</sub> and <italic>A</italic><sub>2</sub> could be coded as 0 and 1 or as 1 and 0) while the value of <italic>r</italic><sup>2</sup> is invariant to this choice. Therefore, <italic>r</italic><sup>2</sup> is increasingly used as a metric to quantify LD in the literature.</p>
<p>In a finite population, allele frequencies at two loci are subject to random fluctuations due to the stochastic sampling of a finite number of gametes across generations. This random change of the gene frequency between one generation and the next is genetic drift. Due to the randomness of the genetic drift, one initial random-mating population can evolve into one of a finite collection of sub-populations or lines with different allele frequencies at the two loci, and thus with different LD. Usually, this finite collection of possible lines is jointly considered to understand the distribution of the allele frequencies and LD of two linked loci over generations. This is equivalent to considering a finite collection of pairs of loci in one line. The consequences of the genetic drift, mutation, and selection for the collection of lines of two loci equally apply to a collection of pairs of loci in one line. The former conceptualization is used in this paper to understand the effects of genetic drift, mutation, selection, and minor allele frequency (MAF) cutoff on the distribution of LD.</p>
<p>Generally, in a finite population, the opposite effects on LD of genetic drift and recombination lead to an equilibrium value for the expectation of <italic>r</italic><sup>2</sup> when these two forces are counterbalanced. It has been shown that for large values of <italic>N</italic><sub><italic>e</italic></sub><italic>c</italic>, the equilibrium value for the expectation of <italic>r</italic><sup>2</sup> (<inline-formula><mml:math id="M3"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>) is approximately:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mn>4</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mi>c</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>c</italic> is the recombination rate and <italic>N</italic><sub><italic>e</italic></sub> is the effective population size (Sved, <xref ref-type="bibr" rid="B10">1971</xref>; Sved and Feldman, <xref ref-type="bibr" rid="B11">1973</xref>; Hill, <xref ref-type="bibr" rid="B5">1975</xref>).</p>
<p>The above formula, however, does not consider the variation introduced by mutation in the population. When mutation is considered, a balance is expected between the loss of variation by the fixation of one or other allele, and the replenishment of an extinct allele by mutation. An approximation for the equilibrium value of <italic>r</italic><sup>2</sup> in the presence of mutation in a finite population has been developed in Ohta and Kimura (<xref ref-type="bibr" rid="B9">1971</xref>) and Hill (<xref ref-type="bibr" rid="B5">1975</xref>). Instead of using <italic>E</italic>(<italic>r</italic><sup>2</sup>), they used a quantity (<inline-formula><mml:math id="M5"><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>) named the standard linkage deviation (Ohta and Kimura, <xref ref-type="bibr" rid="B8">1969</xref>) to describe the status of this equilibrium, where <inline-formula><mml:math id="M6"><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> is an approximation of the expected value of <italic>r</italic><sup>2</sup> expressed as the ratio of the expectations of the numerator and the denominator of <italic>r</italic><sup>2</sup> (Ohta and Kimura, <xref ref-type="bibr" rid="B9">1971</xref>):</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>E</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The equilibrium value of <inline-formula><mml:math id="M8"><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> is approximated as:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>10</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>4</mml:mn><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mn>22</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mn>13</mml:mn><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>32</mml:mn><mml:mi>&#x003B8;</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003C1;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:mn>6</mml:mn><mml:mi>&#x003C1;</mml:mi><mml:mi>&#x003B8;</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>8</mml:mn><mml:msup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003C1; &#x0003D; 4<italic>N</italic><sub><italic>e</italic></sub><italic>c</italic>, &#x003B8; &#x0003D; 4<italic>N</italic><sub><italic>e</italic></sub><italic>u</italic> and <italic>u</italic> is the mutation rate per site (Ohta and Kimura, <xref ref-type="bibr" rid="B9">1971</xref>). The equilibrium value of <inline-formula><mml:math id="M10"><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> as an approximation of <inline-formula><mml:math id="M11"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> has been described in Walsh and Lynch (<xref ref-type="bibr" rid="B14">2018</xref>). Similarly, this value is considered in this paper as an approximation of <inline-formula><mml:math id="M12"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> in the presence of mutation.</p>
<p>Selection is another important force affecting LD that was not considered in either Sved&#x00027;s or Hill&#x00027;s formulas. Selection occurs at causal variants or QTL during evolution (natural selection) or selective breeding (artificial selection). The impact of selection can either decrease or increase the LD in a population given different scenarios (Mitchell-Olds et al., <xref ref-type="bibr" rid="B7">2007</xref>). The equilibrium among mutation and selection is of special interest in terms of the distribution of allele frequencies and of <italic>r</italic><sup>2</sup>.</p>
<p>In contrast to the approximate methods mentioned above, in this paper we will derive the exact distribution of allele frequencies and LD of two linked loci at equilibrium in the absence or presence of mutation, selection, and MAF cutoff. The distribution of LD in the presence of filtering by MAF is considered in our study due to the prevalent practice of filtering out marker loci with low MAF in genomic analyses used for genetic evaluation or QTL discovery (GWAS).</p>
<p>Our exact distributions are first used to validate the approximate deterministic formulas for <inline-formula><mml:math id="M13"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>, such as Sved&#x00027;s and Hill&#x00027;s approximations. Next we describe, quantify and compare <inline-formula><mml:math id="M14"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> at different equilibria: including equilibrium in the presence of mutation; equilibrium in the presence of mutation and selection; with or without filtering by MAF. Finally, we calibrate Sved&#x00027;s deterministic formula for <inline-formula><mml:math id="M15"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> in the presence of mutation based on the exact distribution of LD. The objective of this paper is to present a computational approach to derive the exact distribution of LD over generations and use it to study the distribution of LD whether or not there is mutation, selection, or filtering by MAF.</p>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>2. Materials and Methods</title>
<p>In this section, we will show how to compute the exact distribution of allele frequencies at two linked loci in generation <italic>t</italic>, given their distribution in generation <italic>t</italic>&#x02212;1, as a function of the effective population size, the recombination rate between the loci, the mutation rate, the selection coefficient, and the MAF threshold. As LD is a function of the allele frequencies at the two loci, its distribution over generations can be calculated based on the distribution of the allele frequencies.</p>
<sec>
<title>2.1. Transition-Matrix Approach in the Presence of Mutation, Selection, or Filtering by Minor Allele Frequency</title>
<p>Two diallelic loci <italic>A</italic> and <italic>B</italic>, with alleles <italic>A</italic><sub>1</sub> and <italic>A</italic><sub>2</sub> at the <italic>A</italic> locus and alleles <italic>B</italic><sub>1</sub> and <italic>B</italic><sub>2</sub> at the <italic>B</italic> locus, are considered. To incorporate selection and MAF threshold, we consider the A locus as the causal variant under selection and B locus as the molecular marker under the MAF filtering. Four possible haplotypes (i.e., <italic>A</italic><sub>1</sub><italic>B</italic><sub>1</sub>, <italic>A</italic><sub>1</sub><italic>B</italic><sub>2</sub>, <italic>A</italic><sub>2</sub><italic>B</italic><sub>1</sub>, and <italic>A</italic><sub>2</sub><italic>B</italic><sub>2</sub>) are possible at these two loci. In a population of size <italic>N</italic><sub><italic>e</italic></sub>, the frequency counts of these four haplotypes can take on <inline-formula><mml:math id="M16"><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mn>3</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>!</mml:mo></mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mo>!</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>!</mml:mo></mml:mrow></mml:mfrac></mml:math></inline-formula> possible values, where for each of these possibilities, the sum of the four counts is 2<italic>N</italic><sub><italic>e</italic></sub>. For example, when <italic>N</italic><sub><italic>e</italic></sub> is equal to two, all possible combinations of haplotype frequency counts are given in <xref ref-type="supplementary-material" rid="SM1">Figure S1</xref>. In general, let <bold>X</bold> be a <italic>k</italic> &#x000D7; 4 matrix with each row representing a possible combination of haplotype frequency counts for some value of <italic>N</italic><sub><italic>e</italic></sub>. Thus, the rows of <bold>X</bold> represent the collection of <italic>k</italic> lines with different allele frequencies at two loci. Let <bold>P</bold><sub><italic>t</italic></sub> denote a <italic>k</italic> &#x000D7; 1 vector with element <italic>i</italic> indicating the probability of the frequency counts in row <italic>i</italic> of <bold>X</bold> at generation <italic>t</italic>. Thus, <bold>P</bold><sub><italic>t</italic></sub> gives the distribution of allele frequencies at generation <italic>t</italic>. We show below that the distribution in generation <italic>t</italic>&#x0002B;1 can be written in general as</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M17"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>P</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle mathvariant="bold"><mml:mtext>A</mml:mtext></mml:mstyle><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>P</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <bold>A</bold> is a <italic>k</italic> &#x000D7; <italic>k</italic> transition matrix that is derived below in various circumstances of recombination, mutation, and selection.</p>
<p>First, consider a line with haplotype frequency counts <inline-formula><mml:math id="M18"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> from <italic>i</italic>th row of <bold>X</bold> at generation <italic>t</italic>, e.g., <inline-formula><mml:math id="M19"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>22</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, where the frequency counts for the four haplotypes, <italic>A</italic><sub>1</sub><italic>B</italic><sub>1</sub>, <italic>A</italic><sub>1</sub><italic>B</italic><sub>2</sub>, <italic>A</italic><sub>2</sub><italic>B</italic><sub>1</sub>, and <italic>A</italic><sub>2</sub><italic>B</italic><sub>2</sub>, are denoted by <italic>f</italic><sub>11</sub>, <italic>f</italic><sub>12</sub>, <italic>f</italic><sub>21</sub> and <italic>f</italic><sub>22</sub>. Ignoring the effects of recombination, mutation, or selection, sampling of 2<italic>N</italic><sub><italic>e</italic></sub> gametes from this line can be modeled by a multinomial process with sample size <italic>n</italic> &#x0003D; 2<italic>N</italic><sub><italic>e</italic></sub> and probability vector <inline-formula><mml:math id="M20"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-itlaic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula> for the four haplotype probabilities. Thus, the distribution of the frequency count in the next generation is given by the <italic>k</italic> &#x000D7; 1 vector <bold>m</bold><sub><italic>i</italic></sub>, where the element <italic>j</italic> of <bold>m</bold><sub><italic>i</italic></sub> is the probability of getting <inline-formula><mml:math id="M21"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> from the Multinomial(<italic>n</italic>, <bold><italic>&#x003B8;</italic></bold><sub><italic>i</italic></sub>) distribution. Given that element <italic>i</italic> of <bold>P</bold><sub><italic>t</italic></sub> is the probability of <inline-formula><mml:math id="M22"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> in generation <italic>t</italic>, the distribution of allele frequencies in the next generation is given by:</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M23"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>P</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle mathvariant="bold"><mml:mtext>M</mml:mtext></mml:mstyle><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>P</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <bold>M</bold> is the <italic>k</italic> &#x000D7; <italic>k</italic> matrix with columns <bold>m</bold><sub><italic>i</italic></sub>, as described above, for <italic>i</italic> &#x0003D; 1, &#x02026;, <italic>k</italic>. The above formula shows how the distribution of allele frequencies change due to genetic drift, ignoring recombination, mutation, and selection. Now we accommodate recombination in computing <bold>P</bold><sub><italic>t</italic>&#x0002B;1</sub>. In gametes produced from generation <italic>t</italic>, the probability of a non-recombinant <italic>A</italic><sub>1</sub><italic>B</italic><sub>1</sub> is <inline-formula><mml:math id="M24"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula>, where <italic>c</italic> is the recombination rate between locus <italic>A</italic> and locus <italic>B</italic>. On the other hand, a recombinant <italic>A</italic><sub>1</sub><italic>B</italic><sub>1</sub> can be produced in one of four ways. They and their associated probabilities are:</p>
<list list-type="order">
<list-item><p>Alleles <italic>A</italic><sub>1</sub> and <italic>B</italic><sub>1</sub> originate from two different <italic>A</italic><sub>1</sub><italic>B</italic><sub>1</sub> haplotypes with probability <inline-formula><mml:math id="M25"><mml:mi>c</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula>.</p></list-item>
<list-item><p>Allele <italic>A</italic><sub>1</sub> originates from an <italic>A</italic><sub>1</sub><italic>B</italic><sub>1</sub> haplotype and <italic>B</italic><sub>1</sub> originates from an <italic>A</italic><sub>2</sub><italic>B</italic><sub>1</sub> haplotype with probability <inline-formula><mml:math id="M26"><mml:mi>c</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula>.</p></list-item>
<list-item><p>Allele <italic>A</italic><sub>1</sub> originates from an <italic>A</italic><sub>1</sub><italic>B</italic><sub>2</sub> haplotype and <italic>B</italic><sub>1</sub> originates from an <italic>A</italic><sub>1</sub><italic>B</italic><sub>1</sub> haplotype with probability <inline-formula><mml:math id="M27"><mml:mi>c</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula>.</p></list-item>
<list-item><p>Allele <italic>A</italic><sub>1</sub> originates from an <italic>A</italic><sub>1</sub><italic>B</italic><sub>2</sub> haplotype and <italic>B</italic><sub>1</sub> originates from an <italic>A</italic><sub>2</sub><italic>B</italic><sub>1</sub> haplotype with probability <inline-formula><mml:math id="M28"><mml:mi>c</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula>.</p></list-item>
</list>
<p>Thus, accounting for recombination, the probability of a <italic>A</italic><sub>1</sub><italic>B</italic><sub>1</sub> haplotype changes from <inline-formula><mml:math id="M29"><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula> to:</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M30"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>Pr</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>c</mml:mi><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>c</mml:mi><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Similarly, due to recombination, the probabilities of the other three haplotypes become:</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M31"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>Pr</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>c</mml:mi><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>22</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>c</mml:mi><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>22</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E10"><label>(10)</label><mml:math id="M32"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>Pr</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>c</mml:mi><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>c</mml:mi><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>22</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>and</p>
<disp-formula id="E11"><label>(11)</label><mml:math id="M33"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>Pr</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>22</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>c</mml:mi><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>22</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>c</mml:mi><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>22</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>22</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Now, to see how mutation alters these probabilities, we let <inline-formula><mml:math id="M34"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> be a vector of the four probabilities from Equations (8) through (11) computed using the four haplotype frequencies in <inline-formula><mml:math id="M35"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. In modeling mutation, we assume that an <italic>A</italic><sub>1</sub> or <italic>B</italic><sub>1</sub> allele mutates to an <italic>A</italic><sub>2</sub> or <italic>B</italic><sub>2</sub> allele with probability <italic>u</italic> and an <italic>A</italic><sub>2</sub> or <italic>B</italic><sub>2</sub> allele mutates to an <italic>A</italic><sub>1</sub> or <italic>B</italic><sub>1</sub> allele with probability <italic>v</italic>. Then, haplotype probabilities following mutation can be computed as:</p>
<disp-formula id="E12"><label>(12)</label><mml:math id="M36"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle mathvariant="bold-italic"><mml:mi>T</mml:mi></mml:mstyle><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where</p>
<disp-formula id="E13"><mml:math id="M37"> <mml:mstyle mathvariant="bold"><mml:mtext>T</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>u</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>u</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>v</mml:mi></mml:mtd><mml:mtd><mml:mi>v</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>u</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:msup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>u</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>u</mml:mi></mml:mtd><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>u</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>v</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:mtd><mml:mtd><mml:mi>v</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>v</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>u</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>u</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mi>u</mml:mi><mml:mi>v</mml:mi></mml:mtd><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>v</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>u</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>v</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>v</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mi>u</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>v</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>v</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>u</mml:mi></mml:mtd><mml:mtd><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>v</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula>
<p>Furthermore, to incorporate selection in the model, a selection coefficient <italic>s</italic> that reduces the allele frequency of <italic>A</italic><sub>1</sub> at locus <italic>A</italic>, which is a causal variant, is considered. Conditional on the haplotype probabilities in <inline-formula><mml:math id="M38"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, i.e., <inline-formula><mml:math id="M39"><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula><mml:math id="M40"><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula><mml:math id="M41"><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <inline-formula><mml:math id="M42"><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, the haplotype probabilities following selection can be computed as</p>
<disp-formula id="E14"><mml:math id="M43"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>s</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>s</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr></mml:mtr><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>s</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>s</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr></mml:mtr><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>s</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr></mml:mtr><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>s</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Finally, to compute the distribution of allele frequencies in generation <italic>t</italic>&#x0002B;1 using Equation (6), where the forces of the genetic drift, recombination, mutation and selection are simultaneously considered, the matrix <bold>A</bold> is defined such that element <italic>j</italic> of column <italic>i</italic> contains the probability of getting <inline-formula><mml:math id="M44"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> from a Multinomial <inline-formula><mml:math id="M45"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo><mml:mo>*</mml:mo><mml:mo>*</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> distribution, for <italic>i, j</italic> &#x0003D; 1, &#x02026;, <italic>k</italic>.</p>
<p>The value of <italic>r</italic><sup>2</sup> can be computed for a sub-population or line with haplotype frequency counts in any row of <bold>X</bold>. For example, consider the frequency counts in row 11 of the <bold>X</bold> matrix given in <xref ref-type="supplementary-material" rid="SM1">Figure S1</xref>, where <italic>f</italic><sub>11</sub> &#x0003D; 0, <italic>f</italic><sub>12</sub> &#x0003D; 2, <italic>f</italic><sub>21</sub> &#x0003D; 1 and <italic>f</italic><sub>22</sub> &#x0003D; 1. From these frequency counts, <italic>Pr</italic>(<italic>A</italic><sub>1</sub>) &#x0003D; 1/2, <italic>Pr</italic>(<italic>B</italic><sub>1</sub>) &#x0003D; 1/4, and <italic>Pr</italic>(<italic>A</italic><sub>1</sub><italic>B</italic><sub>1</sub>) &#x0003D; 0 are obtained, and <italic>r</italic><sup>2</sup> calculated using Equations (1) and (2) is 1/3 for a line with haplotype frequency counts in <inline-formula><mml:math id="M46"><mml:msubsup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>x</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> from <xref ref-type="supplementary-material" rid="SM1">Figure S1</xref>. The distribution of <italic>r</italic><sup>2</sup> in generation <italic>t</italic> is given by values of <italic>r</italic><sup>2</sup> corresponding to the frequency counts in each row of <bold>X</bold> together with the probabilities for haplotype frequency counts given in <bold>P</bold><sub><italic>t</italic></sub>. Note that haplotype frequency counts in rows of <bold>X</bold> that have indeterminate values of <italic>r</italic><sup>2</sup> (i.e., when one or other allele is extinct) are not used to compute the distribution of <italic>r</italic><sup>2</sup>. Similarly, the distribution of <italic>r</italic><sup>2</sup> with MAF cutoff is given by considering only the values of <italic>r</italic><sup>2</sup> and their probabilities corresponding to the rows of <bold>X</bold> with MAF at the B locus &#x02265;5%. MAF of 0.05 is used as a threshold in the present study.</p>
<p>Starting with an allele frequency of 0.5 at each locus and linkage equilibrium between the two loci, the expected value of <italic>r</italic><sup>2</sup> was computed over generations given some values of the mutation rate, recombination rate, selection coefficient and effective population size. Mutation rate of <italic>u</italic> &#x0003D; <italic>v</italic> &#x0003D; 0 is used to represent the absence of mutation, while <italic>u</italic> &#x0003D; <italic>v</italic> &#x0003D; 1 &#x000D7; 10<sup>&#x02212;9</sup> is used to represent the existence of mutation. Similarly, a selection coefficient of <italic>s</italic> &#x0003D; 0 is used to represent the absence of selection, while <italic>s</italic> &#x0003D; 0.1 or <italic>s</italic> &#x0003D; 0.01 is used to represent the existence of selection.</p>
</sec>
<sec>
<title>2.2. Data Analysis</title>
<p>In our analysis, different population parameters are considered. Populations with <italic>N</italic><sub><italic>e</italic></sub> ranging from 5 to 50 in intervals of 5 are compared under different recombination rates, mutation rates (i.e., <italic>u</italic> &#x0003D; 0 or <italic>u</italic> &#x0003D; 1<italic>e</italic><sup>&#x02212;9</sup>) and selection coefficients for the causal variant (i.e., <italic>s</italic> &#x0003D; 0.1 or <italic>s</italic> &#x0003D; 0.01 for locus <italic>A</italic>), in the absence or presence of filtering by MAF (MAF&#x02265;0.05 for locus <italic>B</italic>). The recombination rates ranged from 0.01 to 0.1 in intervals of 0.01 and from 0.1 to 0.5 in intervals of 0.1.</p>
<p>The recombination rate for adjacent markers in a bovine 777 k SNP panel was considered to mimic a realistic scenario. The recombination rate for adjacent markers was estimated to be 6.25 &#x000D7; 10<sup>&#x02212;5</sup> using the Kosambi map function given that the mean distance between adjacent markers is around 5 kb (Espigolan et al., <xref ref-type="bibr" rid="B3">2013</xref>), and the average genetic distance per unit of physical distance in bovine genome is 1.25 cM/Mb (Arias et al., <xref ref-type="bibr" rid="B1">2009</xref>). A period of more than 3,000 generations was used to ensure that the haplotype frequencies had reached their equilibrium values. At equilibrium, the distribution of LD stays constant over generations (i.e., <italic>P</italic><sub><italic>t</italic>&#x0002B;1</sub> &#x0003D; <italic>P</italic><sub><italic>t</italic></sub>). That distribution was used to describe, quantify, and compare <inline-formula><mml:math id="M48"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> at different equilibria in the absence or presence of mutation, selection or MAF filtering. Results from a population similar to the international black and white Holstein dairy cattle are presented in these cases, for which the effective population size is estimated to be about 50 (Kim and Kirkpatrick, <xref ref-type="bibr" rid="B6">2009</xref>; Wray et al., <xref ref-type="bibr" rid="B15">2019</xref>).</p>
<p>Furthermore, the exact distributions of <inline-formula><mml:math id="M49"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> derived under different scenarios were used to validate the approximate deterministic formulas from Sved and Feldman (<xref ref-type="bibr" rid="B11">1973</xref>) and Hill (<xref ref-type="bibr" rid="B5">1975</xref>). To generalize our results to larger effective population size, a non-linear regression model of the form <inline-formula><mml:math id="M50"><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>b</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>c</mml:mi></mml:mrow></mml:mfrac></mml:math></inline-formula>, following Sved&#x00027;s formula, was considered. The parameters in that model were estimated by non-linear least squares using the data generated from our transition-matrix approach.</p>
<p>The authors state that all data necessary for confirming the conclusions presented in the article are represented fully within the article.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<sec>
<title>3.1. Comparison of Sved&#x00027;s and Hill&#x00027;s Approximations to the Exact Distributions of <inline-formula><mml:math id="M51"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> Derived From the Transition-Matrix Approach</title>
<p>In this section, Sved&#x00027;s and Hill&#x00027;s approximations were compared to the exact distribution of <inline-formula><mml:math id="M52"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> derived from our transition-matrix approach. The relationship between equilibrium value of <italic>r</italic><sup>2</sup> (<inline-formula><mml:math id="M53"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>), recombination rate and mutation rate is shown in <xref ref-type="fig" rid="F1">Figure 1</xref> for these three approaches. In the absence of mutation and selection, Sved&#x00027;s formula showed consistency with the exact values from our transition-matrix approach (<xref ref-type="fig" rid="F1">Figure 1A</xref>). On the other hand, Hill&#x00027;s formula had a better fit to the exact values of <inline-formula><mml:math id="M54"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> (<xref ref-type="fig" rid="F1">Figure 1B</xref>) in the presence of mutation. However, neither Sved&#x00027;s nor Hill&#x00027;s deterministic formulae are accurate enough to describe <inline-formula><mml:math id="M55"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> in the presence of mutation. Therefore, we proposed an adjusted non-linear regression model to correct Sved&#x00027;s approximation to predict <inline-formula><mml:math id="M56"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> for larger effective population sizes in the presence of mutation. The mean square of errors is used to evaluate the fit of these three models and shows that our calibrated non-linear regression model was significantly better than Sved&#x00027;s or Hill&#x00027;s approximations.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Comparison of Sved&#x00027;s and Hill&#x00027;s approximation to exact distribution of <inline-formula><mml:math id="M47"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> (scatter points) derived from transition-matrix approach. Mean square errors are shown in the parentheses. &#x0201C;Calibrated&#x0201D; denotes the non-linear regression formula derived from the transition-matrix approach. <bold>(A)</bold> In the absence of mutation and selection. <bold>(B)</bold> In the presence of mutation but no selection.</p></caption>
<graphic xlink:href="fgene-11-00362-g0001.tif"/>
</fig>
</sec>
<sec>
<title>3.2. Trajectory of LD Over the Generations Until Equilibrium Is Reached</title>
<p>The trajectory of evolving LD over generations computed from our transition-matrix approach is shown in this section, in the absence or presence of mutation and selection, where selection was applied with either a low selection coefficient of 0.01 or a high selection coefficient of 0.1. Starting with an allele frequency of 0.5 at each locus and linkage equilibrium between the two loci, expected values of LD [i.e., <italic>E</italic>(<italic>r</italic><sup>2</sup>)] over generations are displayed in <xref ref-type="fig" rid="F2">Figure 2</xref>. Note that the trajectory may be quite different, depending on the initial conditions used.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Expected value of LD <italic>E</italic>(<italic>r</italic><sup>2</sup>) over generations under four different conditions at different recombination rates (c) for a population of effective size <italic>Ne</italic> &#x0003D; 50.</p></caption>
<graphic xlink:href="fgene-11-00362-g0002.tif"/>
</fig>
<p>In the absence of mutation and selection, the expected value of <italic>r</italic><sup>2</sup> over generations increases to its maximum and reaches its &#x0201C;equilibrium&#x0201D; status. Value of <inline-formula><mml:math id="M57"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> increases as c decreases due to the decreasing breakdown of LD with lower recombination rates. Typically, when the recombination rate is relatively small (e.g., <italic>c</italic> &#x0003D; 6.25 &#x000D7; 10<sup>&#x02212;5</sup>), the expected value of LD at &#x0201C;equilibrium&#x0201D; is almost equal to 1. However, at this &#x0201C;equilibrium&#x0201D; stage, frequencies of lines (<bold>P</bold><sub><italic>t</italic></sub>) keep changing though the distribution of allele frequencies remain unchanged. This is because forces of recombination and drift are balanced in lines with determinate <italic>r</italic><sup>2</sup>, and frequencies of lines with determinate <italic>r</italic><sup>2</sup> almost proportionally decrease toward 0.</p>
<p>In the absence of selection but with mutation, the expected value of <italic>r</italic><sup>2</sup> increases and stays at an apparent equilibrium for several generations and then decreases to its true equilibrium value, where the mutation-drift equilibrium is reached. During the apparent equilibrium that is initially observed, the expected value of <italic>r</italic><sup>2</sup> is identical to the equilibrium value in the absence of mutation and selection, where the forces of drift and recombination are balanced. When this stage is first reached, the effect of mutation is negligible because the mutation rate is low relative to the frequency of lines with segregating loci. The apparent equilibrium ends when the frequency is high for lines where one or both loci are fixed (<italic>r</italic><sup>2</sup> is indeterminate) and the frequency of lines with segregating loci becomes close to the rate of mutation. Then, alleles introduced due to mutation into these lines have a noticeable effect on the distribution of allele frequencies and expected LD changes. At the end of the apparent equilibrium, the most frequent lines have only one haplotype. In the lines with two haplotypes, which have low frequencies, <italic>r</italic><sup>2</sup> is either indeterminate or has value 1.0. When mutation introduces a third haplotype in a line where <italic>r</italic><sup>2</sup> is 1.0, it drops in value. Further, when mutation introduces a third haplotype in a line where <italic>r</italic><sup>2</sup> is indeterminate, the value of <italic>r</italic><sup>2</sup> will be close to zero. Thus, as mutation becomes noticeable, the expected value of LD decreases. When the true equilibrium is reached, the loss of variation by fixation is balanced by its replenishment by mutation. In other words, the frequency of lines where loci are segregating will remain non-zero. Therefore, when mutation is present, the probabilities of allele frequencies (<italic>P</italic><sub><italic>t</italic></sub>) stay constant over generations when the true equilibrium is reached.</p>
<p>In the presence of both mutation and selection, lines with segregating loci become low in frequency at an earlier stage due to selection, particularly when the selection coefficient is large (e.g., s = 0.1), and thus, mutation starts to have a noticeable effect on allele frequency at an earlier stage in the presence of selection relative to when selection is absent. When selection coefficient is relatively large (e.g., s = 0.1), the expected value of <italic>r</italic><sup>2</sup> reaches its maximum value early and then decreases sharply to two plateaus. When the recombination rate is low, e.g., <italic>c</italic> &#x02264; 0.01, the difference between these two plateaus is small.</p>
<p>In the absence of selection, allele frequencies at two loci tend to be similar. However, when locus A is under selection, allele frequencies at these two loci tend to be different such that LD is lower than when selection is absent. This effect, however, is negligible when the selection coefficient is small (e.g., s = 0.01). Thus, a selection coefficient of 0.1 is used to present results when mutation and selection are present in the next section.</p>
<p>The expected values of <italic>r</italic><sup>2</sup> over generations in the presence of filtering by MAF (i.e., MAF &#x02265;0.05 for locus B) are shown in <xref ref-type="supplementary-material" rid="SM1">Figure S2</xref>, where a similar pattern as for <italic>E</italic>(<italic>r</italic><sup>2</sup>) over generations are observed.</p>
</sec>
<sec>
<title>3.3. Distributions of LD at Equilibrium</title>
<p>In a population where mutation is present and <italic>N</italic><sub><italic>e</italic></sub> &#x0003D; 50, the equilibrium value of <italic>r</italic><sup>2</sup> (<inline-formula><mml:math id="M59"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>) is shown in <xref ref-type="fig" rid="F3">Figure 3</xref>, when selection is absent or present, for different recombination rates. As explained in last section, the equilibrium values of LD in the absence of selection are always higher than those in the presence of selection. The equilibrium value of LD (i.e., <inline-formula><mml:math id="M60"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>) with filtering by MAF (i.e., MAF &#x02265;5% at locus B) is higher than its corresponding value in the absence of filtering.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Relationship between the expectation of LD at equilibrium (<inline-formula><mml:math id="M58"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>) and recombination rate (<italic>c</italic>) in the absence or presence of selection or filtering by MAF at locus B for a population of effective size <italic>N</italic><sub><italic>e</italic></sub> &#x0003D; 50 with mutation rate <italic>u</italic> &#x0003D; 1.0 &#x000D7; 10<sup>&#x02212;9</sup>. Only locus A is under selection with selection coefficient <italic>s</italic> &#x0003D; 0.1.</p></caption>
<graphic xlink:href="fgene-11-00362-g0003.tif"/>
</fig>
<p>To understand why filtering by MAF increases <inline-formula><mml:math id="M61"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>, the <italic>k</italic> lines with different haplotype frequencies were divided into two groups: lines where MAF at the locus B is &#x0003C;5% and lines where MAF &#x02265;5%. Next, the transition-matrix approach was used to compute the frequency of each of these lines at equilibrium. In lines where <italic>r</italic><sup>2</sup> is defined, the equilibrium frequency of each line is plotted against its <italic>r</italic><sup>2</sup> value in <xref ref-type="fig" rid="F4">Figure 4</xref>; the color blue is used for the first group of lines where MAF at the locus B is &#x0003C;5% and red is used for the second group of lines where MAF is &#x02265;5%; frequencies in the absence of selection are given in <xref ref-type="fig" rid="F4">Figure 4A</xref> and those in the presence of selection are given in <xref ref-type="fig" rid="F4">Figure 4B</xref>.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Distribution of LD (<italic>r</italic><sup>2</sup>) with recombination rate (<italic>c</italic>) of 6.25 &#x000D7; 10<sup>&#x02212;5</sup> at equilibrium in the presence of mutation but no selection or in the presence of mutation and selection in a population of effective size <italic>Ne</italic> &#x0003D; 50. Only locus A is under selection. <bold>(A)</bold> Presence of mutation but no selection (s = 0). <bold>(B)</bold> Presence of mutation and selection (s = 0.1).</p></caption>
<graphic xlink:href="fgene-11-00362-g0004.tif"/>
</fig>
<p>In the absence of selection (<xref ref-type="fig" rid="F4">Figure 4A</xref>), a proportion of about 0.41 of the lines were in the first group and a proportion of about 0.59 were in the second group. It can be seen from this figure that most of the <italic>r</italic><sup>2</sup> values in the first group were small, where around 45% of the lines had an <italic>r</italic><sup>2</sup> value &#x0003C; 0.001. In contrast, the second group had many lines with large <italic>r</italic><sup>2</sup> values; around 15% of the lines had <italic>r</italic><sup>2</sup> &#x0003D; 1. This difference in the distribution of <italic>r</italic><sup>2</sup> in the two groups shows why filtering by MAF at the B locus (removing lines from the first group, which had an abundance of low <italic>r</italic><sup>2</sup> values) would increase the expected value of <italic>r</italic><sup>2</sup>.</p>
<p>The reason for the lower value of <italic>r</italic><sup>2</sup> in the first group is that, due to filtering by MAF at the B locus, this group has a large proportion of lines with recent mutations at the B locus. Consider a line where alleles are segregating at the A locus but fixed at the B locus. Such a line would belong to neither of the groups because <italic>r</italic><sup>2</sup> is not defined in a line where one locus is fixed. However, the introduction of a new allele at the B locus due to mutation will result in <italic>r</italic><sup>2</sup> becoming defined, but, typically, at a very low level because it results from the association in a single haplotype. The MAF in this line with the new mutation will be <inline-formula><mml:math id="M62"><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula>, and it will belong to the first group provided that <inline-formula><mml:math id="M63"><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>05</mml:mn></mml:math></inline-formula>.</p>
<p><xref ref-type="fig" rid="F4">Figure 4B</xref> gives the distribution of <italic>r</italic><sup>2</sup> for the two groups in the presence of selection, and it can be seen that still there is a greater abundance of low <italic>r</italic><sup>2</sup> values in the first group. The difference between the two groups, however, is smaller than in the absence of selection, and this explains the smaller effect of filtering on equilibrium value of <italic>r</italic><sup>2</sup> when selection is present (<xref ref-type="fig" rid="F3">Figure 3</xref>).</p>
</sec>
<sec>
<title>3.4. Extrapolation of Exact <inline-formula><mml:math id="M64"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> to Larger Population Size by Non-linear Modeling</title>
<p>The two existing deterministic formulas (i.e., Sved&#x00027;s and Hill&#x00027;s formula) for <inline-formula><mml:math id="M65"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> are not accurate as shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. Thus, a non-linear regression formula by recombination rate was calibrated using the exact values of <inline-formula><mml:math id="M66"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> calculated from our transition-matrix approach. <inline-formula><mml:math id="M67"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> is computed as the expectation of <italic>r</italic><sup>2</sup> at equilibrium (i.e., mean <italic>r</italic><sup>2</sup> weighted by the frequency of each of the k possible lines at equilibrium). In the following, this formula is referred to as the calibrated non-linear regression formula. The form of the non-linear regression formula follows that of Sved&#x00027;s formula as presented below:</p>
<disp-formula id="E15"><label>(13)</label><mml:math id="M68"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>To study the extrapolation of <inline-formula><mml:math id="M69"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> from small population sizes (i.e., <italic>N</italic><sub><italic>e</italic></sub> &#x02264; 50) to predict those for a larger population size, we split the values of <inline-formula><mml:math id="M70"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> calculated from our transition-matrix approach into training and validation sets. Exact values of <inline-formula><mml:math id="M71"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> with <italic>N</italic><sub><italic>e</italic></sub> &#x02264; 40 are included in the training set, and the remaining with <italic>N</italic><sub><italic>e</italic></sub> &#x0003E;40 are used for validation. The prediction accuracy is assessed using mean square error (MSE). As shown in <xref ref-type="supplementary-material" rid="SM1">Table S1</xref>, prediction accuracy from the calibrated non-linear regression formula is substantially higher than those from Sved&#x00027;s or Hill&#x00027;s formulas. Finally, parameters in Equation (13) are estimated using all available values of <inline-formula><mml:math id="M72"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>, and parameter estimates are shown in <xref ref-type="table" rid="T1">Table 1</xref>. Typically, the estimates of &#x003B2;<sub>0</sub> and &#x003B2;<sub>1</sub> are inversely related to the values of recombination rates.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Calibration (estimation of &#x003B2;<sub><italic>i</italic></sub><italic>s</italic>) of the non-linear regression model under different recombination rates (c).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>c</bold></th>
<th valign="top" align="center"><bold>&#x003B2;<sub>0</sub></bold></th>
<th valign="top" align="center"><bold>&#x003B2;<sub>1</sub></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">6.25 &#x000D7; 10<sup>&#x02212;5</sup></td>
<td valign="top" align="center">2.26</td>
<td valign="top" align="center">1557.11</td>
</tr>
<tr>
<td valign="top" align="left">0.01</td>
<td valign="top" align="center">1.75</td>
<td valign="top" align="center">21.41</td>
</tr>
<tr>
<td valign="top" align="left">0.02</td>
<td valign="top" align="center">1.49</td>
<td valign="top" align="center">15.13</td>
</tr>
<tr>
<td valign="top" align="left">0.03</td>
<td valign="top" align="center">1.31</td>
<td valign="top" align="center">12.58</td>
</tr>
<tr>
<td valign="top" align="left">0.04</td>
<td valign="top" align="center">1.17</td>
<td valign="top" align="center">11.10</td>
</tr>
<tr>
<td valign="top" align="left">0.05</td>
<td valign="top" align="center">1.05</td>
<td valign="top" align="center">10.09</td>
</tr>
<tr>
<td valign="top" align="left">0.06</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">9.33</td>
</tr>
<tr>
<td valign="top" align="left">0.07</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">8.74</td>
</tr>
<tr>
<td valign="top" align="left">0.08</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">8.25</td>
</tr>
<tr>
<td valign="top" align="left">0.09</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">7.84</td>
</tr>
<tr>
<td valign="top" align="left">0.1</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">7.49</td>
</tr>
<tr>
<td valign="top" align="left">0.2</td>
<td valign="top" align="center">0.28</td>
<td valign="top" align="center">5.48</td>
</tr>
<tr>
<td valign="top" align="left">0.3</td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">4.49</td>
</tr>
<tr>
<td valign="top" align="left">0.4</td>
<td valign="top" align="center">&#x02212;0.19</td>
<td valign="top" align="center">3.85</td>
</tr>
<tr>
<td valign="top" align="left">0.5</td>
<td valign="top" align="center">&#x02212;0.38</td>
<td valign="top" align="center">3.38</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To further validate the accuracy of the extrapolation approach, the expected value of LD at equilibrium between adjacent markers in a 777k SNPs chip (i.e.,<italic>c</italic> &#x0003D; 6.25 &#x000D7; 10<sup>&#x02212;5</sup>) was predicted for a Nellore cattle population, for which the effective population size is about 100 (Brito et al., <xref ref-type="bibr" rid="B2">2013</xref>), using Equation (13) with &#x003B2;<sub>0</sub> &#x0003D; 2.26 and &#x003B2;<sub>1</sub> &#x0003D; 1557.11. In a real genotyped Nellore cattle population (Espigolan et al., <xref ref-type="bibr" rid="B3">2013</xref>), the estimated mean and standard deviation for <inline-formula><mml:math id="M73"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> was 0.17 and 0.20, respectively. The predicted <inline-formula><mml:math id="M74"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> is 0.08, 0.449, and 0.976 using the calibrated non-linear regression formula, Hill&#x00027;s formula, and Sved&#x00027;s formula, respectively. Only the predicted <inline-formula><mml:math id="M75"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> from the calibrated non-linear regression formula falls in the range 0.17&#x000B1;0.2, while the other two are out of the measured range from this real population.</p>
<p>We further studied the accuracy of the extrapolation approach for very large effective population size, e.g., <italic>Ne</italic> &#x0003D; 10, 000 in human. However, <inline-formula><mml:math id="M76"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> did not align with the biological expectation, which indicates that our non-linear formulas for extrapolation of <inline-formula><mml:math id="M77"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> to extremely large population size does not work because the non-linear regression formulas are derived from data with <italic>N</italic><sub><italic>e</italic></sub> &#x02264; 50.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>A contribution of this article is to propose a computational transition-matrix approach to deriving the distribution of LD between two loci over generations in the presence of multiple genetic forces including drift, mutation, and selection. These distributions of LD are also studied in the presence of filtering by MAF (i.e., MAF &#x02265;0.05). In addition to deriving exact distribution of LD, several critical results emerge from the proposed approach.</p>
<list list-type="order">
<list-item><p><bold>The expected value of LD at equilibrium decreases in the presence of selection</bold>. Decrease of LD caused by mutation is magnified in the presence of selection due to a higher fixation rate of the favorable allele. In the presence of mutation but no selection (or in the presence of mutation and weak selection), allele frequencies at both loci are similar due to genetic drift, and the LD between them tends to remain high. Conversely, in the presence of mutation and strong selection, this phenomenon is disrupted by the selection process resulting in diverse gene frequencies at the two loci,and low LD between them is observed.</p></list-item>
<list-item><p><bold>Caution is needed when LD between a causal variant and a marker is inferred after filtering out marker loci with low MAF</bold>. LD between a causal variant and a marker variant with high MAF (i.e., MAF &#x02265;5%) is higher than that between a causal variant and all marker variants, especially in the presence of mutation but no selection. That is, <inline-formula><mml:math id="M78"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> attributed to LD between a causal variant and marker variants with high MAF is relatively higher than that between a causal variant and marker variant with low MAF which leads to the reduction of the overall expected LD. In practice, LD is sometimes inferred from molecular marker panel with MAF cutoff applied. The inferred value is usually used to estimate important population parameters (e.g., effective population size). Here we have demonstrated the potential increase of LD brought by MAF cutoff and caution is needed when inferring LD in the presence of filtering by MAF.</p></list-item>
<list-item><p><bold>&#x0201C;Fake&#x0201D; equilibrium may appear in the presence of mutation</bold>. Two equilibrium stages are observed in the presence of mutation. The first &#x0201C;equilibrium&#x0201D; indicates the balance between fixation increasing LD and recombination decreasing LD in lines with determinant <italic>r</italic><sup>2</sup> given the effect of mutation is negligible. When most loci become fixed at the end of the first &#x0201C;equilibrium&#x0201D; stage, change of allele frequency caused by mutation is of importance and reduces the expected LD. At the second equilibrium stage, the balance between mutation decreasing LD and fixation increasing LD is reached. Note that linkage equilibrium between two loci of allele frequency 0.5 is assumed in the initial population in <xref ref-type="fig" rid="F2">Figure 2</xref>, and distribution of allele frequencies in the initial population affects the existence of the &#x0201C;fake&#x0201D; equilibrium.</p></list-item>
</list>
<sec>
<title>4.1. Exact <inline-formula><mml:math id="M79"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> for Large Effective Population Sizes</title>
<p>The maximum effective population size (<italic>N</italic><sub><italic>e</italic></sub>) presented in this paper is 50 (Kim and Kirkpatrick, <xref ref-type="bibr" rid="B6">2009</xref>; Wray et al., <xref ref-type="bibr" rid="B15">2019</xref>), which is the estimated <italic>N</italic><sub><italic>e</italic></sub> for the international black and white Holstein dairy cattle population. Larger <italic>N</italic><sub><italic>e</italic></sub> was not studied in this paper due to computational limitations. When <italic>N</italic><sub><italic>e</italic></sub> &#x0003D; 50, the transition matrix <bold>A</bold> is square of order <inline-formula><mml:math id="M80"><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>4</mml:mn><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:mn>8</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>176</mml:mn><mml:mo>,</mml:mo><mml:mn>851</mml:mn></mml:math></inline-formula>, which requires more than 250 gigabytes memory to store. In addition, it takes around 5 h to complete the computational process for the analysis of 3,500 generations. Note that the memory complexity to store <bold>A</bold> is O(<inline-formula><mml:math id="M81"><mml:msubsup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>). When <italic>N</italic><sub><italic>e</italic></sub> &#x0003D; 100, the estimated processing memory to store the transition matrix <bold>A</bold> is more than 14 terabytes. This computational problem may be addressed by parallel computing. This idea, however, needs further investigation and is out of the scope of this paper. Thus, to generalize the transition-matrix approach a non-linear regression model was calibrated using exact values of <inline-formula><mml:math id="M82"><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>. The proposed calibrations may provide a better description of the relationship between LD, effective population size, recombination rate, and mutation.</p>
</sec>
</sec>
<sec sec-type="data-availability-statement" id="s5">
<title>Data Availability Statement</title>
<p>The scripts used to generate the data can be found at <ext-link ext-link-type="uri" xlink:href="https://github.com/Jiayi-Qu/Distribution_of_LD">https://github.com/Jiayi-Qu/Distribution_of_LD</ext-link>.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>HC and RF conceived the study. JQ, HC, and RF undertook the analysis and wrote the draft. DG and SK contributed to the analysis. All authors contributed to the final version of manuscript, read, and approved the final manuscript.</p>
</sec>
<sec id="s7">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ack><p>RF is grateful to Prof. William G. Hill for extensive discussions of early results from the approach presented here. This manuscript has been released as a Pre-Print at <ext-link ext-link-type="uri" xlink:href="https://www.biorxiv.org/content/10.1101/794347v1">https://www.biorxiv.org/content/10.1101/794347v1</ext-link>.</p>
</ack>
<sec sec-type="supplementary-material" id="s8">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2020.00362/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fgene.2020.00362/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.PDF" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arias</surname> <given-names>J. A.</given-names></name> <name><surname>Keehan</surname> <given-names>M.</given-names></name> <name><surname>Fisher</surname> <given-names>P.</given-names></name> <name><surname>Coppieters</surname> <given-names>W.</given-names></name> <name><surname>Spelman</surname> <given-names>R.</given-names></name></person-group> (<year>2009</year>). <article-title>A high density linkage map of the bovine genome</article-title>. <source>BMC Genet.</source> <volume>10</volume>:<fpage>18</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2156-10-18</pub-id><pub-id pub-id-type="pmid">19393043</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brito</surname> <given-names>F.</given-names></name> <name><surname>Sargolzaei</surname> <given-names>M.</given-names></name> <name><surname>Braccini Neto</surname> <given-names>J.</given-names></name> <name><surname>Cobuci</surname> <given-names>J.</given-names></name> <name><surname>Pimentel</surname> <given-names>C.</given-names></name> <name><surname>Barcellos</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>In-depth pedigree analysis in a large Brazilian nellore herd</article-title>. <source>Genet. Mol. Res.</source> <volume>12</volume>, <fpage>5758</fpage>&#x02013;<lpage>5765</lpage>. <pub-id pub-id-type="doi">10.4238/2013.November.22.2</pub-id><pub-id pub-id-type="pmid">24301944</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Espigolan</surname> <given-names>R.</given-names></name> <name><surname>Baldi</surname> <given-names>F.</given-names></name> <name><surname>Boligon</surname> <given-names>A. A.</given-names></name> <name><surname>Souza</surname> <given-names>F. R.</given-names></name> <name><surname>Gordo</surname> <given-names>D. G.</given-names></name> <name><surname>Tonussi</surname> <given-names>R. L.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Study of whole genome linkage disequilibrium in nellore cattle</article-title>. <source>BMC Genomics</source> <volume>14</volume>:<fpage>305</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2164-14-305</pub-id><pub-id pub-id-type="pmid">23642139</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hill</surname> <given-names>W.</given-names></name> <name><surname>Robertson</surname> <given-names>A.</given-names></name></person-group> (<year>1968</year>). <article-title>Linkage disequilibrium in finite populations</article-title>. <source>Theor. Appl. Genet.</source> <volume>38</volume>, <fpage>226</fpage>&#x02013;<lpage>231</lpage>. <pub-id pub-id-type="doi">10.1007/BF01245622</pub-id><pub-id pub-id-type="pmid">24442307</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hill</surname> <given-names>W. G.</given-names></name></person-group> (<year>1975</year>). <article-title>Linkage disequilibrium among multiple neutral alleles produced by mutation in finite population</article-title>. <source>Theor. Popul. Biol.</source> <volume>8</volume>, <fpage>117</fpage>&#x02013;<lpage>126</lpage>. <pub-id pub-id-type="doi">10.1016/0040-5809(75)90028-3</pub-id><pub-id pub-id-type="pmid">1198348</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>E. S.</given-names></name> <name><surname>Kirkpatrick</surname> <given-names>B. W.</given-names></name></person-group> (<year>2009</year>). <article-title>Linkage disequilibrium in the North American Holstein population</article-title>. <source>Anim. Genet.</source> <volume>40</volume>, <fpage>279</fpage>&#x02013;<lpage>288</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-2052.2008.01831.x</pub-id><pub-id pub-id-type="pmid">19220233</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mitchell-Olds</surname> <given-names>T.</given-names></name> <name><surname>Willis</surname> <given-names>J. H.</given-names></name> <name><surname>Goldstein</surname> <given-names>D. B.</given-names></name></person-group> (<year>2007</year>). <article-title>Which evolutionary processes influence natural genetic variation for phenotypic traits?</article-title> <source>Nat. Rev. Genet.</source> <volume>8</volume>, <fpage>845</fpage>&#x02013;<lpage>856</lpage>. <pub-id pub-id-type="doi">10.1038/nrg2207</pub-id><pub-id pub-id-type="pmid">17943192</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ohta</surname> <given-names>T.</given-names></name> <name><surname>Kimura</surname> <given-names>M.</given-names></name></person-group> (<year>1969</year>). <article-title>Linkage disequilibrium at steady state determined by random genetic drift and recurrent mutation</article-title>. <source>Genetics</source> <volume>63</volume>, <fpage>229</fpage>&#x02013;<lpage>238</lpage>.<pub-id pub-id-type="pmid">5365295</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ohta</surname> <given-names>T.</given-names></name> <name><surname>Kimura</surname> <given-names>M.</given-names></name></person-group> (<year>1971</year>). <article-title>Linkage disequilibrium between two segregating nucleotide sites under the steady flux of mutations in a finite population</article-title>. <source>Genetics</source> <volume>68</volume>, <fpage>571</fpage>&#x02013;<lpage>580</lpage>.<pub-id pub-id-type="pmid">5120656</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sved</surname> <given-names>J.</given-names></name></person-group> (<year>1971</year>). <article-title>Linkage disequilibrium and homozygosity of chromosome segments in finite populations</article-title>. <source>Theor. Popul. Biol.</source> <volume>2</volume>, <fpage>125</fpage>&#x02013;<lpage>141</lpage>. <pub-id pub-id-type="doi">10.1016/0040-5809(71)90011-6</pub-id><pub-id pub-id-type="pmid">5170716</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sved</surname> <given-names>J.</given-names></name> <name><surname>Feldman</surname> <given-names>M.</given-names></name></person-group> (<year>1973</year>). <article-title>Correlation and probability methods for one and two loci</article-title>. <source>Theor. Popul. Biol.</source> <volume>4</volume>, <fpage>129</fpage>&#x02013;<lpage>132</lpage>. <pub-id pub-id-type="doi">10.1016/0040-5809(73)90008-7</pub-id><pub-id pub-id-type="pmid">4726005</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sved</surname> <given-names>J. A.</given-names></name> <name><surname>Hill</surname> <given-names>W. G.</given-names></name></person-group> (<year>2018</year>). <article-title>One hundred years of linkage disequilibrium</article-title>. <source>Genetics</source> <volume>209</volume>, <fpage>629</fpage>&#x02013;<lpage>636</lpage>.<pub-id pub-id-type="pmid">29967057</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sved</surname> <given-names>J. A.</given-names></name> <name><surname>McRae</surname> <given-names>A. F.</given-names></name> <name><surname>Visscher</surname> <given-names>P. M.</given-names></name></person-group> (<year>2008</year>). <article-title>Divergence between human populations estimated from linkage disequilibrium</article-title>. <source>Am. J. Hum. Genet.</source> <volume>83</volume>, <fpage>737</fpage>&#x02013;<lpage>743</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2008.10.019</pub-id><pub-id pub-id-type="pmid">19012875</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Walsh</surname> <given-names>B.</given-names></name> <name><surname>Lynch</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <source>Evolution and Selection of Quantitative Traits</source>. <publisher-loc>Oxford: Oxford University Press</publisher-loc>.</citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wray</surname> <given-names>N. R.</given-names></name> <name><surname>Kemper</surname> <given-names>K. E.</given-names></name> <name><surname>Hayes</surname> <given-names>B. J.</given-names></name> <name><surname>Goddard</surname> <given-names>M. E.</given-names></name> <name><surname>Visscher</surname> <given-names>P. M.</given-names></name></person-group> (<year>2019</year>). <article-title>Complex trait prediction from genome data: contrasting EBV in livestock to PRS in humans</article-title>. <source>Genetics</source> <volume>211</volume>, <fpage>1131</fpage>&#x02013;<lpage>1141</lpage>. <pub-id pub-id-type="doi">10.1534/genetics.119.301859</pub-id><pub-id pub-id-type="pmid">30967442</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> Financial support was provided by the United States Department of Agriculture, Agriculture and Food Research Initiative National Institute of Food and Agriculture Competitive grant no. 2018-67015-27957.</p>
</fn>
</fn-group>
</back>
</article> 