ORIGINAL RESEARCH article

Front. Genet., 18 January 2021

Sec. Human and Medical Genomics

Volume 11 - 2020 | https://doi.org/10.3389/fgene.2020.594318

SNP-Density Crossover Maps of Polymorphic Transposable Elements and HLA Genes Within MHC Class I Haplotype Blocks and Junction

  • 1. Faculty of Health and Medical Sciences, Medical School, The University of Western Australia, Crawley, WA, Australia

  • 2. Division of Basic Medical Science and Molecular Medicine, Department of Molecular Life Science, Tokai University School of Medicine, Isehara, Japan

Abstract

The genomic region (~4 Mb) of the human major histocompatibility complex (MHC) on chromosome 6p21 is a prime model for the study and understanding of conserved polymorphic sequences (CPSs) and structural diversity of ancestral haplotypes (AHs)/conserved extended haplotypes (CEHs). The aim of this study was to use a set of 95 MHC genomic sequences downloaded from a publicly available BioProject database at NCBI to identify and characterise polymorphic human leukocyte antigen (HLA) class I genes and pseudogenes, MICA and MICB, and retroelement indels as haplotypic lineage markers, and single-nucleotide polymorphism (SNP) crossover loci in DNA sequence alignments of different haplotypes across the Olfactory Receptor (OR) gene region (~1.2 Mb) and the MHC class I region (~1.8 Mb) from the GPX5 to the MICB gene. Our comparative sequence analyses confirmed the identity of 12 haplotypic retroelement markers and revealed that they partitioned the HLA-A/B/C haplotypes into distinct evolutionary lineages. Crossovers between SNP-poor and SNP-rich regions defined the sequence range of haplotype blocks, and many of these crossover junctions occurred within particular transposable elements, lncRNA, OR12D2, MUC21, MUC22, PSORS1A3, HLA-C, HLA-B, and MICA. In a comparison of more than 250 paired sequence alignments, at least 38 SNP-density crossover sites were mapped across various regions from GPX5 to MICB. In a homology comparison of 16 different haplotypes, seven CEH/AH (7.1, 8.1, 18.2, 51.x, 57.1, 62.x, and 62.1) had no detectable SNP-density crossover junctions and were SNP poor across the entire ~2.8 Mb of sequence alignments. Of the analyses between different recombinant haplotypes, more than half of them had SNP crossovers within 10 kb of LTR16B/ERV3-16A3_I, MLT1, Charlie, and/or THE1 sequences and were in close vicinity to structurally polymorphic Alu and SVA insertion sites. These studies demonstrate that (1) SNP-density crossovers are associated with putative ancestral recombination sites that are widely spread across the MHC class I genomic region from at least the telomeric OR12D2 gene to the centromeric MICB gene and (2) the genomic sequences of MHC homozygous cell lines are useful for analysing haplotype blocks, ancestral haplotypic landscapes and markers, CPSs, and SNP-density crossover junctions.

Introduction

The human major histocompatibility complex (MHC), also referred to as human leukocyte antigen (HLA), is investigated continuously because of its importance in the regulation of the innate and adaptive immune system, autoimmunity, and transplantation (Dawkins et al., 1999; Vandiedonck and Knight, 2009; Lokki and Paakkanen, 2019). The genomic region of the human MHC encompasses approximately 160 coding genes including three distinct structural regions: class I with the classical and non-classical HLA class I genes (HLA-A, -B, -C, -F, -G, and -E) and ~39 non-HLA genes, class II with the classical and non-classical HLA class II genes (HLA-DRB1, -DRA, -DQA1, -DQB1, -DQA2, -DQB2, -DPA1, and -DPB1) and class III that harbours more than 60 genes including the complement genes, TNF, NFKBIL2, and many other genes that code for cytokines, transcription factors, structural and developmental proteins (Shiina et al., 2004, 2009). The MHC class I and class II gene clusters contain numerous sequence duplications, insertions and deletions and considerable sequence diversity or polymorphisms (Trowsdale and Knight, 2013) that have accumulated into distinct multilocus haplotypes with relatively high population frequencies (>1%) (Awdeh et al., 1983; Degli-Esposti et al., 1992; Dawkins et al., 1999; Yunis et al., 2003; Goodin et al., 2018). These date from at least the beginning of human expansion and dispersal out of Africa, 50,000–100,000 years ago (Henn et al., 2012; López et al., 2015). The MHC multilocus haplotypes have been associated strongly with many diseases (Lokki and Paakkanen, 2019). On the basis of the large number of known HLA-B alleles, more than 20,000 different MHC multilocus haplotypes might be distributed worldwide in human populations, with less than a hundred in certain localised populations such as the Europeans (Steele and Lloyd, 2015; Jensen et al., 2017; Goodin et al., 2018). The common Northern European HLA haplotype HLA-A1-B8-C7-DRB3-DQ2 (8.1AH) is estimated to have diverged from a single common ancestor about 23,500 years ago (Smith et al., 2006).

Although the MHC is highly polymorphic for single-nucleotide polymorphisms (SNPs), the degree of polymorphism (SNP density per 100 kb) depends on which haplotypes (haploid genotypes) are compared. There are at least two main types of genomic haplotype blocks that are studied for SNP variations: (1) those that are constructed on the basis of linkage disequilibrium (LD) statistical tests of a contiguous set of SNP markers (Ahmad et al., 2003; Walsh et al., 2003; Miretti et al., 2005; Blomhoff et al., 2006) and (2) those constructed from alignments of SNP density maps or genotyped alleles that identify well-defined haplotype blocks or segmental structures without using LD tests (Alper et al., 1983, 2006; Degli-Esposti et al., 1992; Dawkins et al., 1999; Aly et al., 2006; Smith et al., 2006; Lam et al., 2015; Alper and Larsen, 2017). If employed independently of each other, the two methods can result in unrelated single and/or multilocus haplotype block patterns (Yunis et al., 2003; Alper et al., 2006; Jensen et al., 2017). Homologous haplotype sequences have only a few detectable SNPs extended over a long-range of multilocus regions (Smith et al., 2006), whereas a large number of SNPs of varying density are detected in comparisons between different MHC class I haplotypes (Gaudieri et al., 1999, 2000; Miretti et al., 2005; Shiina et al., 2006, 2009; Jensen et al., 2017; Norman et al., 2017). The absence of SNPs over megabases of continuous sequence within the same MHC haplotypes is described as conserved sequence polymorphisms (CSPs) within conserved extended haplotypes (CEHs) (Yunis et al., 2003; Alper et al., 2006) and/or ancestral haplotypes (AHs) (Degli-Esposti et al., 1992; Dawkins et al., 1999), such as 8.1CEH/AH (Price et al., 1999; Aly et al., 2006; Smith et al., 2006; Gambino et al., 2018), 7.1CEH/AH (Gaudieri et al., 1997; Dunn et al., 2005), 57.1CEH/AH (Dunn et al., 2005), 38.1CEH/AH (Romero et al., 2007) and the Sardinian haplotype 18.2CEH/AH (Contu et al., 1989; Bilbao et al., 2006). The non-LD, SNP-poor, long-range haplotypic sequences are mainly contained within polymorphic frozen blocks (PFBs) (Gaudieri et al., 1997; Dawkins et al., 1999) or fixed (conserved) haplospecific blocks (Alper et al., 2006; Barquera et al., 2020).

SNPs within many different recombinant haplotypes are absent for relatively much shorter distances ranging between 10 and 1,000 kb such as those found within PFBs of 60–300 kb (Gaudieri et al., 1997; Dawkins et al., 1999) and/or SNP-LD-blocks of ~18–50 kb (Daly et al., 2001; Jeffreys et al., 2004; Miretti et al., 2005; Blomhoff et al., 2006). The SNP-LD-block based on statistical associations between the frequencies of two or more genotyped loci in population studies cannot map the classical CEH/AH or PFB architectural structures directly or reliably (Schaid et al., 2002; Alper et al., 2006; Slatkin, 2008), whereas reliable linkage mapping is usually dependent on pedigree studies of particular genotyped markers to evaluate their linkage or segregation in meiosis or on phased genomic sequences (Alper and Larsen, 2017) such as those that have been sequenced or genotyped using multilocus HLA-captured haplotype phasing (Guo et al., 2006), de novo assembled trios (Jensen et al., 2017), MHC homozygous cell lines (Dorak et al., 2006; Horton et al., 2008; Norman et al., 2017), sperm (Cullen et al., 2002; Kirkness et al., 2013) or single chromosomes (Murphy et al., 2016). SNP-LD analyses often fail to detect linkage of multiple loci within the conserved haplotype structure as effectively as the genes that may be involved in disease susceptibility or resistance because of the use of non-haplotypic SNP markers (Alper et al., 2006; Slatkin, 2008; Alper and Larsen, 2017). Nevertheless, the SNP-LD-blocks that were identified by LD or long-range haplotype (LRH) and extended haplotype homozygosity (EHH) tests (Traherne, 2008) of the MHC genomic regions include a variety of genotyped haplotypic microsatellites (Karell et al., 2000; Doxiadis et al., 2007), SNPs (Ahmad et al., 2003; de Bakker et al., 2006; Shiina et al., 2006; Smith et al., 2006; Romero et al., 2007; Lam et al., 2013), and indels (WGS500 Consortium et al., 2014; Jensen et al., 2017; Huang et al., 2019) as well as structural dimorphic retroelements (REs), such as Alu, SVA, LTR, and HERVs (Kulski and Dunn, 2005; Kulski et al., 2011).

Segmental shuffling is a meiotic recombination or crossing over process between different haplotypes (Gaudieri et al., 1997; Traherne et al., 2006) that often occurs within nucleotide sequences in regions between the alpha, beta, epsilon and delta frozen polymorphic blocks (Dawkins et al., 1999; Traherne et al., 2006; Romero et al., 2007), although breakpoints have been reported also within and between the HLA class I genes within the alpha (Lam et al., 2013) and beta blocks (Nair et al., 2006) and between the HLA class II genes within the delta block (Jeffreys et al., 2004; Larsen et al., 2014). Many MHC recombinant haplotypes appear to have originated in relatively recent times (Smith et al., 2006; Lam et al., 2013) due to the pressures of bottlenecks, migrations and gene flow, inbreeding, and outbreeding in various times of abundance and deprivation (van Oosterhout, 2009; Lobkovsky et al., 2019; Wang et al., 2020). With the formation of new human MHC haplotypes during and/or after speciation, many of the high-frequency (>1%) AHs were preserved over numerous generations and migrations; even across different ethnic populations as deduced from the European (Contu et al., 1989; Aly et al., 2006; Bilbao et al., 2006; Smith et al., 2006) and Asian haplotypes (Lam et al., 2013, 2015, 2017).

Much of genomic sequence diversity, including haplotype diversity, is driven by molecular mechanisms such as DNA repair, replication, single point mutations, indels, recombination, duplication, conversion, transposition, and segmental rearrangements (Gu et al., 2008; Brawand et al., 2014; Lin and Gokcumen, 2019). In addition, interspersed repeat sequences that contribute to >50% of the human genomic content (de Koning et al., 2011) have been implicated in a variety of these DNA molecular processes (Moolhuijzen et al., 2010; George and Alani, 2012; Raviram et al., 2018; Lu et al., 2020). The identification of transposable elements (TEs) near the junctions of duplicated genes (Kulski et al., 1997, 1999b, 2004) and at ectopic and meiotic recombination sites (Myers et al., 2010; Altemose et al., 2017; Kent et al., 2017) emphasise their role in driving genomic diversity. Interspersed REs, because of their mobility, hypermutability, and potential role in meiotic recombination, are an integral part of molecular drive (Dover, 1982) that together with point mutations, gene conversion (Madrigal et al., 1993; Adamek et al., 2015) and balancing selection (van Oosterhout, 2009) with a component of multiplicative fitness (Lobkovsky et al., 2019) probably have generated and maintained haplotypic polymorphisms in the MHC class I regions. This multifunctional role for active REs is evidenced in part by the structural biallelic Alu, SVA, LTR, and HERVs located near to or within putative recombination hotspots throughout the MHC class I, II, and III genomic regions (Kulski et al., 2011). In recent years, the proposed broad roles for TE and polymorphisms in the regulation of meiotic recombination (a mechanism that undoubtedly generated the MHC haplotype diversity in humans) has gained increasing attention (Myers et al., 2008; Zamudio et al., 2015; Altemose et al., 2017; Kent et al., 2017; Bourgeois and Boissinot, 2019).

Both first-generation and second-generation sequencing methods have produced phased genomic sequences of representative MHC haplotypes by using MHC homozygous cell lines (Horton et al., 2004, 2008; Stewart et al., 2004; Traherne et al., 2006; Norman et al., 2017). These phased MHC genomic sequences are important reference DNA sequences that provide representative haplotypes for better informed large population studies and for mapping heterozygous sequence reads such as by inference graphs (Dilthey et al., 2016), SNP-LD based haplotype frequencies (Romero et al., 2007) and EEH tests (Lam et al., 2015), especially for disease associations (Alper and Larsen, 2017; Lokki and Paakkanen, 2019). Although Norman et al. (2017) produced an important database for 95 MHC homozygous cell lines of assembled and resolved MHC genomic sequences, they limited their own analysis to the multilocus alleles and haplotypes of the HLA classical class I and class II genes, MUC22 and the structural diversity of C4 duplications. Missing from their analysis are the many REs, repeats and retrotransposable subfamilies, as well as the amplified and duplicated members of the genomic DNA that make up >50% of the human DNA content and that contribute to disease (Ayarpadikannan and Kim, 2014; Payer et al., 2017; Payer and Burns, 2019), gene regulation and recombination (Moolhuijzen et al., 2010; Myers et al., 2010; Altemose et al., 2017; Chuong et al., 2017) and to the duplicated segmental organisation of the human and other primate MHC genomic structures (Kulski et al., 1997, 1999a,b; Anzai et al., 2003; Kulski et al., 2004).

The purpose of the present study was to extend the Norman et al. (2017) analysis by investigating the haplotypic linkages between the MHC class I genic and intergenic regions including HLA-F, HLA-G, MICA, MICB, eight HLA pseudogenes (HLA-V, -P, -H, -T, -K, -U, -W, and -J) and a set of previously published biallelic REs, AluOR (Kulski et al., 2014), AluHF, AluHG, AluHJ, AluTF, AluMICB (Kulski and Dunn, 2005; Kulski et al., 2019), HERVK9 (Kulski et al., 2008), MER9 (Kulski et al., 2009) and four biallelic SVA haplotypic markers, SVA-HA, SVA-HC, SVA-HB, and SVA-HF (Kulski et al., 2010, 2011). A further aim was to identify and characterise the ancestral SNP-density crossover (XO) loci in DNA sequence alignments of different haplotype blocks or segments from the GPX5 gene in the OR gene region telomeric of HLA-F to the centromeric MICB gene within the MHC class I genomic region. The overall results of the study suggest that the SNP XOs are indicators of haplotype XO, which in turn point to putative ancestral recombination sites that are widely distributed across the 2-Mb-MHC class I genomic region from telomeric of HLA-F to centromeric of MICB.

Materials and Methods

The haplotype data of 95 MHC genomic sequences sequenced and assembled from HLA-homozygous cell lines by Norman et al. (2017) at NCBI BioProject with the accession number PRJEB6763 (https://www.ncbi.nlm.nih.gov/bioproject/) were downloaded as Fasta files and used for the analyses described below. The other MHC genomic sequences used in haplotype analyses were the GRChr38.p13 (GCF_000001405.39) of the chromosome 6 reference NC_000006.12 at the NCBI (https://www.ncbi.nlm.nih.gov/assembly/GCF_000001405.39/), UCSC (https://genome.ucsc.edu/cgi-bin/hgGateway) and eEnsembl (http://asia.ensembl.org/Homo_sapiens/Info/Index) browsers and databases, the eight human reference haplotypes described by Horton et al. (2008), the chimpanzee sequence of Anzai et al. (2003) and the gorilla sequence of Wilming et al. (2013). All of the Fasta sequences downloaded from the public archives were submitted to the RepeatMasker webserver (http://www.repeatmasker.org/cgi-bin/WEBRepeatMasker) for output files of annotated members of the interspersed repetitive DNA families, their locations in the sequence and their relative similarity or identity in comparison to reference sequences of SINEs, LINEs, LTRs, ERVs, DNA elements, small RNA, and simple repeats. For the online analysis, RepeatMasker used the Dfam database (3.0) for the repeat sequence comparisons (Hubley et al., 2016) (http://www.dfam.org) because since 20th May 2019, it no longer had access to the RepBase library of repetitive elements (Bao et al., 2015) previously provided by GIRI (https://www.girinst.org/repbase/). The main difference between the Dfam database and RepBase for our analysis was that Dfam listed many Alu-like short sequences as SVA, whereas we were interested only in the SVA mosaic of 500 to 1,800 bp in RepBase with structures similar to those described by Shen et al. (1994). Thus, we used four dimorphic SVA sequences (SVA-HA, SVA-HC, SVA-HB, and SVA-MIC) previously reported by Kulski et al. (2010) and added three new dimorphic SVA sequence markers to this analysis (Table 1).

Table 1

Retroelement or microsatNearest flanking (/) genesLocation within genome reference Ch38/hg38, Chr 6Popln frequencies caucasian/Japanese (n= 88–260)References
AluOROR12D2 intron29396132–293962630.140.32Kulski et al., 2014
AluOR13′OR12D129416044Present study
AluHFZFP57/HLA-F~29710985*0.230.06Dunn et al., 2002
AluHGHLA-G/HLA-H~29850749*0.30Kulski et al., 2001
0.300.21Dunn et al., 2002
AluHJHLA-J/ETF1P1~30030620*0.250.38Dunn et al., 2002
AluTFMUC21/MUC22~31003947*0.110.08Dunn et al., 2003
AluP5MICA/MICB~31470733*Present study
AluMICBMICB intron 1~31498446*0.12Kulski et al., 2002a
0.160.12Kulski and Dunn, 2005
HERVK9HLA-G/HLA-H29875649–298818290.370.59Kulski et al., 2008
sMER9 (1)HLA-G/HLA-H~29881317*0.660.41Kulski et al., 2008
LTR13HLA-K/HLA-U29929971–29930908Present study
sMER9 (2)HLA-U/HLA-A29936175–29936676Kulski et al., 2009
LTR5LTRIM26/HLA-L~30221451Present study
MER5/LTR33HLA-C/HLA-B~31313186Present study
HAL1/MER5AMICA/MICB31418238–31418519Present study
LTR9MICA/MICB31423445–31424086Present study
SVAOR3′GPX628501515–28503131Present study
SVA-HFLTR16/HLA-F29717873–29720770.140.00Kulski et al., 2010
SVA-16**HLA-H/HLA-T29895386–29896449fixedPresent study
SVA-HAHLA-K/HLA-A29932087–299337530.260.06Kulski et al., 2010
SVA-T26**TRIM26/HLA-L30221503–30222724Present study
SVA-ER**MICC/HLA-E30474489–30475999Present study
SVA-EG**HLA-E/GNL130498159–30499333Present study
SVA-M21**MUC21/MUC2230992538–30993994Present study
SVA-M22**MUC22/C6orf1531066602–31068056Present study
SVA-HCHCG27/HLA-C31243860–312453220.100.03Kulski et al., 2010
SVA-CBHLA-C/HLA-B~31310982*Present study
SVA-HBHLA-C/HLA-B~31329940*0.650.25Kulski et al., 2010
SVA-MICMICA/MICB31453745–31456553Kulski et al., 2010
9.5-kb delHLA-C/HLA-B~31298645*Present study
(ATAG)nHLA-G/MICF29838629–29838750Present study
(CAGAGA)nHLA-G/MICF29838997–29839045Present study
(ATAA)nHLA-A/HLA-W29949553–29949592Present study
(ATTT)nHLA-A/HLA-W29949590–29949639Present study
(TTTA)nTRIM26/HLA-L30221462–30221500Present study
(GAGG)nMUC22/C60rf15~ 31066254Present study
(TTTC)nHCG27/HLA-C31236405–31236481Kulski et al., 1997
(ACA)nHCG27/HLA-C31239846–31239880Kulski et al., 1997
(TTCC)nHCG27/HLA-C31241314–31241352Kulski et al., 1997
(TTAT)nHLA-C/HLA-B31321185–31321227Kulski et al., 1997
(CTG)nwithin MICA31412369–31412393Mizuki et al., 1997
(TGT)nwithin MICA~31412394*Present study

Dimorphic retroelements (absent or present) and STR analysed in this study.

*

Approximate location because these deletions, retroelements or STR are absent from the Ch38/hg38 Genome Reference that has the HLA haplotype of HLA-A*03:01:01:01/ B*07:02:01:01/C*07:02:01:03/.

**

SVA were present in all 95 haplotypes of the Norman et al. (2017) sequences and chimpanzee (Anzai et al., 2003), and therefore are fixed in humans.

Norman et al. (2017) provided the alleles of the HLA-A, -B, and -C class I genes for all the 95 cell line sequences shown in Supplementary Table 1. We confirmed the alleles of the HLA class I genes and included the alleles of HLA-E, -F, and -G, and the MICA and MICB genes and eight HLA-A class I pseudogenes (Supplementary Table 2) in the 95 cell line sequences by comparing them to the IMGT HLA allele sequences (IMGTRelease 3.38.0) using the DNA sequence assembly software Sequencher ver.5.0 (Gencode http://www.genecodes.com). The alleles that were not in the IMGT HLA allele databases (Robinson et al., 2019) at https://www.ebi.ac.uk/ipd/imgt/hla/ are reported here as “new” without providing any further information about the novel nucleotide or amino acid differences. We also found that the FTQW01000001.1 sequence provided by Norman et al. (2017) as the chimpanzee “Clint” (Pan troglodytes genome assembly, contig: 1_COX_Oct2016_Scaffold, whole genome shotgun sequence) has strong identity with the COX cell line sequence that harbours the 8.1AH haplotype A*01:01:01:01/ B*08:01:01:01/ C*07:01:01:01 (Horton et al., 2008).

We added a laboratory identifier number (ID_1 to ID_95) to each of the Norman et al. (2017) sequences (Supplementary Table 2) for ease of identification in comparative sequence analysis. A shorthand identifier for the MHC CEH/AH haplotypes based on the HLA-B allele such as 7.1CEH/AH, 8.1CEH/AH, 13.1CEH/AH was used as previously described (Degli-Esposti et al., 1992, Dawkins et al., 1999, Dorak et al., 2006). The alleles of the HLA class I genes, MIC genes, HLA class I pseudogenes and HLA class II genes were determined also for the GRChr38p13 genomic reference sequence, which corresponds to the 7.1AH of the PGF homozygous cell line (Horton et al., 2008), shown in Supplementary Table 3. The dimorphic RE and microsatellite markers that were searched for and identified by RepeatMasker in the 95 MHC genomic sequences are shown in Table 1. The RE dimorphisms (absence or presence) were easily recognised in each of the RepeatMasker outputs because of their positions within or close proximity to other TE elements and short tandem repeats (STRs). For example, the MER9/HERVK9-int/MER9 insertion at nucleotide positions (nts) 160655 to 166834 in Supplementary Table 4 is flanked by a string of telomeric LTR16B2/MLT1F1/ STR/AluY/STR/L1ME3 elements and a string of centromeric Charlie9 (nts, 89–303)/Charlie9 (nts, 1283–1803)/L1PA10/LMLT1F1/THE1C elements that are easily identified in the RepeatMasker outputs with the solitary MER5 and HERVK9-int deletion at their corresponding locations.

Comparative sequence alignments between two or more sequences to evaluate SNP densities and determine XO regions between SNP-poor regions (SPR) of <20 SNPs per 100 kb and SNP-rich regions (SRR) of >100 SNPs per 100 kb were performed with the web-based MultiPipMaker alignment program (http://pipmaker.bx.psu.edu/cgi-bin/multipipmaker) by uploading the Fasta sequence files, a RepeatMasker output file and using the MultiPipMaker setting for single coverage as described by Schwartz et al. (2000) to generate the optimal sequence alignment. SNPs in the alignments were counted twice manually, and an average number was presented in the results. Obvious assembly errors, polynucleotides, simple microsatellite repeats and indels were not counted as SNPs. Also, a series of many adjoining SNPs (e.g., >5 SNPs in a string of 50 nucleotides) or SNPs within 50 bp of obvious sequencing errors with runs of unspecified nucleotides (Ns) and/or inconsistent long strings of deletions were not counted. The length of sequence alignments usually ranged between 50 and 500 kb depending on (1) the segments targeted for the analysis and the ease of SNP manual counting in the pdf outputs of the nucleotide alignments and/or (2) the length of the Percentage Identity Plot (PIP) output for reproduction as a convenient and readable image. The targeted sequences were selected and trimmed from the Fasta files that had been previous downloaded from the NCBI BioProject, accession number PRJEB6763. The software program Genetyx ver.20 (GENETYX Co., Tokyo, Japan) was used with the Selector function set to select and trim to obtain the required Fasta file sequences with the genomic sequence target positions taken from those listed in the RepeatMasker output text file (Supplementary Table 3). The T-Coffee multiple sequence alignment tool at EMBL-EBI (https://www.ebi.ac.uk/Tools/msa/tcoffee/) was used to construct multiple sequences of ERV3-16A3-int in the Fasta format and the CLUSTALW (1.83) format.

Results

MHC Haplotype Sequences

Of the 95 human MHC haplotypes sequenced by Norman et al. (2017), 82 differed at least at one of the 9 loci, HLA-A, -C, -B, -DRB1, -DRB345, -DQA1, -DQB1, -DPA1, and -DPB1. However, there were 46 sequences representing 18 haplotypes that had the same combination of HLA-A, -C, and -B alleles for at least one haplotype pair. Furthermore, 70 sequences represented 19 different HLA-C/-B haplotypes and 67 sequences represented 23 different HLA-A/-C haplotypes (Supplementary Table 1) with a homologous alignment for at least one haplotype pair.

In this study, the haplotypic alleles of 56 loci were analysed, ranging between the OR gene region and the MHC class I region including the classical HLA-A, -B, and -C loci, the non-classical HLA-F, -G, and -E loci, 8 HLA pseudogenes, MICA and MICB, 8 Alu loci, 13 SVA loci, 2 MER9 loci, the HERVK9 locus, 5 LTR or MER5 loci, and 12 STR loci (Tables 13, Supplementary Table 2). The 9.5-kb MER5/LTR33 indel between the HLA-C and HLA-B loci also contained within its sequence a string of different L1 fragments, ERVL-E-int fragments, MER3, MIR, AluJ, MLT1B, LTR84b, MLT1G3, AluSx, and MLT2C1 beside the MER5 and LTR33 elements (Supplementary Figure 1). There were numerous other indels ranging between 1 and 40 kb within the beta block sequences (Supplementary Figures 24) that were not included as allelic markers in this study. To assess the MHC class I haplotypic integrity of the 95 cell lines, the additional allelic haplotype combinations that we typed were sorted and grouped according to alpha block haplotypes (Table 2, Supplementary Table 5) and beta block haplotypes (Table 3, Supplementary Table 6) and then used for SNP XO studies across ~3 Mb of sequence between GPX5 and MICB (Tables 48).

Table 2

Hap IDNo. HapAluHFHLA-FHLA-GAluHGERVK9HLA-HHLA-AHLA-JAluHJ
15101:01:01:0901:01:021102:0101:01:0101:01:01:022
29101:01:01:0901:061102:0101:01:0101:01:01:022
313101:01:01:0101:01:012101:0102:01:0101:01:01:051
41101:01:01:0101:01:012101:0102:01:0101:01:01:041
51101:01:01:0101:01:012101:0102:01:0101:01:01:051
61101:01:01:0401:01:012101:0102:01:0101:01:01:051
731/201:01:01:0801:01:012101:0102:01:0101:01:01:051
821/201:01:01:0801:01:012101:0102:01:0101:01:01:021
91101:01:01:0901:01:012101:0102:01:0101:01:01:051
101101:04:01:0201:01:012101:0102:01:0101:01:01:051
111101:01:02:0701:03:0111New02:05:01New1
121101:01:01:0901:01:012101:0102:01:0101:01:01:051
131101:04:01:0201:01:012101:0102:01:0101:01:01:051
141101:01:02:0701:03:0111New02:05:01New1
151101:01:01:18/1901:01:221102:0403:01:0101:01:01:022
166101:03:01:01/0401:01:011202:0403:01:0101:01:01:041
172101:01:02:09/1201:01:0312New11:01:0101:01:01:041
181101:03:01:0301:04:0412Deletion23:01:0101:01:01:041
191101:01:01:0801:04:0112Deletion24:02:0101:01:01:022
203101:01:01:0901:04:0112Deletion24:02:0101:01:01:022
211101:01:01:0901:04:0112Deletion24:02:0101:01:01:041
221101:01:02:1001:04:0112Deletion24:02:0101:01:01:022
233101:01:01:18/1901:01:0212Deletion24:02:0101:01:01:022
2431/201:01:01:0801:01:021101:0226:01:0101:01:01:081
254201:01:01:0801:01:011102:0229:02:0101:01:01:011
261101:01:01:0901:01:0211New30:01/A68New1
301101:01:01:0901:05N11New30:01:01New1
312101:01:01:01/1701:01:0121New30:02:0101:01:01:041
321101:01:0101:03:012101:0131:01:0201:01:01:051
333101:01:01:1101:03:0112New31:01:02New1
342101:01:02:06/1001:01:221102:0332:01:0101:01:01:062
351101:01:02:0601:01:121102:0332:01:0101:01:01:062
362101:01:02:0701:03:0112New33:01:01New1
371101:01:02:1001:04:0112New33:01:01New1
381101:01:01:0801:01:021101:0266:01:01New1
391201:01:01:18/1901:01:0211New68:02:01New1

Alpha block haplotypes and alleles from AluHF to AluHJ including HLA-F, -G, -H, -A, and -J alleles.

For AluHF, AluHG, AluHJ, and ERVK9, allele 1 is the absence of the element and allele 2 is the presence of the element. This nine-marker table is a summary of the more detailed Supplementary Table 4 with 21 markers.

Table 3

Hap IDNo. HapsSVA-HCHLA-CSVA-BC9.5 kb indelSVA-HBHLA-BSVA-MICMICAMICBCEH/AH
1201:02:0122146:01:012010:01005:0246
21101:02:0122151:01:012010:01005:0251
31101:02:0122154:01:011012:01005:0254
41101:02:0122156:01:011012:01005:0256
51101:02:0122115:01:01:012010:0100615
61103:04:01:0112140:01:022008:04002:0140
75105:01:01:0112118:01:01:011/2001005:0218.2
85105:01:01:0212144:02:01:011008:01005:0244.1
92107:01:01:011/22118:01:01:021/2018:01002:0118.
101107:01:01:0122149:01:011004005:0249.x
111107:01:01:0122157:01:01201700357.1
128207:02:01:0312107:02:012008:04004:017.1
131108:02:01:0112114:01:012019:01005:0214.x
142108:02:01:0112114:02:011011005:0214.y
151112:02:0212052:01:012009:01002:0152.1
161114:02:0122151:01:012049005:0251
171114:0322144:03:012004005:0244
181115:02:0112151:01:012009:01002:0151.y
192115:02:0112151:01:012009:01005:0251.x
202101:02:01122*27:05:021007:01005:0227.1
211102:02:02:0112227:05:021007:01005:0227.x
222102:02:02:0112240:02:012027005:0240.x
231102:02:02:0112240:02:01202701340.y
242103:03:0112215:01:01:011/2010:01002:0115.x
251103:03:0112215:01:01:012010:01005:0215.y
263103:04:01:01122*15:01:01:012010:01002:0162.1
271103:04:01:0112240:01:022008:04002:0160.x
281103:04:01:0112240:01:022008:04004:0160.y
291103:04:01:0112240:01:022008:0401460.z
3011C*04:01:01:01122*15:26N2010:01005:0215.n
3111C*04:01:01:0112235:01:01:010002:01005:0235.x
3211C*04:01:01:0110235:01:01:01201700335.y
3311C*04:01:01:01022*35:01:01:021002:01002:0135.2
3421C*04:01:01:0112235:02:012016005:0135.z
3511C*04:01:01:0112235:03:011002:01005:0235.w
3611C*04:01:01:01122*35:08:012016002:0135.v
3711C*04:01:01:01122*53:01:011002:0100653.x
383106:02:01:0110213:02:012008:01005:0213.1
391106:02:01:01102*37:01:012010:01002:0137.x
401106:02:01:0110240:01:022008:04004:0140.x
411106:02:01:0110247:01:01:2008:01004:0147.1
424106:02:01:0110257:01:01201700357.1
431106:02:01:0210250:01:011009:02005:0650.1
445107:01:01:01122/2*08:01:011008:010088.1
451107:01:01:0112208:01:012008:04004:018.x
461107:18122*58:01:011002:0100858.1
473112:02:02122/2*52:01:012009:01005:0352.1
481112:03:01:0110235:03:011002:01005:0235.2
493112:03:01:01102/2*38:01:011002:01002:0138.x
501112:03:01:01102*51:01:012006005:0251.x
513116:01:0112244:03:011004005:0244.2
521116:01:0112245:01:012015002:0145.x
531117:01:01:0212241:01:011004005:0241.x
541117:01:01:0212242:01:011004002:0142.1

Beta block haplotypes.

SVA-HB allele 2* is a SVA-HB duplicated sequence within LTR10/HERVI/LTR10 rearrangements that do not correlate with any particular HLA-C lineage and therefore may be sequence assembly errors. For SVA, Alu, and the indel, allele 1 is the absence of the element and allele 2 is the presence of the element.

Table 4

Lab IDCEHHaplotypeSP regionSNPs/100 kbSequence length kb
47.1A*03:01/C*07:02/B*07:022,939
67.1A*03:01/C*07:02/B*07:02SP across whole region0.8161,962
517.1A*03:01/C*07:02/B*07:02SP across whole region0.7492,938
757.1A*03:01/C*07:02/B*07:02SP across whole region1.1242,937
907.1A*03:01/C*07:02/B*07:02SP across whole region1.0542,942
278.1A*01:01/C*07:01/B*08:012,996
118.1A*01:01/C*07:01/ B*08:01SP across whole region0.7162,933
128.1A*01:01/C*07:01/B*08:01SP across whole region0.863,023
168.1A*01:01/C*07:01/B*08:01SP across whole region1.6622,948
198.1A*01:01/C*07:01/B*08:01SP across whole region1.1812,963
2518.2A*30:02/C*05:01/B*18:012,988
2618.2A*30:02/C*05:01/B*18:01SP across whole region0.682,940
6751.xA*02:04/C*15:02/B*51:012,952
7651.xA*02:04/C*15:02/B*51:01SP across whole region0.4432,937
3757.1A*02:01/C*06:02/B*57:012,984
5857.1A*02:01/C*06:02/B*57:01SP across whole region0.5462,932
1762.xA*02:01/C*03:03/B*15:012,944
3262.xA*02:01/C*03:03/B*15:01SP across whole region0.242,921
4062.1A*02:01/C*03:04/B*15:012,941
8562.1A*02:01/C*03:04/B*15:01SP across whole region0.7692,990
4162.1A*02:01/C*03:04/B*15:01SR/SP at LINC00243SR & SP2,922
4965.1A*33:01/C*08:02/B*14:022,940
8765.1A*33:01/C*08:02/B*14:02SP from MASIF to MICA9.11,923
7844.2A*29:02/C*16:01/B*44:032,920
7944.2A*29:02/C*16:01/B*44:03SP from MASIF to MICB0.8471,772
8344.2A*29:02/C*16:01/B*44:03SP from HLA-F to MICBSR & SP2,975
2335.5A*01:01/C*04:01/B*35:022,938
4535.5A*01:01/C*04:01/B*35:02SP from HLA-F to MICBSR & SP2,938
3927.1A*02:01/C*01:02/B*27:05:022,943
4727.1A*02:01/C*01:02/B*27:05:02SP from MUC21 to MICBSR & SP2,944
2444.1A*02:01/C*05:01/B*44:022,921
6044.1A*02:01/C*05:01/B*44:02SP from HLA-C to MICBSR & SP1,792
7444.1A*02:01/C*05:01/B*44:02SP from HLA-F to MICBSR & SP2,937
3018.xA*02:01/C*07:01/B*18:011,912
3318.xA*02:01/C*07:01/B*18:01SP from HLA-E to MICBSR & SP2,929
6252.1A*24:02/C*12:02/B*52:012,893
9352.1A*24:02/C*12:02/B*52:01SP from HLA-L to MICBSR & SP2,887
1360.1A*02:01/C*03:04/B*40:01:022,749
8660.1A*02:01/C*03:04/B*40:01:024 SNP crossover regionsSR & SP2,947
944.xA*32:01/C*05:01/B*44:022,937
7244.xA*32:01/C*05:01/B*44:024 SNP crossover regionsSR & SP2,976

SNP-poor (SP) and SNP-rich (SR) haplotypes in the MHC class I region from GPX5 to MICB.

For details of sequence alignments between Lab ID and CEH, see Supplementary Table 8.

Allelic Lineages Within Alpha Block Haplotypes

Supplementary Table 5 shows the 46 alpha block haplotypes and HLA-A and RE allelic lineages of 95 homozygous cell lines (Norman et al., 2017) and the MHC sequence on chromosome 6 of the reference human genome (NC_000006.12, NCBI) using 21 genic and non-genic allelic markers from the telomeric locus of AluHF to the centromeric locus of AluHJ including HLA-F, HLA-G, eight HLA pseudogenes and 14 HLA-A allelic lineages. Of the HLA-A allelic lineages, only seven represented more than three sequence samples: HLA-A*01 (n, 14), -A*02 (n, 29), -A*03 (n, 8), -A*24 (n, 9), -A*29 (n, 4), -A*30 (n, 4), and -A*31 (n, 4). All seven HLA-A haplotype lineages were differentiated by haplotypic and/or haplospecific markers: most of the Alu, SVA, HERVK9, and MER9 within the alpha block (Table 1) were haplotypic and linked to particular HLA-A allelic lineages as well as to those of HLA-F, HLA-G and the HLA pseudogenes (HLA-V, -P, -H, -T, -K, -U, -W, and -J). Table 2 presents a summary of Supplementary Table 5 and shows the linkages of AluHF, AluHG, AluHJ, and HERVK9 with the HLA-F, -G, -H, -A and -J alleles in 39 alpha block haplotypes.

(1) All 14 HLA-A*01:01:01:01 alleles were linked to the haplospecific AluHJ insertion, HLA-J*01:01:01:02, HLA-H*02:01:01:01, HLA-F*01:01:01:09, and the ERV3-16/(ATAA)42/(ATTT)34 microsatellite.

(2) Twenty-seven of 28 HLA-A*02 haplotype lineages were linked to the haplotypic AluHG insertion, the (CAGAGA)n microsatellite deletion, the ERV3-16/(ATAA)46/(ATTT)35 microsatellite, HLA-G*01:01:01:01 and HLA- H*01:01:01:01.

(3) A single sequence sample with the HLA-A allele A*02:05:01 had no AluHG insertion, but had a variant (CAGAGA)n microsatellite number, and different alleles for all the alpha block HLA pseudogenes except for HLA-U*01:03.

(4) The AluHG insertion linked to the (CAGAGA)n microsatellite deletion was haplospecific for HLA-A*02/ G*01:01:01:01/H*01:01:01:01, whereas the AluHJ insertion was haplotypic for HLA-A*01/ G*01:01:02:01 or HLA-A*24/ G*01:01:02:01/G*01:04:01:01, respectively.

(5) The AluHG insertion with the (CAGAGA)n microsatellite deletion was linked also to HLA-A*30 in one haplotype and to HLA*A31 with the HERVK9 deletion in another haplotype, but not to the other three with the HERVK9 insertion, probably as a result of past recombinations or conversions.

(6) The AluHJ insertion in the HLA-A*01, HLA*24 haplotypes and the occasional HLA-A*02 or HLA-A*03 haplotypes was linked to all of the J*01:01:01:02 alleles (25 samples) and to J*01:01:01:06 in three samples of the A*32:01:01 haplotype.

(7) The AluHF was linked to 8 of 14 HLA-F*01:01:01:08 alleles, in 2 of 29 HLA-A*02 haplotypes, 2 of 3 HLA-A*26 haplotypes, all 4 HLA-A*29 haplotypes and 1 HLA-A*68:02 haplotype.

(8) The HERVK9 insertion was present in 25 cell lines, whereas the other 70 cell lines had the signatory deletion marker, a solitary MER9 that is the deletion product of a recombination between the 5MER9 and 3MER9 flanking the 6-kb HERVK9 internal sequence.

(9) The HERVK9 insertion was haplotypic for seven of eight HLA-A*03/G*01:01:01:05/H*02:04 samples, all nine HLA-A*24 samples, three of four HLA-A*31, both HLA-A*11 and HLA-A*33 samples, and the single HLA-A*23 sample.

(10) Both HLA-A*11 sequence lineages from the cell lines WT100BIS (Lab ID1) and KGU (Lab ID21) were linked to the HERVK9 insertion, and to C*04:01 and B*35:03:01 as extended haplotypes.

(11) All nine HLA-A*24:02:01:01 lineages and the single A*23:01:01 lineage had a ~55-kb deletion of HLA-H, SVA-16, HLA-T, HLA-K, LTR13A, HLA-U, and sMER9, ranging from centromeric of the HERVK9 in the HLA-H segment to the Charlie9 element at the telomeric end of the MER9 sequence of the HLA-A segment (Supplementary Figure 5).

(12) Eight of the nine HLA-A*24 haplotypes acquired HLA-J*01:01:01:02 with the AluHJ insertion, while the other acquired J*01:01:01:04 without the AluHJ insertion that is similar to the two HLA-A*11 haplotypes and the one HLA-A*23 haplotype.

Allelic Lineages Within Beta Block Haplotypes

Supplementary Table 6 shows the genic and non-genic allelic markers for the beta block haplotype sequences of 95 homozygous cell lines from the telomeric locus of HERVK9/MER9 microsatellite (TTTC)n known as M13 (Kulski et al., 1997) to the centromeric locus of MICB including five other microsatellite loci (M11, M9, Msx, MSa, and Msb), six dimorphic indels (SVA-HC, SVA-BC, 9.5-kb indel, SVA-HC, and AluP5) and four SNP loci (HLA-C and -B alleles and MICA and MICB alleles). There are 12 HLA-C, 21 HLA-B, 16 MICA, and 7 MICB allelic lineages that are linked together to form at least 54 HLA-C/HLA-B/MIC haplotype lineages. These haplotype lineages were sorted in the sequential order for the absence (allele 1) and presence (allele 2) of the SVA-HB insertion, and the alleles of HLA-C, HLA-B, MICA, and MICB, respectively (Table 3). The SVA-HB insertion (allele 2) is missing from the chimpanzee and gorilla MHC (data not shown), and its absence is assumed to be the ancestral allele.

Fifty-six of the 95 sequenced cell lines had the SVA-HB insertion. The HLA-C haplotypic lineages with no SVA-HB insertion were 6 C*01:02:01, 1 C*03:04:01 linked to HLA-B*40:01:02, 10 C*05, 2 C*07:01, 8 C*07:02, 3 C*08, 1 C*12:02 linked to HLA-B*52, 2 C*14 and 3 C*15. The HLA-C lineages with the SVA-HB insertion were 2 C*01:02:01 linked to HLA-B*27, 3 C*02, 9 C*03, 9 C*09, 11 C*06, 6 C*07:01 and 1 C*07:18, 9 C*12, 4 C*16, and 2 C*17. Only C*01, C*03 and C*07 had crossed over to be represented by both the absence and presence of the SVA-HB. Eighteen of 56 SVA-HB positive sequences contained SVA-HB duplications and LTR10/HERVI/LTR10 rearrangements that did not correlate with any particular HLA-C lineage or haplotype, suggesting that these variants were likely sequencing assembly errors. Nevertheless, there are two distinct haplotype evolutionary histories for the beta block that are based on the absence or presence of the SVA-HB insertion.

The SVA-HC insertion was specific for the eight HLA-C*07:02:01/B*07:02:01/MICA*008:04/ MICB*008:04 haplotypes, whereas the SVA-BC insertion was linked to six C*01:02 samples with various HLA-B alleles, three of four C*07:01 alleles and both C*14 alleles (Table 3, Supplementary Table 1). The SVA-BC and the SVA-HC insertions were present only in samples without the SVA-HB insertion. In comparison, the SVA-MIC insertion was linked to various HLA-B alleles both with and without linkage to the SVA-HB insertion. The (ACACAT)101 and the (ACACAT)161 simple repeats located between HLA-C and HLA-B further subdivide these HLA-B haplotypic lineages (data not shown). Three different haplotype families with (ACACAT)101 had no SVA-HB insertion, three HLA-B*14, seven HLA-B*18 and five of nine HLA-B*44 (Supplementary Table 6). The five lineage haplotypes with the microsatellite (ACACAT)161, but without the SVA-HB insertion, were B*07, B*46, B*51, B*54 and B*56. Single examples of B*15, B*40, B*44, B*52 and B*57 with (ACACAT)161 were either with or without SVA-HB (Supplementary Table 6). The different HLA-B lineage haplotypes with the SVA-HB insertion were partitioned further into another two lineages: those with the 9-kb deletion between AluY-(AT)n and AluJb-(TTAT)n and those without the 9-kb deletion (Table 3, Supplementary Figure 1). The 18 HLA-B haplotypes with the 9-kb deletion were all linked to either HLA-C*06:02:01 or -C*12:03:01 (Table 3).

Segmental Exchanges and SNP XO Within the MHC Class I Region

Supplementary Table 7 shows 39 examples of segmental shuffling between HLA class I genes A, B, C and E, pseudogene HLA-J and the MICA and MICB genes of 59 different representative AH and subtypes using HLA-B alleles as AH anchor points. Most of the MHC haplotypes within the homozygous cell lines are Caucasoids from Europe, North America, South Africa and Australia. The exceptions are one cell line from a North American Hispanic (MGAR), five Oriental cell lines (SA, ISH3, HOR, AKIBA, and KAWASAKI) and five South American Indian cell lines (LZL, AMALA, SPL, RML, and KRC005). The four RE dimorphic structural markers AluHG, AluHJ, and HERK9 within the alpha block and SVA-HB within the beta block further subdivided some of these AH. It is noteworthy that the Sardinian 18.2AH, HLA-A*30:C*05:B*18 (Contu et al., 1989, Bilbao et al., 2006) in the cell lines EJ32B and DUCAF has two specific dimorphic Alu insertions, AluOR and AluOR1 (Supplementary Table 2), located ~300 kb from the HLA-F gene (Figure 1). This finding confirms that the CPS of some MHC class I haplotypes and AH like the 18.2AH extend well into the OR gene cluster telomeric of the HLA-F gene and the MHC alpha block by at least 1,185 kb. The AluOR insertion was found also in one of two HLA-A*29 sequenced samples (cell lines PITOUT and MOU, respectively), the HLA-A*02:05:01 cell line WT49, the HLA-A*11:01 cell line WT100BIS and the HLA-A*23:01 cell line WT51 (Supplementary Table 2).

Figure 1

The four Warao South American Indian cell lines LZL, AMALA, SPL and RML (Supplementary Tables 1, 2) present an interesting ethnic contrast to the Caucasoid cell lines (Supplementary Table 7). The Warao people who inhabit the rainforests of Orinoco Delta of northeastern Venezuela and western Guyana are an ancient ethnic minority with an extant population of ~50,000 people. The Warao 62.xAH and 51.xAH have the HLA-A alleles A*2:17:02, A*02:04, and A*02:12 rather than the common Caucasoid A*02:01:01, but they also carry the AluHG insertion that is linked to most of the Caucasoid A*02 lineages (Table 2). One of the Warao cell lines (SPL, ID73) has the HLA-A*31:02:02 allele linked to the HERVK9 insertion and SVA-HB deletion, which is markedly different to the Caucasoid and Oriental 62.xAHs, but with an alpha block haplotypic structure similar to two Caucasoid HLA-A*31:02 lineages represented by the English cell line JHAF (ID46) and the Australian cell line MT14B (ID82) (Supplementary Tables 1, 2). The other HLA-A*31:02 haplotype represented by the European Caucasoid cell line DEU (ID35) is with an AluHG insertion, a HERVK9 deletion (Table 2) and an SVA-HB insertion (Table 3), suggesting that a more modern HLA-A*31 AH was subsequently generated by segmental shuffling exchanges.

Since the exact SNP XO regions between different MHC haplotypes are poorly defined in regard to the intergenic and genic distribution of repeat elements, a detailed comparative examination of DNA sequence alignments of similar and different haplotypes was undertaken using the PIP method of Schwartz et al. (2000). We started with an examination of ~3 Mb of genomic sequence of the same haplotypes that included the class I region from HLA-F to MICB and the OR gene cluster that included the GPX5, ZNF311, OR2H2, GABBR1, and MOG genes (Table 4, Supplementary Table 8). We then performed a more detailed examination of SNP densities and XOs within the 1.2-Mb OR gene region (Table 5), the 310-kb alpha block (Table 6), the 1,172-kb inter alpha and beta blocks (Supplementary Tables 9, 10) and the 307-kb beta block (Tables 7, 8) of the same and/or different HLA-A, -C and -B haplotypes. The alignments and SNP counts were analysed manually across the entire 3 Mb and in 50-kb to 500-kb segments connecting the various segments as a sliding window. Figure 1 summarises the findings of our analysis of more than 250 sequence alignments between different and the same haplotypes with the identification of at least 38 ancestral SNP XO sites between SNP-poor and SNP-rich regions within ~2.8 Mb from GPX5 to MICB.

Table 5

Haplotype sequence alignmentOR GENE CLUSTER REGION/GABBR1/MOGTotal SNPsXO point in haplotype sequence ID_HAP1Within or between (/) REXO between (/) genes
GPX5 to OR2H1MAS1L & LINC01015USB & OR2H2GABBR1MOG & ZFP57
AAABCDA-D
ID_HAP1ID-HAP2960 kbSNP/50 kbSNP/50 kbSNP/50 kbSNP/60 kbSNP/210 kb
49_A*33-C*08:0250_A*33-C*14:03SRR875617216 (XO)376208524 A/GERV3-16A3F segment
15_A*26-C*0564_A*26-C*12:03SRR394421214 (XO)318211345 A/GERV3-16A3F segment
15_A*26-C*0569_A*66-C*12:03SRR45341119 (XO)109210115 G/ALTR43/ERV3F segment
2_A*01-C*0611_A*01-C*07SRR10568262 (XO) 1202165747 T/CCharlie4/L3bZFP57/HLA-F
2_A*01-C*0616_A*01-C*07SRR9364212 (XO) 1181165747 T/CCharlie4/L3bZFP57/HLA-F
4_A*03-C*07:027_A*03-C*06:02SRR97434 (XO)3147121130 A/CL3b/MER5GABBR1/MOG
9_A*32-C*0522_A*32-C*12:03SRR905117(XO) 02160143183 T/CAluY/L2GABBR1/MOG
9_A*32-C*0572_A*32-C*05SRR914915 (XO) 02157143183 T/CAluY/L2GABBR1/MOG
10_A*02-C*12:0317_A*02:17-C*03SRR9330 (XO) 02513092214 G/AAluY(Sc8)OR2H2/GABBR1
10_A*02-C*12:0392_A*02:12-C*01SRR9130 (XO) 02312692214 G/AAluY(Sc8)OR2H2/GABBR1
5_A*02-C*0510_A*02-C*12:03SRR7732 (XO) 00211175758 T/CMER21CUSB/OR2H2
10_A*02-C*12:035_A*02-C*05SRR7732 (XO) 00211175792 C/TMER21CUSB/OR2H2
34_A*24-C*03:0468_A*24-C*04SRR9715 (XO) 01011359547 C/TLTR10AUSB/OR2H2
46_A*31-C*15:0273_A*31-C*01:02SRR589 (XO) 2267559532 T/CLTR10AUSB/OR2H2
15_A*26-C*0584_A*26-C*07SRR3720 (XO) 1205966738 A/TTigger2b-PriUSB/OR2H2
22_A*32-C*12:0372_A*32-C*05(XO)21407136748 T/CL2/L1OR12D3/OR12D2
4_A*03-C*07:026_A*03-C*07:026_missing seq4004UndetectedSPR from block AA to block D
78_A*29-C*1679_A*29-C*1679_missing seq0011UndetectedSPR from block AA to block D
11_A*01-C*0716_A*01-C*07700101UndetectedSPR from block AA to block D
4_A*03-C*07:0251_A*03-C*07:02800011undetectedSPR from block AA to block D
34_A*24-C*03:0494_A*24-C*03:04000213UndetectedSPR from block AA to block D
25_A*30-C*0526_A*30-C*05001001UndetectedSPR from block AA to block D
46_A*31-C*15:0282_A*31-C*03:04002147UndetectedSPR from block AA to block D
15_A*26-C*0578_A*29-C*16SRR4537 (XOa) 034 (XOb)8991066 G/AL2/MSOR2H2/GABBR1
2_A*01-C*0617_A*02:17-C*03SRR115848197314
25_A*30-C*0534_A*24-C*03:04SRR494085191365
25_A*30-C*0546_A*31-C*15:02SRR493527287398
25_A*30-C*0522_A*32-C*12:03SRR83421351189

SNP crossover (XO) loci in the extended MHC class I OR gene region.

SRR is SNP-rich region, SPR is SNP-poor region, XO is crossing over, and numbers in the block columns AA to A–D are SNP counts per block.

Table 6

Segment number123456789101 to 10
Segment nameOR endFVPGHTKAWJF to JPosition and XO SNPXO within or between (/) RE
Segment size2.2 kb57 kb25 kb23 kb53 kb23.4 kb20.4 kb19 kb15 kb51 kb34 kb320.8
Hap sequence comparisons
ID_Hap 1ID_Hap 2
11_A*01:014_A*03:0112284134198382340330245107180402,240
2_A*01:014_A*03:0111280144200388367374248110187382,336
10_A*02:014_A*03:01122623235133328321248343380302,112
2_A*01:0110_A*02:01934013519437110355164346380352,123
5_A*02:011_A*11:0115280353051735625980240211272,035
15_A*26:0178_A*2904 XO 30137182208692032321487272172,157[F] 56590 G/CCharlie20a
25_A*3046_A*31132061111522643552602973354911732,644
SNP densityaverage: SNPs/kb4.74.24.26.26.111.712.611.415.57.22.47.0
4_A*03:0146_A*3113SRRSRRSRRSRRSRRSRRSRRSRRSRRSRRSRR
4_A*03:0149_A*332279191166236651062683314761692287
46_A*3149_A*3315245196164 XO 0167911113315707[P] 97693 C/TMICG
4_A*03:0148_A*24/B*15:26N11255182750292/XO/deldeldeldel/XO/175171221262[H] 168684 delL1/AAAGA/MLT1F1
4_A*03:0194_A*241027913519444091/XO/deldeldeldel/XO/175219381571[H] 168684 delL1/AAAGA/MLT1F1
5_A*02:018_A*02:051126810615637127 XO149Assembly errors28952[H] 168492 G/CL2/HLA-H/L2
10_A*02:018_A*02:0511SRRSRRSRRSRRSRR XO5713XO 19924SRR/SPR[H] 168959 G/CL2/HLA-H/L2
10_A*02:018_A*02:0513XO 19924SPR/SRR[W] 224700 C/TERV3-16A3
34_A*2448_A*24973139200229/XO/01/XO/deldeldeldel/XO/28239763[G] 141380 A/GHAL1/MICF
34_A*2459_A*242132141199229/XO/10/XO/deldeldeldel/XO/251710[G] 141380 A/GHAL1/MICF
48_A*2468_A*24986132199331/XO/10/XO/deldeldeldel/XO8941879[G] 144028 G/AHAL1/MICF
48_A*2494_A*24988132199334/XO/10/XO/deldeldeldel/XO8842884[G] 144028 G/AHAL1/MICF
59_A*2494_A*242132136197284/XO/13/XO/deldeldeldel/XO41760[G] 141500 G/AHAL1/MICF
10_A*02:0113_A*02:0110171+XO003211000178[F] 53540 G/TCharlie20a
10_A*02:0167_A*02:0410155+XO112200100152[F] 53540 G/TCharlie20a
10_A*02:0176_A*02:0410155+XO112200100152[F] 53540 G/TCharlie20a
5_A*02:0113_A*02:0111166+XO006133Assembly errors37[F] 53275 G/TCharlie20a
4_A*03:0118_A*03/A*24513778126179+XO0027/XO13137692[G} 150337 C/TTigger1/Charlie20a
[A] 229198 G/TL2/HLA-A/L2
48_A*2459_A*2411139 XO 01000/XO/deldeldeldel/XO8742269[F} 53185 T/CTigger1/Charlie20a
34_A*2468_A*2400011XO/deldeldeldel/XO102[H] del [A]
34_A*2494_A*2400011XO/deldeldeldel/XO001[H] del [A]
10_A*02:0132_A*02:17Seq missing0141003009OR cluster
5_A*02:0110_A*02:0113005322Assembly errors318OR cluster
11_A*01:0116_A*01:01000000000121.5 kb del1OR cluster
2_A*01:0116_A*01:010110133332321.5 kb del47OR cluster
2_A*01:0111_A*01:010110133313218OR cluster
10_A*02:0117_A*02:17000121003007OR cluster
10_A*02:0192_A*02:12010010012106OR cluster
4_A*03:016_A*03:01010001000103OR cluster
4_A*03:017_A*03:01000100110216OR cluster
15_A*2664_A*260101230104113OR cluster
15_A*2684_A*26010031000207OR cluster
15_A*2669_A*66210311576139863OR cluster
78_A2979_A29000010000001OR cluster
25_A*3026_A*30000010000001OR cluster
46_A*3173_A*31111001000104OR cluster
46_A*3182_A*31100000000011OR cluster
9_A*3222_A*32000000000000OR cluster
9_A*3272_A*32001011210118OR cluster
49_A*3350_A*33000010000001OR cluster

SNP counts and crossovers (XO) within alpha block segments 1 to 10 between different haplotype sequence alignments (ID_Hap1 and ID_Hap 2).

*

The SRR to SPR XO in Seg G (4) at A/G 141380 HAL1 and (ATAAT)n is near the AluHG insertion locus and the MICF pseudogene. XO is the abbreviation for crossover; SRR, SNP-rich region; SPR, SNP-poor region. The numbers before and/or after XO are the number of counted SNPs before and/or after the observed XO. There are XO points outside the alpha block in the telomeric OR gene region and the centromeric non-HLA region between HLA-J and HLA-E. The bold values here show the SNP density and the average number of SNPs/kb for the top 7 haplotype sequence comparisons in the table.

Table 7

Alignments between haplotypesNumber of SNPs per sectionXOSNP orclosestXO nt distanceXO in or
Lab ID numbers precede haplotypes0–7 k7–10 kb HLA-C10–20 k20–40 k40–60 k60–80 k80–93 k HLA-B0–93 kLocation bpIndel at XORepeatTo end of HLA-B exon 8Between (/) HLA-C and -B
Haplotype 1Haplotype 2
(A) Different HLA-C/HLA-B haplotypes
66_C*02/B*27:0549_C*08:02/B*14:02333285114455805 + MI2691,811none0nd0SRR
55_C*06:02/B*13:0295_C*07/B*57158885112247 + MI7681851,500none0nd0SRR
(B) Same HLA-C/ HLA-B haplotypes
28_C*03:03/B*1517_C*03:03/B*150001022500nd0SPR
54_C*04/B*3523_C*04/B*350001100200nd0SPR
55_C*06:02/B*13:0291_C*06:02/B*130010000 + indels100nd0SPR
6_C*07/B*07:024_C*07/B*07:020000000000nd0SPR
11_C*07/B*0812_C*07/B*080000000000nd0SPR
65_C*08:02/B*14:0149_C*08:02/B*14:02000000777700nd0SPR
(C) Same HLA-C allele–different HLA-B allele
6_C*07/B*07:0230_C*07/B*18129 + XO6889SRRSRRSRRSRRHLA-C*70ndndHLA-C
30_C*07/B*186_C*07/B*07:02129 + XO6889SRRSRRSRRSRRHLA-C*70ndndHLA-C
6_C*07/B*07:028_C*07:18/B*58199 + XO122163SRRSRRSRRSRRHLA-C*70ndndHLA-C
8_C*07:18/B*586_C*07/B*07:02199 + XO122163SRRSRRSRRSRRHLA-C*70ndndHLA-C
94_C*12:03/B*5156_C*12:02/B*5241XO + 2476195SRR + MISRR41SRRHLA-C*120ndndHLA-C
94_C*12:03/B*5153__C*12:02/B*52412 + XO70217SRRSRR41SRRHLA-C*120ndndHLA-C
11_C*07/B*088_C*07:18/B*5816942+XO+216332277 + MI387103619961A/GL1PA1372,286HLA-C/HLA-B
11_C*07/B*0830_C*07/B*1815900+XO+291327 + MI348 + MI341126819961A/GL1PA1372,286HLA-C/HLA-B
11_C*07/B*0842_C*07/B*4915900+XO+289329 + MI348 + MI376130319961A/GL1PA1372,286HLA-C/HLA-B
11_C*07/B*0895_C*07/B*57100+XO+191327 + MI348 + MI373114119961A/GL1PA1372,286HLA-C/HLA-B
92_C*01/B*5139_C*01/B*27:05001XO+30367578 + 5k MI240121633326C/TMIR/L1MB855,903HLA-C/HLA-B
39_C*01/B*27:0592_C*01/B*51001XO+SRRSRRSRRSRRSRR32511C/TMIR/L1MB860,958HLA-C/HLA-B
39_C*01/B*27:0570_C*01/B*46212XO+SRRSRRSRRSRRSRR35360T/AAluY/MLT1D58,109HLA-C/HLA-B
39_C*01/B*27:0589_C*01/B*542602XO+SRRSRRSRRSRRSRR35360T/AAluY/MLT1D58,109HLA-C/HLA-B
39_C*01/B*27:0559_C*01/B*56202XO+SRRSRRSRRSRRSRR35360T/AAluY/MLT1D58,109HLA-C/HLA-B
39_C*01/B*27:0573_C*01/B*15112XO+SRRSRRSRRSRRSRR35360T/AAluY/MLT1D58,109HLA-C/HLA-B
11_C*07/B*084_C*07/B*07:0215876700 + XO +217SRR + MISRRSRR47841C/TMIR44,406HLA-C/HLA-B
11_C*07/B*086_C*07/B*07:0215876700 + XO + 217SRR + MISRRSRR47841C/TMIR44,406HLA-C/HLA-B
55_C*06:02/B*13:0231_C*06:02/B*4000000XO+52921073960351G/THERVI21,449HLA-C/HLA-B
5_C*05/B*189_C*05/B*44:02000100XO+878889798T/CMLT1N23,8273′HLA-B
95_C*07/B*5730_C*07/B*1816001012XO+14731987226A/CMLT1N2/MER52,7673′HLA-B
28_C*03:03/B*1513_C*03:04/B*40:01512253XO+18420488616indel 36bpMLT1N2/MER52,4783′HLA-B
28_C*03:03/B*1534_C*03:04/B*40:0152225MI + 3XO+18520588616indel 36bpMLT1N2/MER52,4783′HLA-B
54_C*04/B*3548_C*04/B*15:26N00MI + 01017XO+15217088190C/TMLT1N2/MER52,4403′HLA-B
54_C*04/B*3561_C*04/B*53001101XO+71088194T/CMLT1N2/MER52,4363′HLA-B
50_C*14:03/B*44:0388_C*14:02/B*510026125+XO+15817499804C/AMLT1N2/MER52,2733′HLA-B
30_C*07/B*1895_C*07/B*5716001012XO+14731993101A/CMLT1N2/MER52,2683′HLA-B
94_C*12:03/B*5122__C*12:03/B*38000000XO+18218280294G/AMLT1N2/MER52,2263′HLA-B
66_C*02:02/B*27:0538_C*02/B*40:02001000XO+333488896indel 5bpMLT1N2/MER52,1653′HLA-B
55_C*06:02/B*13:0277_C*06:02/B*47000010XO+11912080217G/CL21,5833′HLA-B
55_C*06:02/B*13:0236_C*06:02/B*57000002XO+18718981180G/AL26203′HLA-B
55_C*06:02/B*13:022_C*06:02/B*57000003XO+18719081180G/AL2/HLA-B6203′HLA-B
55_C*06:02/B*13:0243_C*06:02/B*37000003XO+19019381180G/AL2/HLA-B6203′HLA-B
92_C*01/B*5173_C*01/B*15001000XO+12512688668G/AL2/HLA-B5603′HLA-B
92_C*01/B*5159_C*01/B*56001001XO+24124288668G/AL2/HLA-B5603′HLA-B
92_C*01/B*5170_C*01/B*46001100XO+12712988668G/AL2/HLA-B5603′HLA-B
92_C*01/B*5189_C*01/B*541800000XO+16820788787indel 22bpL2/HLA-B4413′HLA-B
83_C*16/B*44:0380_C*16/B*45000202XO+13313791097indelL2/HLA-B5033′HLA-B
29_C*17/B*4114_C*17/B*42000102XO+20520886647G/AHLA-B exon 3−2,542HLA-B (ex 3)
(D) Different HLA-C allele/same HLA-B allele
92_C*01/B*5194_C*12/B*5117613279 + MI791 + MI190 + XO136889594G/AHLA-B exon 8−186HLA-B exon 8
5_C*05/B*1830_C*07/B*18SRRSRRSSR + MI (5.2kb)SRRSRR +XOSRR + MI93767G/AHLA-B exon 8−142HLA-B exon 8
9_C*05/B*440283_C*16/B*4403SRRSRRSRRSRR + MISRR +XOSRR + MI92080C/AL2/HLA-B1393′HLA-B
92_C*01/B*5146_C*15/B*51141295222149140 + XO94789065A/TL2/HLA-B1633′HLA-B
9_C*05/B*440250_C*14/B*4403SRRSRRSRR + MISRR + MISRR +XOSRR + MI91971T/CL2/HLA-B2483′HLA-B
73_C*01/B*1528_C*03/B*15SRRSRRSRRSRR + MISRR + XOSRR + MI87246A/GL2/HLA-B6403′HLA-B
73_C*01/B*1548_C*04/B*1525NSRRSRRSRRSRR + MISRR + XOSRR + MI87246A/GL2/HLA-B6403′HLA-B
2_C*06/B*5795_C*07/B*57SRRSRRSRR + MISRRSRR +XOSRR + MI83390A/GL2/HLA-B6503′HLA-B
54_C*04/B*3520_C*12/B*35SRR + MISRR + MISRR + MISRRSRR +XOSRR + MI89662C/AL29683′HLA-B
39_C*01/B*27:0566_C*02/B*27:0517579+XO+000025435005T/AL1MB8/MLT1D58523HLA-C/HLA-B
38_C*02/B*40:0234_C*03/B*40:02SRRSRRSRRSRR + MISRRSRRnd0ndndSRR
38_C*02/B*40:0277_C*06/B*40:01SRRSRRSRR + MISRRSRRSRRnd0ndndSRR

SNP counts and crossover (XO) loci between linked HLA-C and HLA-B alleles using different combinations of haplotype pairs.

SRR is SNP-rich region that is estimated to be >100 SNP without manual counts, SPR is SNP-poor region (<10 SNP), and MI is major indel (>1 kb usually <6 kb). XO in columns is crossover and XO + number presents the number of SNPs before or after the crossover in each of the 20-kb genomic sections. nd, not determined.

Table 8

Alignments between paired haplotype sequencesXO distance from HLA-BXOXO within gene or within and between (/) repeat elements
Lab ID numbers precede haplotypesSNP
Haplotype 1Haplotype 2SPR/SRR
65_B*14:01/MICA*019/MICB*50249_B*14:02/MICA*011/MICB*5028796T/CL1
65_B*14:01/MICA*019/MICB*50287_B*14:01/MICA*11/MICB*005:028822G/AL1
72_B*44/MICA*008/MICB*00583_B*44:03/MICA*004/MICB*00514436A/GL1M5/L1ME3
72_B*44/MICA*008/MICB*00550_B*44:03/MICA*004/MICB*00514436A/GL1M5/L1ME3
25_B*18/MICA*001MICB*500230_B*18/MICA*018/MICB*20116711G/TL1
19_B*08/MICA*008/MICB*00884_B*08/MICA*008/MICB*00430243C/TMLT2C1/Charlie9
94_B*51/MICA*006/MICB*00592_B*51/MICA*010/MICB*00544271A/GLTR8A
94_B*51/MICA*006/MICB*00546_B*51/MICA*009/MICB*00244618C/TLTR8A/AluJb
94_B*51/MICA*006/MICB*00567_B*51/MICA*009/MICB*00552042G/A(CTC)n/L1M3
88_B*51/MICA*049/MICB*00592_B*51/MICA*010/MICB*00554075A/GLTR8A
54_B*35/MICA*002/MICB*00568_B*35/MICA*016/MICB*00266656/MICAC/TL1MB2/MIR in MICA
46_B*51/MICA*009/MICB*00267_B*51/MICA*009/MICB*00595787C/TL1PA3/Tigger3b
86_B*40/MICA*008/MICB*00282_B*40/MICA*008/MICB*00496012C/TL1PA3/Tigger3b
88_B*51/MICA*049/MICB*00594_B*51/MICA*006/MICB*005100093G/ATHE1D/ L1M2
32_B*15/MICA*010/MICB*00228_B*15/MICA*010/MICB*005101267A/CMER2
32_B*15/MICA*010/MICB*00273_B*15/MICA*010/MICB*006101267A/CMER2
86_B*40/MICA*008/MICB*00213_B*40/MICA*008/MICB*014102370C/TSVA-MIC indel
21_B*35/MICA*002/MICB*0051_B*35/MICA*002/MICB*002105577G/CMER21C/MER4
54_B*35/MICA*002/MICB*00521_B*35/MICA*002/MICB*005113437C/GMER21C/MER4
68_B*35/MICA*016/MICB*00223_B*35/MICA*16/MICB*005116829G/A5′-ERV3-16A3_I (HCP5)
54_B*35/MICA*002/MICB*0051_B*35/MICA*002/MICB*002119556A/G5′-ERV3-16A3_I
88_B*51/MICA*049/MICB*00546_B*51/MICA*009/MICB*002123534G/C5′-ERV3-16A3_I
88_B*51/MICA*049/MICB*00567_B*51/MICA*009/MICB*005123570C/T5′-ERV3-16A3_I
38_B*40/MICA*027/MICB*00571_B*40/MICA*027/MICB*013124806A/GTHE1A/LTR33
28_B*15/MICA*010/MICB*00573_B*15/MICA*010/MICB*006undetectedSPRSPR HLA-B to MICB gene
19_B*08/MICA*008/MICB*00827_B*08/MICA*008/MICB*008no XOSPRSPR from HLA-B to MICB
65_B*1401/MICA*019/MICB*50249_B*1402/MICA*011/MICB*502SRR/112318/SPRA/G5′-ERV3-16A3 (HCP5)
65_B*1401/MICA*019/MICB*50287_B*14:01/MICA*11/MICB*005:02SRR/146518/SPRC/TL1
38_B*40/MICA*027/MICB*00586_B*40/MICA*008/MICB*002undetectedSRRSRR from HLA-B to MICB
38_B*40/MICA*027/MICB*00582_B*40/MICA*008/MICB*004undetectedSRRSRR from HLA-B to MICB
54_B*35/MICA*002/MICB*00535_B*35/MICA*017/MICB*003undetectedSRRSRR from HLA-B to MICB

SNP crossover (XO) loci within intergenic regions between HLA-B and MICA or MICB in alignments of different haplotype pairs.

SRR is SNP-rich region, SPR is SNP-poor region, and XO is crossover.

SNP Densities Within MHC Class I Homologous Haplotypes

We grouped and aligned 41 sequences to evaluate the variations of SNP density and the degree of homology within the same CEHs/AHs (Table 4 and Supplementary Table 8). The homologous sequence alignments for 12 of 16 different CEHs/AHs revealed a scarcity of SNPs ranging over ~1.8 Mb from HLA-F to MICB with <150 SNPs over the entire region at an average of 20 SNPs for 17 sequence alignments. Seven of the 16 different CEHs/AHs were SNP poor (<150 SNPs over ~3 Mb) from the GPX5 locus in the OR gene region to the MICB gene in the beta block region. These highly homologous sequence runs represented the seven sequences of 8.1AH, five sequences of 7.1AH, two sequences each of 18.2AH, 51.1AH, 57.xAH and 62.xAH, and two of three sequences classified as 62.1AHs. The SNP counts over the same range for the alignment of different haplotypes such as between the 7.1 CEH/AH and 8.1 CEH/AH were >2,000 for ~3 Mb.

Six CEHs/AHs had regions of substantial diversity that were SNP-rich between the alpha and beta blocks and/or in the OR gene region at the telomeric end of the alpha block. These haplotypes consisted of different-sized, SNP-poor recombinant blocks interspersed between SNP-rich recombinant blocks. The most surprising results were for the comparison between the three sequences of the 62.1 CEH/AH and the three sequences of the 44.1 CEH/AH. The sequence of LAB ID_41 had varying regions of SNP density with three SNP XOs, whereas ID_40 and ID_85 had few detectable SNPs (0.77 SNPs per 100 kb) and no XO SNPs in their alignment from GPX5 to MICB. Similarly, the ID_60 sequence with the 44.1 CEH/AH had at least two SNP XO events, one in the region between HLA-J and HLA-E and another in the region between MUC21 and HLA-C. In comparison, the ID_74 sequence of 44.1 CEH/AH had few SNPs and no detectable SNP XO in the 1.8-Mb sequence block from HLA-F to MICB.

Crossing Over Within the OR Extended Gene Region

The SNP XOs for some CEHs/AHs were detected in the OR gene regions hundreds of kilobases telomeric of the HLA-F gene (Table 5). For the sequence comparison between haplotype pairs, the OR genomic region was divided into four segmental blocks of 210–300 kb each, ranging from the telomeric GPX5 gene to the ZFP57 gene that are 1,118.4 kb and 42.1 kb telomeric of the HLA-F locus, respectively. The average SNP/210 kb for four different haplotype pairs was 317 SNPs within the genomic region between MAS1L and the start of the alpha block F segment. In paired sequence comparisons of 23 similar haplotypes, a SNP XO was found in 15 pairs within 237 kb between MAS1L and HLA-F. A SNP XO was found near to or within ERV3-16A3, an ancient HERV-16 element at the junction of the F segment in five haplotype pairs, within LTR10A of two haplotype pairs, within MER21C of two haplotype pairs, within AluY/Sc8 of two haplotype pairs and within Tigger2b of one haplotype pair. Most SNP XO were found in loci between ZFP57 and HLA-F (five haplotype pairs), USB and OR2H2 (five haplotype pairs), GABBR1 and MOG (three haplotype pairs) and OR2H2 and GABBR1 (two haplotype pairs), revealing the variability of the SNP XO junctions that were involved with ancestral recombinations.

The SNP XO in a region between the OR12D3 and ORD12D2 genes in the sequence alignment of 22_A*32-B*38 and 72_A*32-B*44:02 is ~332.2 kb from the HLA-F gene. Moreover, this SNP XO site is in close proximity to the young Alu indel, AluOR, that was detected in Japanese and Caucasians at a frequency of 0.32 and 0.14, respectively (Table 1). This SNP XO and active Alu insertion site appear to mark a hotspot for meiotic and insertion recombinations. Of the seven other sequence comparisons, no SNP XO was detected in two pairs (ID4 v ID6 and ID78 v ID79) of sequences that ended at the MA1L/LINC01015 segmental block and in five pairs (ID11 v ID16, ID4 v ID51, ID34 v ID94, ID25 v ID26, and ID46 v ID82) that ended near the GPX5 locus because of the absence of sequence for further analysis. The comparisons between two 7.1AH (ID4 v ID51) and two 8.1AH (ID11 v ID16) were striking because the SNP-poor region extended from the alpha block to at least the GPX5 gene that is ~1118 kb from the HLA-F gene. Both of the extended A*30-B*18 haplotypes (ID25 v ID26) carried the haplospecific AluOR insertion and the novel AluOR1 insertion (Table 1).

SNP Density XOs Within the Alpha Block of Different HLA-A Haplotypes

The alpha block was divided into 10 segments containing the duplicated HLA genes and pseudogenes from segment F with the HLA-F gene to segment J with the HLA-J pseudogene (Dawkins et al., 1999; Kulski et al., 1999a,b) for SNP counts and XO analysis (Table 6) as shown in Supplementary Figure 6 and Supplementary Table 3. The SNP counts were zero to less than 20 SNPs over 320.3 kb of sequence for 20 of 35 similar alpha block haplotype pairs (Table 6). In contrast, the SNP counts were much greater in the sequence alignments of seven different haplotype pairs ranging between 2,035 and 2,644 SNPs per 320.3 kb at an average of 7 SNPs/kb. The highest average SNP density was 16 SNPs/kb in the A segment and the lowest density was 2 SNPs/kb in the J segment. Because of deletions or XO within the alpha block, some recombinant haplotypes had intermediate SNP numbers such as 692 to 884 SNPs per 320.8 kb for six of the HLA-A*24 haplotypes, 269 SNPs for the 48_A*24 vs. 59_A*24 haplotypes and 707 SNPs for the 46_A*31 vs. 49_A*33 haplotypes. The smallest amount of SNP diversity within the alpha blocks of different HLA-A allelic lineages was between the HLA-A*26 and the HLA-A*66 haplotypes with only 66 SNP differences across the 320.8 kb alpha block (one SNP per 4.9 kb compared to an average of one SNP per 0.14 kb for an average of seven haplotype pairs). The biggest SNP difference among the different haplotypes was in the 53-kb G-segment and the smallest was within the 15-kb A segment that had only a total of three SNP differences. A previous analysis of the HLA-A*26/A*66 loci indicated that the HLA-A*66 was a product of a gene conversion (Madrigal et al., 1993). Alternatively, the relatively small number of SNP differences across the entire alpha block of the HLA-A*26/A*66 haplotypes suggests that they might have evolved from the same AH and diverged slightly by gene conversions and mutagenesis over time because of their age.

An intermediate amount of SNP diversity was detected between HLA-A*31 and HLA-A*33 with 707 SNPs within the 320.8-kb alpha block. However, the first 100 kb of the alpha block including the F, V and P duplicated segments was 605 SNPs and the remaining 220 kb of the alpha block from the G segment to the J segment was only 102 SNPs including 11 SNPs in the A segment. This segmental division with sequence homogeneity at the centromeric end of the HLA-A*31 and HLA-A*33 haplotype sequences and large diversity at the telomeric end of their alpha blocks has an ancestral SNP XO at the centromeric end of the P segment within the MICG pseudogene (C/T). In comparison to the 707 SNPs across the alpha block of the HLA-A*31 and -A*33 haplotypes, there were 2,287 SNPs across the alpha block of the HLA-A*03 and -A*33 haplotypes. Coincidently, the HLA-A*03, -A*31, and -A*33 haplotypes all have an HERVK9 insertion. Thus, two thirds of the HLA-A*31/*33 alpha blocks have the same haplotype lineage whereas the other third are evidently from different haplotype lineages.

The largest difference among different HLA-A haplotypes was between HLA-A*30 and HLA-A*31 with 2,644 SNPs in the alpha block, whereas on average there were 2,235 SNPs in the alpha block for seven different sequence pairs. The biggest differences between the same HLA-A haplotypes were obtained for the HLA-A*24 pairs, suggesting that their alpha blocks may have undergone numerous shuffling and exchanges with various other HLA-A haplotypes. No SNPs were detected in the alpha blocks of the HLA-A*32 haplotype pair and only one SNP was detected in the alpha blocks of an HLA-A*24 haplotype pair and the HLA-A*29, -A*30, and -A*31 haplotype pairs.

The SNP XOs for the same haplotypes at the telomeric end of the alpha block were variable depending on which haplotypes were compared (Table 6), but mostly involved the F segment with the HLA-F gene and sites within or between the Tigger1 and Charlie20 DNA elements. In genomic sequence comparisons between seven similar HLA-A*24 haplotype pairs, all of them had the 54-kb deletion of the T and K segments between the H and A segments (Supplementary Figure 5). Also, some XOs occurred in the G segment at A/G 141380 between HAL1 and (ATAAT)n, which is near the AluHG insertion locus and the MICF pseudogene (ID34 v ID48, ID34 v ID59, ID48 v ID68, and ID48 v ID94) (Figure 2).

Figure 2

The sequence comparison between the HLA-A*02:01:01 and HLA-A*02:05:01 haplotypes revealed large SNP differences within five segments from F to H of the alpha block (ID5 v ID8 and ID10 v ID8). This difference indicates that an exchange of segments T, K and A had occurred in the HLA-A*02:05:01 haplotype (Italian cell line WT49) that lacks the haplospecific A*02 lineage marker AluHG within the G segment (Supplementary Table 2). The sequence comparison between 4_A*03:01:01:01 and 18_A*03:01:01:01/A*24:02:01 with a low SNP density within the segments H to A is noteworthy because it reveals that the 18_A*03:01:01:01/A*24, if assembled correctly from the heterozygous Australian Caucasoid cell line LO081785, is an atypical and highly divergent A*03 haplotype with a HERVK9 deletion (Hap ID19.A*03.2 in Supplementary Table 6 and ID_18 in Supplementary Table 2). The SNP XO was within the G segment at nucleotide position 150337 C/T located between the Tigger1 and Charlie20a DNA elements (Table 6).

XO in Regions Between HLA-J and HLA-C

(a) SNP XOs within HLA-A haplotypes. The analysis of XO junctions in the genomic sequences between HLA-J and HLA-C outside the alpha and beta blocks using the same HLA-A and different HLA-C alleles was limited to only 15 comparisons (Supplementary Table 9). Of these, the XO was between HLA-A and HLA-J in two comparisons, between HLA-J and HLA-E in six pairs, and between HLA-E and MUC21 in six pairs. In the sequence comparison between 11_A*01-C*07-B*08 and 2_A*01-C*06-B*57 (example 6* in Supplementary Table 9), at least five different XO junctions were detected in different regions of segments B to E, including the transition from SPR to SRR in block B, two separate XO transitions in block C, a SRR to SPR XO in block D and a SPR to SRR XO in block E. Multiple XOs were observed also for analyses numbered 5 to 8 in Supplementary Table 9. The entire SNP-poor region (20 SNPs) was extended from HLA-J to HLA-C in the sequence alignment of cell lines with the same haplotype, A*01-C*06.

(b) SNP XOs within HLA-C haplotypes. Supplementary Table 10 shows the XO regions for 44 recombinant sequence pairs with the same HLA-C, but different HLA-A allelic lineages. A control homologous sequence pair (analysis 45) with the same HLA-A and HLA-C allelic lineages, 2_A*01-C*06-B*57 and 31_A*01-C*06-B*40, was SNP poor with only 20 SNPs counted over 1,322.9 kb of sequence. In comparison, the other sequences were heterologous (SNP rich) over most of the genomic region between the HLA-A and HLA-C genes. XOs occurred from SNP rich to SNP poor within the HLA-C gene or the 3′ non-coding region (NCR) of HLA-C in a comparison of 14 haplotype pairs, implying that recombinations occurred within the coding region of the ancestral HLA-C*07:18 and some other ancestral HLA-C*07 allelic lineages. The HLA-C haplotypic alleles that transitioned from SNP-poor to SNP-rich regions within 12 kb of the 3′end of exon 8 of HLA-C included C*01, C*04, C*05, C*07, C*08, C*12, and C*14 depending on their linkage with HLA-A alleles. The C*03/HLA-A*01, -A*02, -A*24, and the C*17/HLA-A*01, -A*30 combinations were SNP poor until the PSORS1C3 gene that is located about 100 kb from the HLA-C gene. Some of the HLA-C*06 and HLA-C*07 alleles within different recombinant haplotypes (analysis numbers 31 to 34) were homologous (SNP-poor) in the region from the HLA-C locus to the HCG22 locus. This long-range, homologous genomic segment of ~215 kb between HLA-C and HCG22 includes the various psoriasis candidate genes HCG27, PSORS1C3, POU5F1, TCF19, CCHCR1, PSORS1C2, PSORS1C1, CDSN, and C6orf15 (Nair et al., 2006). However, in combinatorial analyses of recombinant risk haplotypes and alleles genotyped in 678 families with early-onset psoriasis, Nair et al. (2006) determined that HLA-C*06 was the solitary risk allele that conferred susceptibility to early-onset psoriasis. Other HLA-C/HLA-A recombinant haplotypes were SNP poor over even greater distances ranging between 156 kb from the HLA-C gene to C6orf15 (C*01) and 730 kb from the HLA-C gene to a region between the LINC02569 and GNL1 loci. XO between SPR and flanking SRRs also occurred within or near to the MUC22 and MUC21 genes that are 234 kb and 279 kb from the HLA-C gene locus, respectively.

SNP-density XOs in Regions Between Different HLA-C and HLA-B Alleles

Table 7 shows the results of SNP counts and XO loci for 59 HLA-C and HLA-B haplotype sequence alignments using various combinations of similar and different haplotypes. As controls for comparing the various HLA-C/HLA-B recombinants, two were different HLA-C and HLA-B haplotypes and six were the same HLA-C and HLA-B haplotypes. The different haplotype pairs yielded an average of 1,656 SNP counts for ~93 kb, whereas the same haplotype pairs produced a few or no SNPs over the same distance between HLA-C and HLA-B loci. There were 39 sequence pairs of HLA-C/HLA-B recombinants with the same HLA-C allele and a different HLA-B allele, whereas 12 pairs of HLA-C/HLA-B recombinants had the same HLA-B allele and different HLA-C allele. Most SNP XO occurred near to or within the HLA-B or HLA-C coding region depending on which HLA-C/HLA-B recombinant haplotype sequences were aligned. SNP XOs occurred within the 3′ NCR or coding region of the HLA-B gene and within a 3-kb portion between a MLT1N2 element and the HLA-B gene (Figure 3) for 28 of 51 HLA-C/HLA-B haplotype alignments (Table 7). The SNP XOs in the intermediate loci regions were (1) within the L1PA13 fragment that is ~9 kb from HLA-C and ~72 kb from HLA-B, (2) between the L1MB8 and MLT1D elements ~53 kb from HLA-B, (3) a MIR element ~44 kb from HLA-B, and (4) within the HERVI ~21.5 kb from the HLA-B gene.

Figure 3

SNP-density XOs in Regions Between Various HLA-B and MIC Gene Alleles

The SNP XOs in the genomic region between HLA-B and MICA and between MICA and MICB for 31 recombinant haplotypes are listed in Table 8. There are 12 examples of a SNP XO within the genomic region of ~46.4 kb between HLA-B and MICA, and 13 examples of a SNP XO in the 94.5 kb genomic region between MICA and MICB. A SNP XO was detected in the MICA gene for the comparison between 54_B*35/MICA*002/MICB*005 and 68_B*35/MICA*016/MICB*002. The SNP XOs in the genomic region between MICA and MICB were within or near to putative recombinant hotspot REs: Tigger3b, MER2, THE1, MER21, LTR33, and ERV3-16A3_I of the HCP5 lncRNA gene (Figure 4). A SNP XO was found also within the insertion sequence of the SVA-MIC indel (Table 1) for the haplotypes 86_B*40/MICA*008/MICB*002 and 13_B*40/MICA*008/ MICB*014. SNP XO locations were not identified in five sequence comparisons because the genomic region from HLA-B to MICB was either too SNP-rich (comparisons 29 to 31) or too SNP-poor (comparisons 25 and 26).

Figure 4

Discussion

The comparative sequence analyses of the 95 genomic sequences using RepeatMasker (Hubley et al., 2016) and PIP (Schwartz et al., 2000) confirmed the identity and HLA class I allelic linkages of 12 haplotypic RE markers (Table 1) that had been genotyped previously in population frequency studies and MHC homozygous and heterozygous cell lines (Kulski and Dunn, 2005; Kulski et al., 2011). We identified another four novel REs and indels; three of them were annotated as AluOR1, AluP5, and SVA-BC indels (Table 1, Figure 1). A fourth indel was 9.5 kb in size and composed of the MER5B and LTR33 subfamily members (Supplementary Figure 1). The deletion variant of the 9.5-kb indel that is located between the SVA-BC and SVA-HB loci within the beta block was linked to all 11 HLA-C*06:02 alleles, all 4 HLA-C*12:03 alleles, and 1 of 9 C*04:01 alleles (Table 3). The 9.5-kb indel, SVA-BC and SVA-HB are all located within a relatively strong recombination hotspot of 82 kb between the HLA-B and HLA-C genes (Figures 1, 3). This intergenic recombination hotspot was identified previously in sperm studies (Cullen et al., 2002) and HapMap studies (Lam et al., 2013). Interestingly, a SNP XO was detected within the SVA-MIC insertion between the MICA and MICB loci (Figure 4) of the A*02/C*03:04/B*40 haplotype pair ID13 and ID86 (Table 3), suggesting that this insertion locus is near a recombination site that exchanged the MICB*002:01 allele with MICB*014.

We included eight HLA class I pseudogenes (V, P, H, T, K, U, W, and J) as haplotypic markers in our analysis of the alpha block haplotype diversity even though most of them may have no physiological functions or regulatory roles. However, the HLA-H gene is polymorphic and has transcriptional activity, and the signal peptide of the membrane-bound HLA-H molecule can mobilise HLA-E to the cell surface of mononuclear cells, bronchial epithelial cells and lymphoblastoid cell lines (Jordier et al., 2020). The HLA-H*02:07 allele was found at a frequency of 19.6% in some East Asian populations (Paganini et al., 2019) and encodes a full-length HLA protein that may have tolerogenic activity (Jordier et al., 2020) in comparison to immunosuppressive activity of its neighbouring gene product, HLA-G (Elliott, 2016). Taken together, the alleles of the pseudogenes confirmed that the alpha block haplotypes consist of multilocus alleles and that the recombinant XOs were mostly at the telomeric and centromeric ends of the block often within the telomeric F segment and/or the centromeric W and J segments (Table 2), but also in MICF fragment of the G segment close to the biallelic AluHG insertion site (Figure 2) and in proximity of HERVK9 of the HLA-H segment associated with the 54-kb deletion of the HLA-A*24 lineage (Supplementary Figure 5). Thus, the haplotypic HERVK9, AluHF, AluHG, and AluHJ insertion loci are located within 2 to 15 kb of these SNP XO junctions that might have co-evolved as a consequence of recombination events between AH. In previous population and homologous cell line genotyping studies, the AluHF insertion was associated with HLA-A*26, AluHJ was associated with HLA-A*01, -A*24 and -A*32 (Dunn et al., 2002; Kulski et al., 2019) and AluHG was associated with HLA-A*02 (Kulski et al., 2001) and HLA-G*01:01 (Santos et al., 2013) allelic lineages. The present study not only confirms these Alu haplospecificities and HLA associations but also has linked them to HLA pseudogene alleles and the non-classical HLA-F and HLA-G alleles (Table 2). Thus, the AluHF insertion, which is 12.4 kb telomeric of the HLA-F genes and within an ERV3-16A3_int sequence, is linked to HLA-F*01:01:01:08 and a number of different HLA-A alleles (A*02, A*03, A*26, A*29, and A*67) revealing the occurrence of frequent recombination events in the proximity of the HLA-F gene. The AluHG insertion is linked consistently to the HLA-G*01:01/HLA-H*01:01 haplotypic alleles and mostly to HLA-A*02. However, the occasional linkage of the AluHG/HLA-G*01:01/HLA-H*01:01 haplotype with HLA-A*30 and HLA-A*31 suggests a recombination site within a region somewhere nearer the HLA-A locus and probably within the A segment. Similarly, the AluHJ insertion at the centromeric terminal end of the alpha block is linked mainly to HLA-A*01/HLA-J*01:01:01:02, HLA-A*24/HLA-J*01:01:01:02 and HLA-A*33/HLA-J*01:01:01:06 of the 25 analysed haplotype sequences. The two exceptions were the HLA-J*01:01:01:02/AluHJ linkage in the heterologous A*03/A*24 sample and the linkage of HLA-J*01:01:01:02/AluHJ to the HLA-A*02 that again provides different examples of recombination activity within and at the borders of the alpha block.

The alpha block HERVK9 insertion in chimpanzee (Kulski et al., 2005), gorilla (Wilming et al., 2013) and human (Kulski et al., 2008) heralds its ancient hominid lineage. The loss and replacement of the HERVK9 sequence with a solitary MER9 sequence presumably generated the most recent alpha block haplotypes. In contrast, the absence of the SVA-HB from the beta block is the ancestral lineage for the SVA-HB indel (Kulski et al., 2010). The ancient HERVK9 insertion allele occurs frequently in the Caucasian, African American and Japanese ranging between 34 and 59% (Kulski et al., 2008) as does the SVA-HB insertion that was genotyped at 65% in Caucasians, 64% in African Americans and 25% in Japanese (Kulski et al., 2010). Both the alpha block HERVK9 insertion and the beta block SVA-HB insertion separate the HLA-A/B/C haplotypes into four distinct ancestral lineages, HERVK9+/SVA-B+, HERVK9+/SVA-B-, HERVK9-/SVA-B+ and HERVK9-/SVA-HB-. Other TE alleles and indels could be added to the TE classification system as more data evolve to assess MHC class I segmental exchanges and recombinants (Kulski et al., 2011), but this would require more annotated data on polymorphic TE and indels than presently available. Also, the question arises whether the nine different alpha-beta block haplotypes (Tables 2, 3) with the HERVK9 insertion and no SVA-HB insertion that includes A*03:01/C*07:02/B*07:02 are older than those with the alpha block HERVK9 deletion and the beta block SVA-HB insertion (28 different haplotypes) (Supplementary Table 5). Although the 7.1AH with the HERVK9+/SVA-B- alleles might be older than 8.1AH, 27.1AH, 57.1AH, 60.1AH, and 62.1AH that are HERVK9-/SVA-HB-, classifying the evolutionary age of the haplotypes according to the presence and absence of HERVK9 and SVA-HB is probably unreliable because of segmental conversions and exchanges between genomic sequences. In addition, the 7.1AH and the 8.1AH are among the most common European haplotypes with population frequencies of 8.1AH up to at least 18% (Sanchez-Velasco et al., 2003; Kiszel et al., 2007; Sanchez-Mazas et al., 2013; Robinson et al., 2019). In addition, seven of the other nine HERVK9+/SVA-B- haplotypes (Supplementary Table 5) are common in Europeans (Suslova et al., 2012; Robinson et al., 2019) or Japanese (Ikeda et al., 2015), whereas A*31/C01:02/B*15:01 is from the cell line of a Warao South American Indian, which may be moderately common in some Amerindians (Watkins et al., 1992; Cadavid and Watkins, 1997), but not in others (Barquera et al., 2020). Taken together, these observations suggest that the human MHC ERVK9+/SVA-HB haplotype might be broadly spread at relatively high frequency within many different worldwide populations.

The most commonly inferred haplotype blocks that are free of genetic recombination are supposedly those that are identified by LD statistical tests of SNP associations (Daly et al., 2001; Ahmad et al., 2003; Miretti et al., 2005; Blomhoff et al., 2006, Traherne, 2008). However, SNP densities vary markedly throughout the human genome (Sachidanandam et al., 2001) and their transition between SNP-rich and SNP-poor regions can be used to identify recombination free haplotypes without the need for LD tests (Myers, 2005; Kauppi et al., 2007; Bairagya et al., 2008). In the present study, we identified the junctions of haplotype blocks on the basis of SNP densities using homozygous haplotype sequences without applying LD statistical tests. We classified the haplotype junctions as XOs between SNP poor and SNP rich regions of haploidentical and haplodiverse genomic sequence comparisons (Tables 58). The SNP XOs between rich and poor SNP regions revealed a clear block structure both for haplotype mapping potential ancestral recombination sites and for objective analysis of haplotypes in different populations for disease associations. Since block patterns vary between common haplotypes, it will be important to construct maps for each different haplotype from a variety of populations to assess the degree of haplotype diversity within and between populations (Goodin et al., 2018). Many of the XOs that we identified are in close vicinity of recombination hotspots (Figure 1) that were previously identified in studies of MHC recombination sites using homozygous sperm DNA (Cullen et al., 2002) and HapMap populations (Lam et al., 2013).

SNP densities in our haplotype alignments were counted manually while avoiding obvious assembly errors, polynucleotides, simple microsatellite repeats, and indels by scanning overlapping windows of 50 to 500 kb of sequence. SNP-poor regions were easy to count manually because of small SNP numbers (<5 SNPs/100 kb), but SNP-rich regions were more difficult because of large SNP numbers at an average of 7 SNPs/kb. In comparison to our findings, Lam et al. (2015) reported an average of 7 SNPs/100 kb for the A33-B58-DR3 haplotype for 4136 kb and 8.7 SNPs/100 kb for the A2-B46-DR9 for 2721 kb, which are much higher than our comparative counts. This difference is possibly due to spikes of diversity in their localised regions of sequence. The SNP counts by Lam et al. (2015) within homologous haplotypes represented as nucleotide diversity values were still at least 38× less than that found between 7.1CEH/AH and 8.1CEH/AH, the two different common European MHC heterologous haplotypes of the cell lines PGF and COX, respectively. The count of 2240 SNPs per 320.8 kb (7 SNPs/kb) for PGF (LAB ID_4) and COX (LAB ID_11) in our study when their alpha blocks were aligned (Table 6) further demonstrates the large SNP diversity between them. In comparison, the average SNP density of the human genome is one SNP per 1.9 kb or between 5 and 9 SNPs per 10 kb (Sachidanandam et al., 2001; Zhao et al., 2003).

The two longest MHC haplotypes in the human genome are considered to be the European 7.1AH and 8.1AH (Horton et al., 2004, 2008) with CPS extending into the telomeric OR gene cluster region beyond the GPX5 gene that is 1.2 Mb telomeric of HLA-F (Figure 1). However, the CPS of at least four other MHC class I haplotype pairs also extended beyond the GPX5 gene: 18.2AH, 51.xAH, 57.1AH, 62.xAH, and 62.1AH (Table 5, Supplementary Table 8). The CPS of another three HLA-A haplotype pairs also extend beyond the GPX5 gene: HLA-A*24/C*03:04, HLA-A*31/C*15:02 and HLA-A*32/C*12:03 presented in Supplementary Table 8. The SNP XO junctions could not be determined for these haplotypes because they are telomeric of the GPX5 gene (NCBI ID:2880) and located somewhere beyond the region of the available genomic sequences provided for these cell lines (Norman et al., 2017). In comparison, the SNP XO junctions for two other HLA-A*01/C*06 and HLA-A*01/C*07 pairs and two HLA-A*03/C*07:02 and HLA-A*03/C*06:02 pairs were closer to the MOG and OR2H2 genes (Table 5) that are 51 kb and 134 kb from the HLA-F gene, respectively (Figure 1). It seems that the recombinant breakpoint of different haplotype segments or blocks generates the SNP XO site even if the classical HLA alleles are the same. For example, a number of HLA-A*02 allelic lineages were represented by distinctly different alpha block haplotypes (Table 2) that evolved from various shuffling events as well as segmental conversions (Table 6). Thus, the SNP XO within the alpha block F segment was strongly correlated with different local HLA-F alleles even if the downstream HLA-A allele within the A segment of the alpha block was the same; that is, the XO between HLA-F*01:01:01:01 of 18.xAH or 38.a2AH and F*01:04:01:02 of 60.1AH are both linked to HLA-A*02. The SNP XO defines the junction of the haplotype XO, which in turn points towards a putative ancestral recombination site. Whether this relationship that is based on a small number of comparative examples can be confirmed in the future will depend on the empirical findings of a much greater number of haplotypes with different HLA-A allelic and haplotypic lineages.

During mammalian and primate evolution, the MHC region transitioned through various genomic rearrangements including recombinant XO events of MIC, HERV16 and HLA class I genes, which resulted in the current structural organisation of the human MHC class I genomic region (Kulski et al., 1999a,b; Kulski et al., 2004). The HERV16, now classified as ERV3-16A3 (repeated throughout the MHC class I and class II regions), along with HLA class I coding and non-coding sequences, seems to have been a recombination site for many of the duplication events by way of unequal XOs (Kulski et al., 1999a,b; Kulski et al., 2004). Ancient hominoid haplotypes undoubtedly were the progenitors to the modern human CEH/AH, but when, how and where is unknown. In addition, each AH is a unique integrated genetic module consisting of many immunologically related protein-coding genes with gene copy number variations, segmental duplications and fragmented or relatively intact transposons and REs that contribute to more than 50% of the genomic content. In this regard, the MHC haplotype comprising a cluster of multilocus, monocistronic expression units is analogous to the polycistronic bacterial operon, which is a functional unit of DNA containing a cluster of genes under the control of a single promoter (Lee and Sonnhammer, 2003; Blumenthal, 2004). However, the MHC haplotype structures are far more complex with their regulatory network of cis-acting multilocus expression units known as haplotype-specific expression quantitative trait loci (eQTL) (Lamontagne et al., 2016; Lam et al., 2017) that are largely controlled by an array of virus-derived REs and DNA transposons, both as binding sites for transcription factors and as sources of regulatory non-coding RNAs (Kulski, 2019; Sznarkowska et al., 2020). Random mutations, methylations and recombinations can generate considerable haplotype diversity that is part of the MHC immune system's response to highly prevalent infectious (Gao et al., 2019; Sanchez-Mazas, 2020) and chemical agents, including those responsible for drug hypersensitivities (Alfirevic and Pirmohamed, 2010).

Various factors have been hypothesised to elucidate the generation and maintenance of MHC CEHs/AHs including recombination suppression, balancing selection and demographic factors such as population bottlenecks, genetic drift, migration and admixture (Aly et al., 2006; van Oosterhout, 2009; Prohaska et al., 2019). If CEHs/AHs were generated in recent population history, for example, during the last 21,000 to 26,000 years as estimated from the mutation rates of the Caucasian A1-B8-DR3-DQ2 haplotype (Smith et al., 2006) and the Asian A2-B46-DR9 and A33-B58-DR3 haplotypes (Lam et al., 2013), then there may have been insufficient time over a period of a few thousand generations to have disrupted the LRHs by recombination. It seems that there is a time-associated equilibrium between the population amplification of the LRH and its meiotic recombinational breakage over periods of human population coalescent times (Song et al., 2017; Wang et al., 2020). Polymorphism and sequence heterozygosity also might suppress crossing over and recombination (Ohta, 1999; Dluzewska et al., 2018) and it is well-known that an excess of heterozygosity (overdominance) contributes to MHC diversity as one form of balancing selection (van Oosterhout, 2009; Lenz et al., 2016; Lobkovsky et al., 2019). However, in the context of molecular mechanisms and our study, the connection between the silencing of TE/REs and recombination suppression warrants greater consideration (Campos-Sánchez et al., 2016). TEs are known to affect recombination rates by acting as recombination modifiers, activators and suppressors in mice and humans (McVean, 2010; Altemose et al., 2017; Yamada et al., 2017). Although TEs can directly modulate the local recombination environment either through silencing-mediated suppression or the recruitment of recombination hotspots, the silencing of TEs and other repetitive sequences may also contribute directly to recombination suppression. The densities of DNA methylation and repressive chromatin marks associated with the silencing of TEs and other repeats are often negatively correlated with recombination rates (Myers, 2005; Kent et al., 2017). A recombination suppression mechanism discovered and studied in recent years involves the PRDM9-mediated recombination machinery that initiates at specific sequence motifs and alters chromatin structure (Myers et al., 2010; Parvanov et al., 2017). In humans, PRDM9 determines the locations of meiotic recombination hotspots and binds multiple motifs including the ATCCATG/CATGGAT motif of the THE1B repeat for both dependent and independent recombination suppression (Myers et al., 2008; Altemose et al., 2017). The active PRDM9 binding sites are also enriched with other classes of human repeat sequences including L2 LINEs and AluY elements (Myers et al., 2008; McVean, 2010). Recent studies with B6 mice demonstrated that a third of meiotic DNA strand breaks occurred within repetitive sequences of different classes, especially within the DNA transposons like TcMar-Mariner and hAT-Charlie, that resulted in their depletion from the PRDM9 binding sites (Yamada et al., 2017). This finding led the authors to hypothesise that the PRDM9 coevolved with meiotic recombination in order to target active transposons and limit their spread by inactivating or eliminating them by creating mutations or deletions at the PRDM9 binding site. Moreover, a proportion of the duplicated MHC ERV3-16A3_I (HERV16 int) sequences (Kulski and Dawkins, 1999; Kulski et al., 1999a,b) have the PRDM9 binding motif ATCCATG/CATGGAT at one or two of its nucleotide positions (Supplementary Figure 7).

Although TEs, methylations and recombinations can influence each other markedly (Myers et al., 2008; McVean, 2010; Moolhuijzen et al., 2010; Jones, 2012; Zamudio et al., 2015; Altemose et al., 2017; Kent et al., 2017), there is a surprising paucity of such studies in the MHC genomic region of different haplotypes (Rakyan et al., 2004; Jongsma et al., 2019). The genomic PRDM9 binding motif ATCCATG/CATGGAT (Altemose et al., 2017) is spread broadly throughout the MHC class I region and can be found in various fragmented TEs: L1, L2, MER5, HUERS-P3-int (2×), HERVK9-int (2×), HERVK14 (2×), LTR73, MER2, MER8, MER41, MER84, Charlie9, MLT1, MLT2B, Tigger1, ERV3-16A3_I, four of ~1450 Alu elements and various coding regions including the MICB sequence (present study, data not shown). The 38 SNP XOs at the borders of the haplotype blocks that are described in our study appear to reflect regions of historical recombination sites that currently might be involved with recombination suppression and haplotype maintenance (Figure 1). Some TEs are likely to have provided the recombination sequence motifs that Cullen et al. (2002) considered could be due to microsatellites and particularly to those with long tracts of GT repeats. Many of the SNP XO junctions are within 10 kb of TEs that are commonly repeated throughout the MHC genomic region including LTR16B/ERV3-16A3_I, L1, L2, MLT1, THE1, Charlie, Tigger, MST, and MER5 sequences. The presence of the LTR16B/ERV3-16A3_I sequences at the XO and recombination sites is not surprising since this RE was associated with various genomic rearrangements including XO events of MIC, HERV16, and HLA class I genes, which influenced the structural organisation of the MHC locus during primate evolution (Kulski et al., 1997, 1999a,b, 2002b, 2004, 2005). The LTR16B/ERV3-16A3_I and THE1 elements often are located together in close proximity to the XOs. Moreover, there are ATCCATG/CATGGAT motifs in the MER2 and L2 that flank the ERV3-16A3-int extension of the HCP5 gene (Kulski, 2019) that is within a recombination hotspot located between the MICA and MICB genes (Table 8, Figure 4). On the other hand, the THE1 elements in the MHC genomic region are represented by different subfamilies, and most of them lack the ATCCATG motif and therefore might not interact with PRDM9. Some active or young TE insertions are found in regions close to meiotic recombination sites. Therefore, the TE insertions in the MHC class I region that have yet to reach fixation such as the structurally polymorphic Alu and SVA elements (Table 1) are regional markers of insertion bias, and they are in close proximity to SNP XOs that could be active recombination hotspots (Figure 1). This implies that young TE insertion polymorphisms are relatively recent ancestral recombination hotspots; that is, the younger and more haplospecific TE insertions represent the integration and recombination sites of younger haplotype segments (Campos-Sánchez et al., 2016), whereas the fixed TE insertions, such as SVA-16, -T26, -ER, -EG, -M21, and -M22 (Table 1, Figure 1), reveal the older haplotype recombination XO spots of our primate progenitors (Anzai et al., 2003, Wilming et al., 2013). The presence of SNP XOs near to or within the HLA-B and HLA-C genes (Tables 7, 8) suggests that these genes and/or neighbouring TE's L1, Alu, L2, MLT1, MST, MER21, MER41, LTR9, MER1, and MER4 (Kulski et al., 1997) have had a role in recombination (Cullen et al., 2002, Lam et al., 2013) along with the occasional gene conversions (Madrigal et al., 1993, Adamek et al., 2015). Many of these old elements may now contribute to recombination suppression. Further detailed sequence multiple alignment studies at SNP XOs using HLA-B/HLA-C recombinant haplotypes such as previously described by Nair et al. (2006) in their psoriasis association studies may help to resolve this consideration.

The genomic sequences that we analysed in this study have important implications in medical research and treatment (Lokki and Paakkanen, 2019). Genotyping SNPs of the five major HLA class I and class II loci for cross matching provides most of the haplotype information needed for successful transplantation outcomes. However, ignoring SNPs outside the five loci when comparing the same or similar haplotypes may be problematic and misleading and result in choosing the wrong SNP markers for GWAS of disease or phenotypes, although GWAS need to be correlated to the MHC haplotype and not to the SNP per se. Nevertheless, particular haplospecific segments can be used to identify likely “disease” genes or regions (Lokki and Paakkanen, 2019) in comparative haplomics as previously performed for the psoriasis gene (Nair et al., 2006). Haplotypic or haplospecific markers like the well-defined HLA alleles and dimorphic RE markers may assist in narrowing genomic segments and loci towards MHC disease regions. Also, the haplospecific and haplotypic regions of the non-HLA coding regions such as the ~39 non-HLA genes between HLA-A and HLA-C are still poorly defined and need to be better characterised as essential components of haplotype regulatory modules. We found some SNP XO junctions in the non-HLA gene region between HLA-E and HLA-C that might have important implications in affecting disease. The systematic comparison of various recombinant haplotypes (Nair et al., 2006) is a promising approach that has been little utilised in GWAS (Alper and Larsen, 2017; Lokki and Paakkanen, 2019). Imputing haplotypes from 3D genome structures of single diploid human cells (Tan et al., 2018) along with tagging regulatory TEs using chromosome conformation capture techniques (Raviram et al., 2018) might be new and productive technical approaches to investigate haplotype regulatory modules.

Conclusion

Our study confirms that the genomic sequences of MHC homozygous cell lines are useful for analysing MHC haplotypic landscapes and characterising unique CPS, haplotypic markers, and XO zones without a need to use LD or other probabilistic statistical imputations. Comparative sequence analyses confirmed the identity of 12 haplotypic RE markers and revealed that the HERVK9 indel within the alpha block and a SVA indel within the beta block divided the HLA-A/B/C haplotypes into a series of distinctly interrelated historical lineages, and we identified numerous ancestral segmental XOs between different haplotypes within various REs, lncRNA, MUC22, C6orf15, and HLA-C and/or HLA–B genes extending over 2 Mb from the HLA-A to the MICB loci. It is evident from this study and previous studies that there is a vast MHC haplotype and allelic diversity in the human and that we have captured only a fraction of the complexity. Here, we analysed and characterised the polymorphic REs and the duplicated copies of MHC class I genes within genomic sequences of 95 haplotypes that were sequenced and assembled previously by Norman et al. (2017) in order to broaden the framework of the reference sequences so that they can be further improved and utilised to interrogate the human MHC in greater detail. More attention than usual was given to the polymorphic TE and RE at the SNP-density XOs as potential recombination hotspots. A greater emphasis on the commonality and differences of MHC class I recombinants may set the scene for better functional studies involving the described MHC alleles and haplotypes and their role in immunity, transplantation and overall health and well-being.

Statements

Data availability statement

All datasets generated for this study are included in the article/Supplementary Material.

Ethics statement

Ethical review and approval was not required for the study on human participants in accordance with the local legislation and institutional requirements. Written informed consent for participation was not required for this study in accordance with the national legislation and the institutional requirements.

Author contributions

JK carried out the analyses of the repeat elements, SNP-density XOs and interpretation of the data, and wrote the manuscript. SS and TS analyzed and interpreted parts of the data and provided the alleles for the non-classical HLA class I genes, HLA pseudogenes, and the MICA and MICB genes. All authors checked the final version of the paper.

Acknowledgments

JK thanks Karen Barfield for her help with word Table formatting and Enrico Rossi for proof-reading.

Conflict of interest

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fgene.2020.594318/full#supplementary-material

References

  • 1

    AdamekM.KlagesC.BauerM.KudlekE.DrechslerA.LeuserB.et al. (2015). Seven novel HLA alleles reflect different mechanisms involved in the evolution of HLA diversity: description of the new alleles and review of the literature. Hum. Immunol.76, 3035. 10.1016/j.humimm.2014.12.007

  • 2

    AhmadT.NevilleM.MarshallS. E.ArmuzziA.Mulcahy-HawesK.CrawshawJ.et al. (2003). Haplotype-specific linkage disequilibrium patterns define the genetic topography of the human MHC. Hum. Mol. Genet.12, 647656. 10.1093/hmg/ddg066

  • 3

    AlfirevicA.PirmohamedM. (2010). Drug induced hypersensitivity and the HLA complex. Pharmaceuticals4, 6990. 10.3390/ph4010069

  • 4

    AlperC. A.LarsenC. E. (2017). “Pedigree-defined haplotypes and their applications to genetic studies,” in Haplotyping Methods in Molecular Biology, eds. Tiemann-BoegeI.BetancourtA. (New York, NY: Springer New York), 113127.

  • 5

    AlperC. A.LarsenC. E.DubeyD. P.AwdehZ. L.FiciD. A.YunisE. J. (2006). The haplotype structure of the human major histocompatibility complex. Hum. Immunol.67, 7384. 10.1016/j.humimm.2005.11.006

  • 6

    AlperC. A.RaumD.KarpS.AwdehZ. L.YunisE. J. (1983). Serum complement ‘supergenes’ of the major histocompatibility complex in man (complotypes). Vox Sang.45, 6267. 10.1111/j.1423-0410.1983.tb04124.x

  • 7

    AltemoseN.NoorN.BitounE.TumianA.ImbeaultM.ChapmanJ. R.et al. (2017). A map of human PRDM9 binding provides evidence for novel behaviors of PRDM9 and other zinc-finger proteins in meiosis. eLife6:e28383. 10.7554/eLife.28383

  • 8

    AlyT. A.EllerE.IdeA.GowanK.BabuS. R.ErlichH. A.et al. (2006). Multi-SNP analysis of MHC region: remarkable conservation of HLA-A1-B8-DR3 haplotype. Diabetes55, 12651269. 10.2337/db05-1276

  • 9

    AnzaiT.ShiinaT.KimuraN.YanagiyaK.KoharaS.ShigenariA.et al. (2003). Comparative sequencing of human and chimpanzee MHC class I regions unveils insertions/deletions as the major path to genomic divergence. Proc. Natl. Acad. Sci. U.S.A.100, 77087713. 10.1073/pnas.1230533100

  • 10

    AwdehZ. L.RaumD.YunisE. J.AlperC. A. (1983). Extended HLA/complement allele haplotypes: evidence for T/t-like complex in man. Proc. Natl. Acad. Sci. U.S.A.80, 259263. 10.1073/pnas.80.1.259

  • 11

    AyarpadikannanS.KimH.-S. (2014). The impact of transposable elements in genome evolution and genetic instability and their implications in various diseases. Genomics Inform.12, 98104. 10.5808/GI.2014.12.3.98

  • 12

    BairagyaB. B.BhattacharyaP.BhattacharyaS. K.DeyB.DeyU.GhoshT.et al. (2008). Genetic variation and haplotype structures of innate immunity genes in eastern India. Infect. Genet. Evol.8, 360366. 10.1016/j.meegid.2008.02.009

  • 13

    BaoW.KojimaK. K.KohanyO. (2015). Repbase update, a database of repetitive elements in eukaryotic genomes. Mob. DNA6:11. 10.1186/s13100-015-0041-9

  • 14

    BarqueraR.ZunigaJ.Flores-RiveraJ.CoronaT.PenmanB. S.Hernández-ZaragozaD. I.et al. (2020). Diversity of HLA class I and class II blocks and conserved extended haplotypes in Lacandon Mayans. Sci. Rep.10:3248. 10.1038/s41598-020-58897-5

  • 15

    BilbaoJ. R.CalvoB.AransayA. M.Martin-PagolaA.Perez de NanclaresG.AlyT. A.et al. (2006). Conserved extended haplotypes discriminate HLA-DR3-homozygous Basque patients with type 1 diabetes mellitus and celiac disease. Genes Immun.7, 550554. 10.1038/sj.gene.6364328

  • 16

    BlomhoffA.OlssonM.JohanssonS.AkselsenH. E.PociotF.NerupJ.et al. (2006). Linkage disequilibrium and haplotype blocks in the MHC vary in an HLA haplotype specific manner assessed mainly by DRB1*03 and DRB1*04 haplotypes. Genes Immun.7, 130140. 10.1038/sj.gene.6364272

  • 17

    BlumenthalT. (2004). Operons in eukaryotes. Brief. Funct. Genomic. Proteomic.3, 199211. 10.1093/bfgp/3.3.199

  • 18

    BourgeoisY.BoissinotS. (2019). On the population dynamics of junk: a review on the population genomics of transposable elements. Genes10:419. 10.3390/genes10060419

  • 19

    BrawandD.WagnerC. E.LiY. I.MalinskyM.KellerI.FanS.et al. (2014). The genomic substrate for adaptive radiation in African cichlid fish. Nature513, 375381. 10.1038/nature13726

  • 20

    CadavidL. F.WatkinsD. I. (1997). Heirs of the jaguar and the anaconda: HLA, conquest and disease in the indigenous populations of the Americas. Tissue Antigens50, 702711. 10.1111/j.1399-0039.1997.tb02940.x

  • 21

    Campos-SánchezR.CremonaM. A.PiniA.ChiaromonteF.MakovaK. D. (2016). Integration and fixation preferences of human and mouse endogenous retroviruses uncovered with functional data analysis. PLOS Comput. Biol.12:e1004956. 10.1371/journal.pcbi.1004956

  • 22

    ChuongE. B.EldeN. C.FeschotteC. (2017). Regulatory activities of transposable elements: from conflicts to benefits. Nat. Rev. Genet.18, 7186. 10.1038/nrg.2016.139

  • 23

    ContuL.CarcassiC.DaussetJ. (1989). The “Sardinian” HLA-A30,B18,DR3,DQw2 haplotype constantly lacks the21-OHA andC4B genes. is it an ancestral haplotype without duplication? Immunogenetics30, 1317. 10.1007/BF02421464

  • 24

    CullenM.PerfettoS. P.KlitzW.NelsonG.CarringtonM. (2002). High-resolution patterns of meiotic recombination across the human major histocompatibility complex. Am. J. Hum. Genet.71, 759776. 10.1086/342973

  • 25

    DalyM. J.RiouxJ. D.SchaffnerS. F.HudsonT. J.LanderE. S. (2001). High-resolution haplotype structure in the human genome. Nat. Genet.29, 229232. 10.1038/ng1001-229

  • 26

    DawkinsR.LeelayuwatC.GaudieriS.TayG.HuiJ.CattleyS.et al. (1999). Genomics of the major histocompatibility complex: haplotypes, duplication, retroviruses and disease. Immunol. Rev.167, 275304. 10.1111/j.1600-065X.1999.tb01399.x

  • 27

    de BakkerP. I. W.McVeanG.SabetiP. C.MirettiM. M.GreenT.MarchiniJ.et al. (2006). A high-resolution HLA and SNP haplotype map for disease association studies in the extended human MHC. Nat. Genet.38, 11661172. 10.1038/ng1885

  • 28

    de KoningA. P. J.GuW.CastoeT. A.BatzerM. A.PollockD. D. (2011). Repetitive elements may comprise over two-thirds of the human genome. PLoS Genet.7:e1002384. 10.1371/journal.pgen.1002384

  • 29

    Degli-EspostiM. A.LeaverA. L.ChristiansenF. T.WittC. S.AbrahamL. J.DawkinsR. L. (1992). Ancestral haplotypes: conserved population MHC haplotypes. Hum. Immunol.34, 242252. 10.1016/0198-8859(92)90023-G

  • 30

    DiltheyA. T.GourraudP.-A.MentzerA. J.CerebN.IqbalZ.McVeanG. (2016). High-accuracy HLA type inference from whole-genome sequencing data using population reference graphs. PLOS Comput. Biol.12:e1005151. 10.1371/journal.pcbi.1005151

  • 31

    DluzewskaJ.SzymanskaM.ZiolkowskiP. A. (2018). Where to cross over?defining crossover sites in plants. Front. Genet.9:609. 10.3389/fgene.2018.00609

  • 32

    DorakM. T.ShaoW.MachullaH. K. G.LobashevskyE. S.TangJ.ParkM. H.et al. (2006). Conserved extended haplotypes of the major histocompatibility complex: further characterization. Genes Immun.7, 450467. 10.1038/sj.gene.6364315

  • 33

    DoverG. (1982). Molecular drive: a cohesive mode of species evolution. Nature299, 111117. 10.1038/299111a0

  • 34

    DoxiadisG. G. M.de GrootN.ClaasF. H. J.DoxiadisI. I. N.van RoodJ. J.BontropR. E. (2007). A highly divergent microsatellite facilitating fast and accurate DRB haplotyping in humans and rhesus macaques. Proc. Natl. Acad. Sci. U.S.A.104, 89078912. 10.1073/pnas.0702964104

  • 35

    DunnD. S.InokoH.KulskiJ. K. (2003). Dimorphic Alu element located between the TFIIH and CDSN genes within the major histocompatibility complex. Electrophoresis24, 27402748. 10.1002/elps.200305524

  • 36

    DunnD. S.NaruseT.InokoH.KulskiJ. K. (2002). The association between HLA-A alleles and young alu dimorphisms near the HLA-J, -H, and -F genes in workshop cell lines and Japanese and Australian populations. J. Mol. Evol.55, 718726. 10.1007/s00239-002-2367-4

  • 37

    DunnD. S.TaitB. D.KulskiJ. K. (2005). The distribution of polymorphic Alu insertions within the MHC class I HLA-B7 and HLA-B57 haplotypes. Immunogenetics56, 765768. 10.1007/s00251-004-0745-3

  • 38

    ElliottR. L. (2016). Cancer immunotherapy “HLA-G an Important Neglected Immunosuppressive Molecule.”SOJ Immunol.4, 13. 10.15226/2372-0948/4/1/00146

  • 39

    GambinoC. M.AielloA.AccardiG.CarusoC.CandoreG. (2018). Autoimmune diseases and 8.1 ancestral haplotype: an update. HLA92, 137143. 10.1111/tan.13305

  • 40

    GaoJ.ZhuC.ZhuZ.TangL.LiuL.WenL.et al. (2019). The human leukocyte antigen and genetic susceptibility in human diseases:J. Bio-X Res.2, 112120. 10.1097/JBR.0000000000000044

  • 41

    GaudieriS.DawkinsR. L.HabaraK.KulskiJ. K.GojoboriT. (2000). SNP profile within the human major histocompatibility complex reveals an extreme and interrupted level of nucleotide diversity. Genome Res.10, 15791586. 10.1101/gr.127200

  • 42

    GaudieriS.KulskiJ. K.DawkinsR. L.GojoboriT. (1999). Extensive nucleotide variability within a 370 kb sequence from the central region of the major histocompatibility complex. Gene238, 157161. 10.1016/S0378-1119(99)00255-3

  • 43

    GaudieriS.LeelayuwatC.TayG. K.TownendD. C.DawkinsR. L. (1997). The major histocompatability complex (MHC) contains conserved polymorphic genomic sequences that are shuffled by recombination to form ethnic-specific haplotypes. J. Mol. Evol.45, 1723. 10.1007/PL00006194

  • 44

    GeorgeC. M.AlaniE. (2012). Multiple cellular mechanisms prevent chromosomal rearrangements involving repetitive DNA. Crit. Rev. Biochem. Mol. Biol.47, 297313. 10.3109/10409238.2012.675644

  • 45

    GoodinD. S.KhankhanianP.GourraudP.-A.VinceN. (2018). Highly conserved extended haplotypes of the major histocompatibility complex and their relationship to multiple sclerosis susceptibility. PLoS ONE13:e0190043. 10.1371/journal.pone.0190043

  • 46

    GuW.ZhangF.LupskiJ. R. (2008). Mechanisms for human genomic rearrangements. PathoGenetics1:4. 10.1186/1755-8417-1-4

  • 47

    GuoZ.HoodL.MalkkiM.PetersdorfE. W. (2006). Long-range multilocus haplotype phasing of the MHC. Proc. Natl. Acad. Sci. U.S.A.103, 69646969. 10.1073/pnas.0602286103

  • 48

    HennB. M.Cavalli-SforzaL. L.FeldmanM. W. (2012). The great human expansion. Proc. Natl. Acad. Sci. U.S.A.109, 1775817764. 10.1073/pnas.1212380109

  • 49

    HortonR.GibsonR.CoggillP.MirettiM.AllcockR. J.AlmeidaJ.et al. (2008). Variation analysis and gene annotation of eight MHC haplotypes: the MHC haplotype project. Immunogenetics60, 118. 10.1007/s00251-007-0262-2

  • 50

    HortonR.WilmingL.RandV.LoveringR. C.BrufordE. A.KhodiyarV. K.et al. (2004). Gene map of the extended human MHC. Nat. Rev. Genet.5, 889899. 10.1038/nrg1489

  • 51

    HuangM.ZhuM.JiangT.WangY.WangC.JinG.et al. (2019). Fine mapping the MHC region identified rs4997052 as a new variant associated with nonobstructive azoospermia in Han Chinese males. Fertil. Steril.111, 6168. 10.1016/j.fertnstert.2018.08.052

  • 52

    HubleyR.FinnR. D.ClementsJ.EddyS. R.JonesT. A.BaoW.et al. (2016). The Dfam database of repetitive DNA families. Nucleic Acids Res.44, D81D89. 10.1093/nar/gkv1272

  • 53

    IkedaN.KojimaH.NishikawaM.HayashiK.FutagamiT.TsujinoT.et al. (2015). Determination of HLA-A, -C, -B, -DRB1 allele and haplotype frequency in Japanese population based on family study: HLA allele and haplotype frequency in Japanese population. Tissue Antigens85, 252259. 10.1111/tan.12536

  • 54

    JeffreysA. J.HollowayJ. K.KauppiL.MayC. A.NeumannR.SlingsbyM. T.et al. (2004). Meiotic recombination hot spots and human DNA diversity. Philos. Trans. R. Soc. Lond. B Biol. Sci.359, 141152. 10.1098/rstb.2003.1372

  • 55

    JensenJ. M.VillesenP.FriborgR. M.The Danish Pan-Genome Consortium Mailund, T. Besenbacher S.. (2017). Assembly and analysis of 100 full MHC haplotypes from the Danish population. Genome Res.27, 15971607. 10.1101/gr.218891.116

  • 56

    JonesP. A. (2012). Functions of DNA methylation: islands, start sites, gene bodies and beyond. Nat. Rev. Genet.13, 484492. 10.1038/nrg3230

  • 57

    JongsmaM. L. M.GuardaG.SpaapenR. M. (2019). The regulatory network behind MHC class I expression. Mol. Immunol.113, 1621. 10.1016/j.molimm.2017.12.005

  • 58

    JordierF.GrasD.De GrandisM.D'JournoX.-B.ThomasP.-A.ChanezP.et al. (2020). HLA-H: transcriptional activity and HLA-E mobilization. Front. Immunol.10:2986. 10.3389/fimmu.2019.02986

  • 59

    KarellK.KlingerN.HolopainenP.LevoA.PartanenJ. (2000). Major histocompatibility complex (MHC)- linked microsatellite markers in a founder population. Tissue Antigens56, 4551. 10.1034/j.1399-0039.2000.560106.x

  • 60

    KauppiL.JasinM.KeeneyS. (2007). Meiotic crossover hotspots contained in haplotype block boundaries of the mouse genome. Proc. Natl. Acad. Sci. U.S.A.104, 1339613401. 10.1073/pnas.0701965104

  • 61

    KentT. V.UzunovićJ.WrightS. I. (2017). Coevolution between transposable elements and recombination. Philos. Trans. R. Soc. B Biol. Sci.372:20160458. 10.1098/rstb.2016.0458

  • 62

    KirknessE. F.GrindbergR. V.Yee-GreenbaumJ.MarshallC. R.SchererS. W.LaskenR. S.et al. (2013). Sequencing of isolated sperm cells for direct haplotyping of a human genome. Genome Res.23, 826832. 10.1101/gr.144600.112

  • 63

    KiszelP.KovácsM.SzalaiC.YangY.PozsonyiÉ.BlaskóB.et al. (2007). Frequency of carriers of 8.1 ancestral haplotype and its fragments in two caucasian populations. Immunol. Invest.36, 307319. 10.1080/08820130701241404

  • 64

    KulskiJ. K. (2019). Long noncoding RNA HCP5, a hybrid HLA class I endogenous retroviral gene: structure, expression, and disease associations. Cells8:480. 10.3390/cells8050480

  • 65

    KulskiJ. K.AlSafarH. S.MawartA.HenschelA.TayG. K. (2019). HLA class I allele lineages and haplotype frequencies in Arabs of the United Arab Emirates. Int. J. Immunogenet.46, 152159. 10.1111/iji.12418

  • 66

    KulskiJ. K.AnzaiT.InokoH. (2005). ERVK9, transposons and the evolution of MHC class I duplicons within the alpha-block of the human and chimpanzee. Cytogenet. Genome Res.110, 181192. 10.1159/000084951

  • 67

    KulskiJ. K.AnzaiT.ShiinaT.HidetoshiI. (2004). Rhesus macaque class I duplicon structures, organization, and evolution within the alpha block of the major histocompatibility complex. Mol. Biol. Evol.21, 20792091. 10.1093/molbev/msh216

  • 68

    KulskiJ. K.DawkinsR. L. (1999). The P5 multicopy gene family in the MHC is related in sequence to human endogenous retroviruses HERV-L and HERV-16. Immunogenetics49, 404412. 10.1007/s002510050513

  • 69

    KulskiJ. K.DunnD. S. (2005). Polymorphic Alu insertions within the major histocompatibility complex class I genomic region: a brief review. Cytogenet. Genome Res.110, 193202. 10.1159/000084952

  • 70

    KulskiJ. K.DunnD. S.HuiJ.MartinezP.RomphrukA. V.LeelayuwatC.et al. (2002a). Alu polymorphism within the MICB gene and association with HLA-B alleles. Immunogenetics53, 975979. 10.1007/s00251-001-0409-5

  • 71

    KulskiJ. K.GaudieriS.BellgardM.BalmerL.GilesK.InokoH.et al. (1997). The evolution of MHC diversity by segmental duplication and transposition of retroelements. J. Mol. Evol.45, 599609.

  • 72

    KulskiJ. K.GaudieriS.InokoH.DawkinsR. L. (1999a). Comparison between two Human Endogenous Retrovirus (HERV)-rich regions within the major histocompatibility complex. J. Mol. Evol.48, 675683. 10.1007/PL00006511

  • 73

    KulskiJ. K.GaudieriS.MartinA.DawkinsR. L. (1999b). Coevolution of PERB11 (MIC) and HLA class I genes with HERV-16 and retroelements by extended genomic duplication. J. Mol. Evol.49, 8497. 10.1007/PL00006537

  • 74

    KulskiJ. K.MartinezP.Longman-JacobsenN.WangW.WilliamsonJ.DawkinsR. L.et al. (2001). The association between HLA-A alleles and an alu dimorphism near HLA-G. J. Mol. Evol.53, 114123. 10.1007/s002390010199

  • 75

    KulskiJ. K.ShigenariA.InokoH. (2010). Polymorphic SVA retrotransposons at four loci and their association with classical HLA class I alleles in Japanese, Caucasians and African Americans. Immunogenetics62, 211230. 10.1007/s00251-010-0427-2

  • 76

    KulskiJ. K.ShigenariA.InokoH. (2011). Genetic variation and hitchhiking between structurally polymorphic Alu insertions and HLA-A, -B, and -C alleles and other retroelements within the MHC class I region. Tissue Antigens78, 359377. 10.1111/j.1399-0039.2011.01776.x

  • 77

    KulskiJ. K.ShigenariA.InokoH. (2014). Variation and linkage disequilibrium between a structurally polymorphic Alu located near the OR12D2 gene of the extended major histocompatibility complex class I region and HLA-A alleles. Int. J. Immunogenet.41, 250261. 10.1111/iji.12102

  • 78

    KulskiJ. K.ShigenariA.ShiinaT.HosomichiK.YawataM.InokoH. (2009). HLA-A allele associations with viral MER9-LTR nucleotide sequences at two distinct loci within the MHC alpha block. Immunogenetics61, 257270. 10.1007/s00251-009-0364-0

  • 79

    KulskiJ. K.ShigenariA.ShiinaT.OtaM.HosomichiK.JamesI.et al. (2008). Human endogenous retrovirus (HERVK9) structural polymorphism with haplotypic HLA-A allelic associations. Genetics180, 445457. 10.1534/genetics.108.090340

  • 80

    KulskiJ. K.ShiinaT.AnzaiT.KoharaS.InokoH. (2002b). Comparative genomic analysis of the MHC: the evolution of class I duplication blocks, diversity and complexity from shark to man. Immunol. Rev.190, 95122. 10.1034/j.1600-065X.2002.19008.x

  • 81

    LamT.TayM.WangB.XiaoZ.RenE. (2015). Intrahaplotypic variants differentiate complex linkage disequilibrium within human MHC haplotypes. Sci. Rep.5:16972. 10.1038/srep16972

  • 82

    LamT. H.ShenM.ChiaJ.-M.ChanS. H.RenE. C. (2013). Population-specific recombination sites within the human MHC region. Heredity111, 131138. 10.1038/hdy.2013.27

  • 83

    LamT. H.ShenM.TayM. Z.RenE. C. (2017). Unique allelic eQTL clusters in human MHC haplotypes. G37, 25952604. 10.1534/g3.117.043828

  • 84

    LamontagneM.JoubertP.TimensW.PostmaD. S.HaoK.NickleD.et al. (2016). Susceptibility genes for lung diseases in the major histocompatibility complex revealed by lung expression quantitative trait loci analysis. Eur. Respir. J.48, 573576. 10.1183/13993003.00114-2016

  • 85

    LarsenC. E.AlfordD. R.TrautweinM. R.JallohY. K.TarnackiJ. L.KunnenkeriS. K.et al. (2014). Dominant sequences of human major histocompatibility complex conserved extended haplotypes from HLA-DQA2 to DAXX. PLoS Genet.10:e1004637. 10.1371/journal.pgen.1004637

  • 86

    LeeJ. M.SonnhammerE. L. L. (2003). Genomic gene clustering analysis of pathways in eukaryotes. Genome Res.13, 875882. 10.1101/gr.737703

  • 87

    LenzT. L.SpirinV.JordanD. M.SunyaevS. R. (2016). Excess of deleterious mutations around HLA genes reveals evolutionary cost of balancing selection. Mol. Biol. Evol.33, 25552564. 10.1093/molbev/msw127

  • 88

    LinY.-L.GokcumenO. (2019). Fine-scale characterization of genomic structural variation in the human genome reveals adaptive and biomedically relevant hotspots. Genome Biol. Evol.11, 11361151. 10.1093/gbe/evz058

  • 89

    LobkovskyA. E.LeviL.WolfY. I.MaiersM.GragertL.AlterI.et al. (2019). Multiplicative fitness, rapid haplotype discovery, and fitness decay explain evolution of human MHC. Proc. Natl. Acad. Sci. U.S.A.116, 1409814104. 10.1073/pnas.1714436116

  • 90

    LokkiM.PaakkanenR. (2019). The complexity and diversity of major histocompatibility complex challenge disease association studies. HLA93, 315. 10.1111/tan.13429

  • 91

    LópezS.Van DorpL.HellenthalG. (2015). Human dispersal out of Africa: a lasting debate. Evol. Bioinform. Online11 (Suppl. 2), 5768. 10.4137/EBO.S33489

  • 92

    LuJ. Y.ShaoW.ChangL.YinY.LiT.ZhangH.et al. (2020). Genomic repeats categorize genes with distinct functions for orchestrated regulation. Cell Rep.30, 32963311.e5. 10.1016/j.celrep.2020.02.048

  • 93

    MadrigalJ. A.HildebrandW. H.BelichM. P.BenjaminR. J.LittleA.-M.ZemmourJ.et al. (1993). Structural diversity in the HLA-A10 family of alleles: correlations with serology. Tissue Antigens41, 7280. 10.1111/j.1399-0039.1993.tb01982.x

  • 94

    McVeanG. (2010). What drives recombination hotspots to repeat DNA in humans?Philos. Trans. R. Soc. B Biol. Sci.365, 12131218. 10.1098/rstb.2009.0299

  • 95

    MirettiM. M.WalshE. C.KeX.DelgadoM.GriffithsM.HuntS.et al. (2005). A high-resolution linkage-disequilibrium map of the human major histocompatibility complex and first generation of tag single-nucleotide polymorphisms. Am. J. Hum. Genet.76, 634646. 10.1086/429393

  • 96

    MizukiN.OtaM.KimuraM.OhnoS.AndoH.KatsuyamaY.et al. (1997). Triplet repeat polymorphism in the transmembrane region of the MICA gene: a strong association of six GCT repetitions with Behcet disease. Proc. Nat Acad. Sci. U.S.A. 94, 12981303. 10.1073/pnas.94.4.1298

  • 97

    MoolhuijzenP.KulskiJ. K.DunnD. S.SchibeciD.BarreroR.GojoboriT.et al. (2010). The transcript repeat element: the human Alu sequence as a component of gene networks influencing cancer. Funct. Integr. Genomics10, 307319. 10.1007/s10142-010-0168-1

  • 98

    MurphyN. M.BurtonM.PowellD. R.RosselloF. J.CooperD.ChopraA.et al. (2016). Haplotyping the human leukocyte antigen system from single chromosomes. Sci. Rep.6:30381. 10.1038/srep30381

  • 99

    MyersS. (2005). A fine-scale map of recombination rates and hotspots across the human genome. Science310, 321324. 10.1126/science.1117196

  • 100

    MyersS.BowdenR.TumianA.BontropR. E.FreemanC.MacFieT. S.et al. (2010). Drive against hotspot motifs in primates implicates the PRDM9 gene in meiotic recombination. Science327, 876879. 10.1126/science.1182363

  • 101

    MyersS.FreemanC.AutonA.DonnellyP.McVeanG. (2008). A common sequence motif associated with recombination hot spots and genome instability in humans. Nat. Genet.40, 11241129. 10.1038/ng.213

  • 102

    NairR. P.StuartP. E.NistorI.HiremagaloreR.ChiaN. V. C.JenischS.et al. (2006). Sequence and haplotype analysis supports HLA-C as the psoriasis susceptibility 1 gene. Am. J. Hum. Genet.78, 827851. 10.1086/503821

  • 103

    NormanP. J.NorbergS. J.GuethleinL. A.Nemat-GorganiN.RoyceT.WroblewskiE. E.et al. (2017). Sequences of 95 human MHC haplotypes reveal extreme coding variation in genes other than highly polymorphic HLA class I and II. Genome Res.27, 813823. 10.1101/gr.213538.116

  • 104

    OhtaT. (1999). A note on the correlation between heterozygosity and recombination rate. Genes Genet. Syst.74, 209210. 10.1266/ggs.74.209

  • 105

    PaganiniJ.Abi-RachedL.GouretP.PontarottiP.ChiaroniJ.Di CristofaroJ. (2019). HLAIb worldwide genetic diversity: new HLA-H alleles and haplotype structure description. Mol. Immunol.112, 4050. 10.1016/j.molimm.2019.04.017

  • 106

    ParvanovE. D.TianH.BillingsT.SaxlR. L.SpruceC.AithalR.et al. (2017). PRDM9 interactions with other proteins provide a link between recombination hotspots and the chromosomal axis in meiosis. Mol. Biol. Cell28, 488499. 10.1091/mbc.e16-09-0686

  • 107

    PayerL. M.BurnsK. H. (2019). Transposable elements in human genetic disease. Nat. Rev. Genet.20, 760772. 10.1038/s41576-019-0165-8

  • 108

    PayerL. M.SterankaJ. P.YangW. R.KryatovaM.MedabalimiS.ArdeljanD.et al. (2017). Structural variants caused by Alu insertions are associated with risks for many human diseases. Proc. Natl. Acad. Sci. U.S.A.114, E3984E3992. 10.1073/pnas.1704117114

  • 109

    PriceP.WittC.AllockR.SayerD.GarleppM.KokC. C.et al. (1999). The genetic basis for the association of the 8.1 ancestral haplotype (A1, B8, DR3) with multiple immunopathological diseases. Immunol. Rev.167, 257274. 10.1111/j.1600-065X.1999.tb01398.x

  • 110

    ProhaskaA.RacimoF.SchorkA. J.SikoraM.SternA. J.IlardoM.et al. (2019). Human disease variation in the light of population genomics. Cell177, 115131. 10.1016/j.cell.2019.01.052

  • 111

    RakyanV. K.HildmannT.NovikK. L.LewinJ.TostJ.CoxA. V.et al. (2004). DNA methylation profiling of the human major histocompatibility complex: a pilot study for the human epigenome project. PLoS Biol.2:e405. 10.1371/journal.pbio.0020405

  • 112

    RaviramR.RochaP. P.LuoV. M.SwanzeyE.MiraldiE. R.ChuongE. B.et al. (2018). Analysis of 3D genomic interactions identifies candidate host genes that transposable elements potentially regulate. Genome Biol.19:216. 10.1186/s13059-018-1598-7

  • 113

    RobinsonJ.BarkerD. J.GeorgiouX.CooperM. A.FlicekP.MarshS. G. E. (2019). IPD-IMGT/HLA Database. Nucleic Acids Res.48, D948D955. 10.1093/nar/gkz950

  • 114

    RomeroV.LarsenC. E.Duke-CohanJ. S.FoxE. A.RomeroT.ClavijoO. P.et al. (2007). Genetic fixity in the human major histocompatibility complex and block size diversity in the class I region including HLA-E. BMC Genet.8:14. 10.1186/1471-2156-8-14

  • 115

    SachidanandamR.WeissmanD.SchmidtS. C.KakolJ. M.SteinL. D.MarthG.et al. (2001). A map of human genome sequence variation containing 1.42 million single nucleotide polymorphisms. Nature409, 928933. 10.1038/35057149

  • 116

    Sanchez-MazasA. (2020). A review of HLA allele and SNP associations with highly prevalent infectious diseases in human populations. Swiss Med. Wkly.150:W20214. 10.4414/smw.2020.20214

  • 117

    Sanchez-MazasA.BuhlerS.NunesJ. M. (2013). A new HLA map of europe: regional genetic variation and its implication for peopling history, disease-association studies and tissue transplantation. Hum. Hered.76, 162177. 10.1159/000360855

  • 118

    Sanchez-VelascoP.Gomez-CasadoE.Martinez-LasoJ.MoscosoJ.ZamoraJ.LowyE.et al. (2003). HLA alleles in isolated populations from North Spain: origin of the Basques and the ancient Iberians. Tissue Antigens61, 384392. 10.1034/j.1399-0039.2003.00041.x

  • 119

    SantosK. E.LimaT. H. A.FelicioL. P.MassaroJ. D.PalominoG. M.SilvaA. C. A.et al. (2013). Insights on the HLA-G evolutionary history provided by a nearby alu insertion. Mol. Biol. Evol.30, 24232434. 10.1093/molbev/mst142

  • 120

    SchaidD. J.McDonnellS. K.WangL.CunninghamJ. M.ThibodeauS. N. (2002). Caution on pedigree haplotype inference with software that assumes linkage equilibrium. Am. J. Hum. Genet.71, 992995. 10.1086/342666

  • 121

    SchwartzS.ZhangZ.FrazerK. A.SmitA.RiemerC.BouckJ.et al. (2000). PipMaker-a web server for aligning two genomic DNA sequences. Genome Res. 10, 577586. 10.1101/gr.10.4.577

  • 122

    ShenL.WuL.SanliogluS.ChenR.MendozaA. R.DangelA. W.et al. (1994). Structure and genetics of the partially duplicated gene RP located immediately upstream of the complement C4A and the C4B genes in the HLA class III region. J. Biol. Chem.269, 84668476.

  • 123

    ShiinaT.HosomichiK.InokoH.KulskiJ. K. (2009). The HLA genomic loci map: expression, interaction, diversity and disease. J. Hum. Genet.54, 1539. 10.1038/jhg.2008.5

  • 124

    ShiinaT.InokoH.KulskiJ. K. (2004). An update of the HLA genomic region, locus information and disease associations: 2004. Tissue Antigens64, 631649. 10.1111/j.1399-0039.2004.00327.x

  • 125

    ShiinaT.OtaM.ShimizuS.KatsuyamaY.HashimotoN.TakasuM.et al. (2006). Rapid evolution of major histocompatibility complex class I genes in primates generates new disease alleles in humans via hitchhiking diversity. Genetics173, 15551570. 10.1534/genetics.106.057034

  • 126

    SlatkinM. (2008). Linkage disequilibrium — understanding the evolutionary past and mapping the medical future. Nat. Rev. Genet.9, 477485. 10.1038/nrg2361

  • 127

    SmithW. P.VuQ.LiS. S.HansenJ. A.ZhaoL. P.GeraghtyD. E. (2006). Toward understanding MHC disease associations: partial resequencing of 46 distinct HLA haplotypes. Genomics87, 561571. 10.1016/j.ygeno.2005.11.020

  • 128

    SongS.SliwerskaE.EmeryS.KiddJ. M. (2017). Modeling human population separation history using physically phased genomes. Genetics205, 385395. 10.1534/genetics.116.192963

  • 129

    SteeleE. J.LloydS. S. (2015). Soma-to-germline feedback is implied by the extreme polymorphism at IGHV relative to MHC: the manifest polymorphism of the MHC appears greatly exceeded at Immunoglobulin loci, suggesting antigen-selected somatic V mutants penetrate Weismann. BioEssays37, 557569. 10.1002/bies.201400213

  • 130

    StewartC. A.HortonR.AllcockR. J.AshurstJ. L.AtrazhevA. M.CoggillP.et al. (2004). Complete MHC haplotype sequencing for common disease gene mapping. Genome Res.14, 11761187. 10.1101/gr.2188104

  • 131

    SuslovaT. A.BurmistrovaA. L.ChernovaM. S.KhromovaE. B.LuparE. I.TimofeevaS. V.et al. (2012). HLA gene and haplotype frequencies in Russians, Bashkirs and Tatars, living in the chelyabinsk Region (Russian South Urals): HLA gene and haplotype frequencies in Russians, Bashkirs and Tatars. Int. J. Immunogenet.39, 394408. 10.1111/j.1744-313X.2012.01117.x

  • 132

    SznarkowskaA.MikacS.PilchM. (2020). MHC class I regulation: the origin perspective. Cancers12:1155. 10.3390/cancers12051155

  • 133

    TanL.XingD.ChangC.-H.LiH.XieX. S. (2018). Three-dimensional genome structures of single diploid human cells. Science361, 924928. 10.1126/science.aat5641

  • 134

    TraherneJ. A. (2008). Human MHC architecture and evolution: implications for disease association studies. Int. J. Immunogenet.35, 179192. 10.1111/j.1744-313X.2008.00765.x

  • 135

    TraherneJ. A.HortonR.RobertsA. N.MirettiM. M.HurlesM. E.StewartC. A.et al. (2006). Genetic analysis of completely sequenced disease-associated MHC haplotypes identifies shuffling of segments in recent human history. PLoS Genet.2:e9. 10.1371/journal.pgen.0020009

  • 136

    TrowsdaleJ.KnightJ. C. (2013). Major histocompatibility complex genomics and human disease. Annu. Rev. Genomics Hum. Genet.14, 301323. 10.1146/annurev-genom-091212-153455

  • 137

    van OosterhoutC. (2009). A new theory of MHC evolution: beyond selection on the immune genes. Proc. R. Soc. B Biol. Sci.276, 657665. 10.1098/rspb.2008.1299

  • 138

    VandiedonckC.KnightJ. C. (2009). The human major histocompatibility complex as a paradigm in genomics research. Brief. Funct. Genomic. Proteomic.8, 379394. 10.1093/bfgp/elp010

  • 139

    WalshE. C.MatherK. A.SchaffnerS. F.FarwellL.DalyM. J.PattersonN.et al. (2003). An integrated haplotype map of the human major histocompatibility complex. Am. J. Hum. Genet.73, 580590. 10.1086/378101

  • 140

    WangK.MathiesonI.O'ConnellJ.SchiffelsS. (2020). Tracking human population structure through time from whole genome sequences. PLOS Genet.16:e1008552. 10.1371/journal.pgen.1008552

  • 141

    WatkinsD. I.McAdamS. N.LiuX.StrangC. R.MitfordE. L.LevineC. G.et al. (1992). New recombinant HLA-B alleles in a tribe of South American Amerindians indicate rapid evolution of MHC class I loci. Nature357, 329333

  • 142

    WGS500 ConsortiumRimmerA.PhanH.MathiesonI.IqbalZ.TwiggS. R. F.WilkieA. O. M.et al. (2014). Integrating mapping-, assembly- and haplotype-based approaches for calling variants in clinical sequencing applications. Nat. Genet.46, 912918. 10.1038/ng.3036

  • 143

    WilmingL. G.HartE. A.CoggillP. C.HortonR.GilbertJ. G. R.CleeC.et al. (2013). Sequencing and comparative analysis of the gorilla MHC genomic sequence. Database2013:bat011. 10.1093/database/bat011

  • 144

    YamadaS.KimS.TischfieldS. E.JasinM.LangeJ.KeeneyS. (2017). Genomic and chromatin features shaping meiotic double-strand break formation and repair in mice. Cell Cycle16, 18701884. 10.1080/15384101.2017.1361065

  • 145

    YunisE. J.LarsenC. E.Fernandez-VinaM.AwdehZ. L.RomeroT.HansenJ. A.et al. (2003). Inheritable variable sizes of DNA stretches in the human MHC: conserved extended haplotypes and their fragments or blocks. Tissue Antigens62, 120. 10.1034/j.1399-0039.2003.00098.x

  • 146

    ZamudioN.BarauJ.TeissandierA.WalterM.BorsosM.ServantN.et al. (2015). DNA methylation restrains transposons from adopting a chromatin signature permissive for meiotic recombination. Genes Dev.29, 12561270. 10.1101/gad.257840.114

  • 147

    ZhaoZ.FuY.-X.Hewett-EmmettD.BoerwinkleE. (2003). Investigating single nucleotide polymorphism (SNP) density in the human genome and its implications for molecular evolution. Gene312, 207213. 10.1016/S0378-1119(03)00670-X

Summary

Keywords

MHC, haplotypes, snps, retroelements, crossovers, polymorphisms, indels

Citation

Kulski JK, Suzuki S and Shiina T (2021) SNP-Density Crossover Maps of Polymorphic Transposable Elements and HLA Genes Within MHC Class I Haplotype Blocks and Junction. Front. Genet. 11:594318. doi: 10.3389/fgene.2020.594318

Received

13 August 2020

Accepted

24 November 2020

Published

18 January 2021

Volume

11 - 2020

Edited by

Shaochun Bai, GeneDx, United States

Reviewed by

Pu-Feng Du, Tianjin University, China; Xingyun Qi, The State University of New Jersey, United States

Updates

Copyright

*Correspondence: Jerzy K. Kulski

This article was submitted to Human and Medical Genomics, a section of the journal Frontiers in Genetics

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics