Abstract
Subtelomeres of most eukaryotes contain fast-evolving genes usually involved in adaptive processes. In common bean (Phaseolus vulgaris), the Co-2 anthracnose resistance (R) locus corresponds to a cluster of nucleotide-binding-site leucine-rich-repeat (NL) encoding sequences, the prevalent class of plant R genes. To study the recent evolution of this R gene cluster, we used a combination of sequence, genetic and cytogenetic comparative analyses between common bean genotypes from two distinct gene pools (Andean and Mesoamerican) that diverged 0.165 million years ago. Co-2 is a large subtelomeric cluster on chromosome 11 comprising from 32 (Mesoamerican) to 52 (Andean) NL sequences embedded within khipu satellite repeats. Since the recent split between Andean and Mesoamerican gene pools, the Co-2 cluster has experienced numerous gene-pool specific NL losses, leading to distinct NL repertoires. The high proportion of solo-LTR retrotransposons indicates that the Co-2 cluster is located in a hot spot of unequal intra-strand homologous recombination. Furthermore, we observe large segmental duplications involving both Non-Homologous End Joining and Homologous Recombination double-strand break repair pathways. Finally, the identification of a Mesoamerican-specific subtelomeric sequence reveals frequent interchromosomal recombinations between common bean subtelomeres. Altogether, our results highlight that common bean subtelomeres are hot spots of recombination and favor the rapid evolution of R genes. We propose that chromosome ends could act as R gene incubators in many plant genomes.
Introduction
Plant immunity is activated after direct or indirect perception of pathogen molecules by resistance (R) genes (; ). The proteins encoded by the largest class of plant R genes contain nucleotide binding site (NB) and leucine rich repeat (LRR) domains, together with a N terminus that usually contains a predicted coiled-coil structure (CC domain) or shares similarity with Drosophila Toll/mammalian interleukin 1 receptors (TIR domain). Genes encoding CC-NB-LRR (CNL) and TIR-NB-LRR (TNL) proteins correspond to two ancient lineages (; ; ) and are often organized in clusters of two to dozens of copies evolving at different rates (; Smith et al., 2004; ; Ratnaparkhe et al., 2011; ; ). Clustering of duplicated R genes, as observed in numerous plant species (Pryor and Ellis, 1993; ), has been suggested to favor R gene evolution against an ever changing array of pathogens (). Indeed, after duplication, mutations occurring in paralogs and/or subsequent recombination between clustered genes can give rise to novel R gene specificities (Sudupak et al., 1993; Richter et al., 1995; ; ; ; , ; ). However, little is known about the dynamics of R gene clusters on a short (infra-species) time scale (; Zhou et al., 2007).
Common bean (Phaseolus vulgaris) is a major food crop in many areas of the Americas, Europe, Africa and Asia that supports the health and income of ∼400 million people in eastern Africa and ∼250 million in Central and South America. Two major gene pools have been identified for cultivated common bean, Andean (South America) and Mesoamerican (Mexico and Central America) (). These two gene pools diverged ∼0.165 million years ago (Mya) (; Schmutz et al., 2014). In the common bean genome, most of the large R gene loci map at the end of linkage groups (LG) (). This is the case for the Co-2 and B4 R loci, localized at one end of LG-B11 and LG-B4, respectively (; ). Many specific R genes and Quantitative Trait Loci (QTL) conferring resistance to a diverse selection of pathogens, including fungi and bacteria, have been mapped to these two loci (; , , ; ; ; ). Sequencing of ∼0.95 Mb of the Co-2 locus in an Andean genotype (G19833) revealed that it corresponds to an ancestral CNL cluster syntenic to the Rpg1 R cluster in soybean (; ). Comparative genomics analyses revealed that a CNL sequence from the Co-2 cluster has migrated to the B4 locus through an ectopic recombination event between non-homologous chromosomes ∼20–54 Mya (). Strikingly, in soybean this relocated CNL sequence has been pseudogenized, while in common bean it underwent an impressive amplification during the last 20 million years, leading to a Phaseolus-specific CNL subfamily comprising more than 29 members clustered within the B4 R cluster. We have hypothesized that the different fates of this CNL sequence in soybean (pseudogenization) and common bean (amplification) was due to divergent genomic environments (; ). Indeed, in soybean the relocated CNL was pseudogenized in a euchromatin region, while in common bean the B4 cluster arose in a region close to the telomere, called the subtelomere.
Subtelomeres, regions proximal to telomeres, are transition regions between the telomere and the chromosome-specific DNA sequence. Subtelomeres are difficult to define, varying in length from 20 kb in some yeast strains to several 100 kb in higher eukaryotes (; ). Sequence assembling of these regions is particularly challenging because they often contain a high density of repeated sequences (; ; ). Thus, subtelomeres are often lacking from so-called whole-genome sequences and remain relatively understudied, especially in plants (; ; ). Tandemly organized repeats, known as satellites, have been identified by cytogenetic analyses in the subtelomeres of several plants species (; ). In other kingdoms, subtelomeres are often viewed as variable loci containing fast-evolving gene families involved in adaptative processes (). Moreover, some plant genomes such as maize, Arabidopsis, or rice contain blocks of heterochromatin, known as “knobs,” containing highly repeated satellite DNA (; ; ). Interestingly, most common bean subtelomeres bear terminal heterochromatic knobs. Fluorescent In Situ Hybridization (FISH) analyses revealed that the B4 R gene cluster is localized adjacent to heterochromatic knobs and associated with a Phaseolus-specific, 528-bp subtelomeric satellite repeat named khipu. The khipu satellite is found at most common bean terminal knobs, indicating frequent exchanges between common bean subtelomeric regions (; Richard et al., 2013).
Given the distal localization of the Co-2 R gene cluster at one end of LG-B11, we wanted to further characterize the organization of this cluster and to understand its recent evolution since the divergence of Andean and Mesoamerican gene pools (0.165 Mya). To this end, we have extended the sequence of the Co-2 cluster from ∼0.95 Mb to ∼1.35 Mb in G19833 (Andean) and sequenced ∼1.1 Mb in BAT93 (Mesoamerican). We then used a combination of comparative genomics approaches including computational analyses and FISH to study the recent molecular evolution of this locus. Altogether, our results highlight that subtelomeres are hot spots of intra- and inter-chromosomal recombination and favor the rapid evolution of R genes.
Materials and Methods
Bacterial Artificial Chromosome (BAC) Contig Assembly and Sequencing
To extend the G19833 BAC contigs previously sequenced in , we selected five additional overlapping BAC clones based on G19833 BAC end sequence (BES) library (Schlueter et al., 2008). To identify the corresponding region in a Mesoamerican common bean genotype (BAT93), we screened a BAT93 BAC library () using DNA hybridization probes derived from low-copy protein-coding genes identified in the soybean Rpg1 locus sequence (). BAC clones that hybridized to two or more probes were fingerprinted and end-sequenced. Identification of a minimum-tiling path for sequencing was done as described in . Sequencing and assembly were performed at Genoscope (Evry, France). Sequenced BAC clones and contigs are listed in Supplementary Table S1.
Annotation Procedure
Gene prediction was done using a combination of gene-finding programs and sequence homology with known genes and proteins. Two ab initio gene prediction programs FGENESH () and GeneMarkhmm () were used. BLAST () analyses against the GenBank non-redundant database were performed and the Phaseolus Expressed Sequence Tags (EST) available at GenBank (Ramirez et al., 2005) were aligned to the predicted genes as described in Samson et al. (2004). The criteria used to define a gene when there was no EST support were, first, a match to a sequence in a protein database using BLASTX (of 1e-3) on the non-redundant protein sequence database (nr) and, second, prediction as a gene by the two prediction programs. All this information was imported into the annotation platform Artemis for further manual analysis (Rutherford et al., 2000). Sequences were considered as pseudogenes when starting with a methionine, but presenting premature stop codons and/or frameshifts. Numerous sequences presenting homology to CNL sequences but that did not start with a methionine were retrieved and annotated as CNL segments.
Identifying Retroelements, Repeats, and Segmental Duplications
Repetitive elements and transposon sequences were identified using BLASTN (of 1e-10) on the P. vulgaris repeat database (), followed by manual inspection. BLASTN results corresponding to short degenerated sequences (<1000 bp) were discarded from the analysis. The program LTR_STRUC was used as the first step in identifying Long Terminal Repeats (LTRs) retrotransposon sequences (). Then, LTRs from the elements identified by LTR-STRUC, plus 134 LTRs previously identified in Glycine and Phaseolus genomes (Wawrzynski et al., 2008) were used as queries in BLASTN searches with a cutoff value of 1e-10 (). LTRs were classified as described in Wawrzynski et al. (2008). In order to optimize LTR finding, we performed a second BLASTN search using the entire dataset of previously identified LTRs. Regions of homology to known retrotransposon-like sequences (e.g., reverse transcriptase, integrase, etc.) were then manually evaluated for the presence of LTRs and Target Site Duplications (TSDs). These additional searches uncovered several intact elements missed by the LTR_STRUC program as well as solo-LTRs (Supplementary Table S9). The khipu satellite DNA was uncovered using hmmsearch “http://hmmer.janelia.org” with a khipu profile previously defined on 92 khipu units as explained in . Segmental duplications were searched by BLASTN analyses within and between sequenced contigs from each genotype after masking khipu repeats and CNL genes, using the criterion of presenting more than 90% nucleotide identity over 500 bp minimum.
Microsynteny Analysis
After annotation, we identified low-copy genes as genes presenting less than 20 hits in the Arabidopsis genome after BLASTP (of 1e-5) analysis using the FLAGdb++ database (Samson et al., 2004). Low-copy genes from BAT93 and G19833 genotypes were used for aligning the sequenced contigs. Orientation of G19833 contigs, previously aligned to the soybean and Medicago truncatula genomes (), was confirmed by BLASTN analyses on the available G19833 genome (Schmutz et al., 2014) and used as reference for the orientation of BAT93 contigs (Supplementary Table S1). Twelve pairs of non-pseudogenized, non-truncated low-copy genes (Supplementary Table S4) were aligned with MAFFT version 6 using the L-INS-I strategy (). Ks values were determined using DNAsp version 5 using default parameters ().
Phylogenetic Analysis of NL Sequences
Multiple alignments of amino acid sequences from the NB portion of CNLs or TNLs (from the P-loop to the MHD motif) were performed with MAFFT version 6 using the L-INS-I strategy () with default parameters. Aligned protein sequences were used as a guide to align the corresponding DNA sequences. Recombination among loci was assessed using several methods implemented in RDP version 3.15 (): RDP (), Geneconv (), Chimera (Posada and Crandall, 2001), and Bootscan () as described in . For CNLs, the alignment was then trimmed to an approximately 166 amino acid region of the NB domain devoid of recombination events. Alignments were then subjected to maximum likelihood (ML) analysis using the JTT model as implemented in PhyML (). Relative support for clades was assessed with 1,000 bootstrap replicates. The resulting phylogenetic trees (Figure 1A and Supplementary Figure S2) were displayed using MEGA 3.1 (Tamura et al., 2007). The TNL tree (Supplementary Figure S2) comprises the TNL sequences from the Co-2 cluster (this study) and a diversity of common bean TNL sequences (), plus TNLs from the soybean Rpg1 cluster () and TNLs from the whole genome of Medicago truncatula (). Arabidopsis TAO1 gene was used as nearest outgroup (). The CNL tree (Figure 1A) comprises sequences from the Co-2 cluster, the soybean and Glycine tomentella Co-2 syntenic clusters (; ) as well as sequences from the B4 cluster (). The two closest CNLs from the Co-2 family in the M. truncatula genome were used as nearest outgroup (). CNL sequences with an intact and non-recombinant NB domain were grouped into subfamilies CNL1 or CNL3/4 () using phylogeny (Figure 1A). Whole CNL sequences were aligned using Needle () and pseudogenes and/or recombinants were assigned to subfamilies CNL1 or CNL3/4 using a nucleotide identity threshold of 72% (Supplementary Table S7). Ks were determined on whole CNL alignments using DNAsp version 5 with default parameters ().
FIGURE 1
Genetic Mapping
A PCR-based approach was used to map BAT93 contigs using specific oligonucleotide primers. For TNL mapping, a NB-specific 496-bp probe (Supplementary Figure S5) was amplified from TNL sequence PvM11B110 from BAC clone Pva1-231b2 using primers M231b2_3S (5′-TGGTTGCTTCCTCACAAATG-3′) and M231b2_3A (5′-TTCACTTTCCCATGCCTCTT-3′). This TNL probe was used in Southern Blot hybridization experiments as described in . Chi-squared (χ2) tests were used to evaluate the goodness of fit of observed and expected segregation ratios. MAPMAKER software version 3.0 () was used to map segregating markers as described in . All markers used in this study are described in Supplementary Table S10 and were mapped using a population of 179 Recombinant Inbred Lines (RILs) derived from a cross between BAT93 (Mesoamerican) and JaloEEP558 (Andean) ().
Cytogenetic Analysis
DNA probes used in FISH experiments are described in Supplementary Table S8. FISH-CNL1, FISH-CNL3/4 and FISH-TNL probes correspond to pools of two to seven BAC subclones of ∼6 kb each, coming from sequenced BACs of the Andean genotype G19833. These probes contain NL sequences as well as few intergenic sequences (Supplementary Figure S5). Therefore, subclones were selected for the absence of repeats by performing BLASTX or BLASTN analyses of the intergenic portions of these subclones against (i) the GenBank non-redundant database, (ii) G19833 WGS sequence, (iii) a library of common bean repeat sequences based on 89,017 BES from G19833, and (iv) 1,404 shotgun sequences from the BAT7 Mesoamerican genotype (Schlueter et al., 2008; Schmutz et al., 2014). Moreover, to select against local repeats, BLASTN analyses were performed against the entire set of BAC-sequenced contigs from G19833 and BAT93. The khipu probe was generated using a pool of five BAC subclones comprising three BAC subclones coming from different khipu blocks spread over the sequenced contigs from the Co-2 cluster, plus one additional subclone from another subtelomeric BAC from the short arm of chromosome 5 (), and the 1H04 subclone from chromosome 4 (). The 45S rDNA probe corresponds to the R2 probe, a 6.5-kb fragment of an 18S–5.8S–25S rDNA repeat unit from Arabidopsis (Wanzenböck et al., 1997). Pachytene chromosomes were prepared from young flower buds fixed in ethanol:acetic acid (3:1, v/v). Buds were macerated in 2% cellulase/2% pectolyase/2% cytohelicase in 0.01 M citric acid-sodium citrate buffer, pH 4.8, for 3 hr at 37°C, incubated in 60% acetic acid up to 2 h, and squashed after removal of petals and sepals and flaming. Slide selection and pretreatment, chromosome and probe denaturation and hybridization, posthybridization washes, detection and image analyses were performed as described in .
Results
Sequencing and Annotation of the Co-2 Resistance Locus in Two Common Bean Genotypes
In a common bean genotype of Andean origin (G19833), we have previously sequenced, annotated and mapped five BAC contigs corresponding to a 0.95-Mb region centered on the Co-2 R locus (; ). Here, we sequenced five additional overlapping BAC clones, leading to a total of 1.35 Mb organized in five contigs (PvA11A to PvA11E, Figure 1B and Supplementary Table S1); however, the four gaps between the contigs remained unclosed. We have compared our BAC-based sequence data with the whole genome shotgun (WGS) sequencing data of G19833 for this region (Schmutz et al., 2014). The five contigs were located at the end of chromosome 11 long arm of the WGS sequence, and the four gaps were found to be 48833, 101827, 18552, and 120389 bp, respectively (Supplementary Table S1 and Figure 1B).
In order to study the evolution of the Co-2 R-gene cluster since the divergence between Andean and Mesoamerican gene pools, we selected and sequenced nine BACs from a Mesoamerican common bean genotype (BAT93), leading to 1.07 Mb organized in four contigs (PvM11A to PvM11D, Figure 1B). PCR-based mapping confirms that macrosynteny between G19833 and BAT93 contigs is consistent with microsynteny (Figures 1B,C). Automatic and manually curated sequence annotation predicted 116 genes and pseudogenes in the G19833 sequence and 90 genes and pseudogenes in the BAT93 sequence. For both genotypes, the genic part represents one third of the total sequence, which corresponds to an average density of one gene per 12 kb (Supplementary Table S2 and Supplementary Figure S1). The overall organization of the Co-2 locus consists of low-copy gene islands separated by numerous genes from the CNL, kinase and sulfotransferase multigenic families, transposable elements, and the khipu 528-bp subtelomeric satellite (; Figure 1B). These numerous repeated sequences impacted the quality of the WGS sequence in G19833. As a result, 76 gaps corresponding to 167 kb were missing in the WGS sequence corresponding to our 1.35 Mb BAC-based sequences.
CNL sequences from the Co-2 family represent around one third of the predicted genes in both genotypes, with 43 CNLs in G19833 and 30 CNLs in BAT93 (Figure 1B and Supplementary Figure S1). For G19833, most CNL sequences were 100% identical between the WGS and our BAC-based sequence. However, 10 CNLs were either missing or contained errors in the WGS while 13 additional CNLs were retrieved in the WGS sequence outside our BAC contigs, leading to a total number of 56 CNLs in G19833 (Supplementary Table S3).
The CNL sequences from the Co-2 cluster were previously classified into four subfamilies (referred to as CNL1 to CNL4) that diverged prior to the split between common bean and soybean, 20 Mya (; ; ). CNL2 sequences were not found in the common bean genome, indicating that these sequences were specifically lost in common bean (). We refer to CNL3/4 to describe CNL3 plus CNL4 sequences, because CNL3- and CNL4-specific probes cross-hybridize on common bean DNA (). As previously shown in , we observed two well-defined blocks of proximal CNL1 and distal CNL3/4 sequences, suggesting that these CNL copies arose mainly by tandem duplications.
Interestingly, in addition to the prevalent CNL sequences, we also identified one TNL sequence in G19833, collinear to two tandemly duplicated TNL sequences in BAT93, and an additional TNL sequence was retrieved from G19833 WGS annotation (in yellow on Figure 1B). Phylogenetic analyses showed that these TNLs belong to the M. truncatula TNL-7 subfamily defined in (Supplementary Figure S2). In total, the number of NL sequences at the Co-2 locus is 58 (56 CNL plus 2 TNL) in G19833 and 32 (30 CNL plus 2 TNL) in BAT93. Notably, in G19833 52 NL sequences are clustered within 1,35 Mbp (Supplementary Table S3). Consequently, Co-2 is one of the largest NL clusters identified to date in plants, together with the maize Rp1 cluster (Smith et al., 2004).
The Co-2 Cluster Arose From CNL Duplications That Occurred 20 Mya to 0.165 Mya Ago, Followed by Gene-Pool Specific CNL Deletions and Pseudogenization
The nucleotide substitution rate at silent sites (Ks) is a measure of neutral evolution for coding DNA because it doesn’t take in account the non-synonymous sites that can be under selection pressure. Therefore, Ks is often used to estimate the time of divergence between two similar genes, these genes being either orthologs (for dating the divergence between taxa) or paralogs (for dating a duplication event). Lower Ks values correspond to shorter times of divergence between sequences. Here, we wanted to know if the duplication of CNL paralogs from the Co-2 cluster occurred before or after the divergence between Andean and Mesoamerican gene pools, 0.165 Mya. To this end, we used the Ks from 12 ortholog pairs of low-copy genes as reference to infer orthology and paralogy relationships between CNL sequences (Figure 1B and Supplementary Table S4). The mean Ks value for these 12 ortholog pairs is 0.0177, with a maximum of 0.0494. Thus, we consider that a Ks below 0.0494 is indicative of events that are either simultaneous or more recent than 0.165 Mya.
Between gene pools, 16 CNL pairs have Ks values under 0.0494 and can thus be considered as orthologs (Figure 1B and Supplementary Tables S5, S6). Consistently, all these orthologs are found in syntenic positions (Figure 1B) and share more than 96% nucleotide identity within a pair while all other CNL sequences share less than 95% identity with each other (Supplementary Table S7). These 16 CNL pairs comprise 11 intact full-length genes and 21 pseudogenes (i.e., CNLs presenting frameshifts and/or truncated before the usual end of CNL coding sequence). As a result, only three pairs are composed of intact CNL genes in both genotypes while one and four CNLs are specifically intact in G19833 and BAT93, respectively. Manual inspection indicated that only four pseudogenization events were common to both members of a pair and consequently occurred before 0.165 Mya, while 13 occurred more recently and independently in the Andean and Mesoamerican gene pools (Supplementary Table S6). Within each genotype, we detected only one pair of CNL paralogs with a Ks value under 0.0494, thus corresponding to a recent duplication that occurred after the divergence of both gene pools (Supplementary Tables S5, S6). This recent event is specific for G19833 and corresponds to the ectopic duplication of PvA11B090 to PvA11D160 (green arrow in Figure 1B). Consequently, except for this case, all CNLs appear to have emerged by duplications predating the divergence between Andean and Mesoamerican gene pools.
Low-copy genes and CNL orthologs allowed us to delineate four collinear regions spanning 865 kb in total (red zones in Figure 1B). No recent paralogs were found within these regions, as testified by the Ks values over 0.0494 for every paralog pair (Supplementary Table S5). Thus, presence of a CNL in only one genotype indicates that the corresponding ortholog has been lost in the other genotype. Following this rule, we identified six G19833-specific losses (corresponding to PvM11A120∗, PvM11A150∗, PvM11B050∗, PvM11B060∗, PvM11B070, PvM11B090∗ in BAT93) and three BAT93-specific losses (corresponding to PvA11C130, PvA11C220∗, PvA11D110 in G19833). In all, we identified nine recent CNL losses within the red zones.
In all, genomic comparison between one Andean and one Mesoamerican genotype revealed intensive evolution of the CNL sequences from the Co-2 cluster within the last 0.165 My. In particular, we detected one gene pool-specific CNL duplication, nine gene pool-specific CNLs losses, and 13 pseudogenization events, indicating a strong erasure of CNLs from 0.165 Mya to the present day. As a result, only three intact CNLs are conserved in both the Andean and Mesoamerican genotypes.
FISH Experiments Show Differentiations in the NL Repertoire of Mesoamerican Versus Andean Subtelomeres
In order to determine the physical distribution of CNL1, CNL3/4 and TNL-7 sequences in the common bean genome, we used specific probes for FISH experiments on pachytene chromosomes from G19833 (Andean), JaloEEP558 (Andean) and BAT93 (Mesoamerican) genotypes (Supplementary Table S8). We also used the subtelomeric DNA satellite khipu () and the subtelomeric 45S rDNA (Pedrosa-Harand et al., 2006) as probes for labeling subtelomeric regions. In all genotypes, the end of chromosome 11 long arm is composed of a large terminal heterochromatic knob, followed by a region not uniformly condensed (Figure 2a). In agreement with the analysis of khipu sequences in G19833 WGS (Richard et al., 2013), chromosome 11 long arm is the genomic region presenting the largest khipu signal in the entire common bean genome. Here, khipu signal is similar in all genotypes, extending from the terminal knob chromosome 11 long arm to a large region proximal to the knob (Figure 2b). In JaloEEP558, the terminal knob is caped with 45S rDNA, while no 45S signal appears in the other genotypes (Figure 2c).
FIGURE 2
FISH-CNL1 and FISH-CNL3/4 probes show large signals overlapping with khipu (Figures 2d,e). These large signals likely correspond to the Co-2 cluster. In BAT93, the Co-2 cluster is located closer to the terminal knob than in the Andean genotypes. Furthermore, the FISH-CNL3/4 signal is wider in G19833 than in the two other genotypes, suggesting that specific duplications of CNL3/4 sequences occurred in this genotype. In contrast to CNL1 sequences that are localized at only one locus in the common bean genome, additional smaller FISH-CNL3/4 signals appear at positions proximal to the Co-2 cluster, indicating that CNL3/4 sequences moved by ectopic duplications (Figure 2e). FISH-TNL-7 signals appear as spots either overlapping with or proximal to the Co-2 cluster (Figures 2f,g), suggesting that TNL-7 sequences amplified more by ectopic duplication than tandem duplication. All these results are consistent with RFLP-based mapping showing that CNL1 sequences map only at one genetic position while CNL3/4 and TNL-7 sequences are distributed at both proximal and distal positions compared to CNL1 sequences (Supplementary Figure S3).
Unexpectedly, the terminal knob of chromosome 11 long arm was labeled not only with khipu but also with the FISH-CNL3/4 probe (Figure 2h). At the whole genome level, 14 terminal knobs were labeled with strong FISH-CNL3/4 signals in BAT93 (Mesoamerican), and 16 in another Mesoamerican genotype (G11051; Figure 3). This is in sharp contrast with Andean genotypes where only two and three very weak FISH-CNL3/4 signals were observed at JaloEEP558 and G19833 terminal knobs, respectively (Figure 3). These FISH signals likely corresponded to a portion of a CNL3/4 gene or to an intergenic portion of the subclones used for FISH-CNL3/4 probe (Supplementary Figure S4). Unfortunately, we were unable to retrieve these sequences in the WGS of G19833, because the pseudomolecules contain numerous gaps at their ends (Schmutz et al., 2014). Interestingly, these results suggest that a sequence linked to CNL3/4 genes has become a subtelomeric repeat that amplified much more in Mesoamerican than Andean genotypes, both locally (stronger signals) and between non-homologous chromosomes (multiple labeled terminal knobs).
FIGURE 3
Altogether, these results indicate that frequent local and inter-chromosomal recombination events occurred at common bean subtelomeres. In particular, gene pool specific ectopic duplications of CNL3/4 and TNL-7 sequences led to different NL repertoires in Mesoamerican versus Andean subtelomeres.
Local and Ectopic Segmental Duplications Involving NL Sequences Have Occurred in Common Bean
Within the subtelomere of chromosome 11 long arm, spots of contiguous FISH-TNL-7, FISH-CNL3/4 and khipu signals at positions proximal to the main Co-2 locus indicate that ectopic duplications involving TNL-7, CNL3/4, and khipu sequences occurred (Figure 2g). The signals from these three different probes are contiguous, suggesting that these ectopic duplications have been due to segmental duplications (SDs) involving TNL-7, CNL3/4, and khipu sequences. We summarize these hypothetical SD events as a model in Figure 2i. SD event A is observed for all genotypes, suggesting that it predates the divergence between Andean and Mesoamerican genotypes, while events B, and C likely occurred specifically in JaloEEP558 and BAT93, respectively. In G19833 WGS, in addition to the two full length TNL-7 present in the annotation, BLASTn analysis allowed us to retrieve four TNL-7 segments that were located in positions congruent with the FISH-TNL-7 signals (Supplementary Figure S5). This allows us to show that event A corresponds to the duplication of a couple of TNL-7 and CNL3/4 sequences in a tail to tail fashion (Supplementary Figure S5d). Two additional ectopic duplications involving these TNL-7 segments and CNL3/4 sequences occurred in G19833 (Supplementary Figure S5d). The intergenic portion between these TNL-7 and CNL3/4 sequences is partially conserved between duplicates, confirming that these sequences have been duplicated by SDs (data not shown). Interestingly, among the eight NL sequences found in these SDs, six were pseudogenized or interrupted by a TE insertion, leading to only one intact TNL-7 and one intact CNL3/4.
Within the BAC clone sequences, computational analysis allowed us to detect additional SDs in the Co-2 cluster. We identified two ancient SDs shared by BAT93 and G19833 at different genomic positions (blue events in Figure 1B). These blue events correspond to the double duplication of a 10 kb region containing a CNL3/4 sequence and share ∼80% nucleotide identity between duplicated regions. Moreover, we identified three recent Andean-specific SDs defined as the green (22 kb), the purple (16 kb), and the gray (23 kb) SDs in G19833 (Figure 1B). The green SD resulted in the aforementioned duplication of PvA11B90 to PvA11D160 (Figures 1A,B). The absence of corresponding SDs in BAT93 indicates that these three events occurred less than 0.165 Mya, which is in agreement with the high nucleotide identity found between duplicated regions (>98%). Interestingly, all these SDs involved a region of only 180 kb, suggesting that this region is a hot-spot of recombination (Figure 1B). Moreover, the blue and green SDs contain intact CNL sequences (blue: PvMC040, PvMC110, PvA11D120; green: PvA11B090, and PvA11D160) suggesting that these SDs may have contributed to the duplication of functional R genes.
Segmental Duplications Involve Both Homologous and Non-homologous Recombination Pathways at the Co-2 Locus
Ectopic recombination can result from aberrant DNA repair by either Non-homologous End-Joining (NHEJ) or non-allelic homologous recombination (HR), which are two general pathways of double-strand break (DSB) repair in plants and other living organisms (; ; ; Puchta, 2005; ). In the case of SD, the broken recipient molecule is invaded at the breakpoint by an ectopic donor molecule, resulting in a hybrid molecule with two junctions (X and Y) corresponding to two repair points for each DSB (Figure 4A). Defining a breakpoint accurately requires comparison of the hybrid sequence to the original donor and recipient sequences. The presence of sequence similarity that extends beyond the homology boundary between duplicated regions is strongly suggestive of non-allelic HR, while the absence of transition or microhomology is indicative of NHEJ (; ). Using these criteria, we deduced the probable repair process at five breakpoints from the green, purple, and gray SDs (Figures 1B, 4).
FIGURE 4
For junction 1X, as the SD is G19833-specific, we retrieved the corresponding donor sequence in BAT93 (Figure 1B). Microhomology of only three identical nucleotides indicates that junction 1X was likely resolved by NHEJ (Figure 4B). This is illustrated by an abrupt drop in identity from more than 98% to around 50% between the hybrid and corresponding recipient and donor sequences. We used this later criterion for analyzing the putative repair pathways for the other breakpoints because we couldn’t find corresponding recipient sequences. Junctions 2X and 3X present the same identity percentage profile as junction 1X and were thus also resolved by NHEJ (Figures 4C,D). Conversely, junction 3Y was resolved by HR-mediated repair as testified by the presence of a 130 bp transition region with 80% nucleotidic identity between the hybrid and the donor sequence (Figure 4E). Junctions 3X and 3Y delineate a single SD indicating that a single SD could involve both HR and/or NHEJ pathways.
Removal of LTR Retrotransposons by Homologous Recombination Is Occurring at a High Rate at the Co-2 Locus
Within the BAC contig sequences, we identified 88 transposable elements in G19833 (1/15 kb) and 58 in BAT93 (1/18 kb), with a predominance of Long Terminal Repeats (LTR) retrotransposons in both G19833 (42) and BAT93 (41). Among them, 20 and 25 comprise identifiable LTR in G19833 and BAT93, respectively (Supplementary Table S9). The ratio between intact LTR retrotransposons and solo-LTRs (I/S ratio) is often used as an indicator of DNA removal caused by unequal HR (presumably intrastrand) between LTR sequences of an individual element (
Discussion
In various organisms (human, yeast, Plasmodium), subtelomeres have been shown to have great potential for facilitating rapid adaptation because of their highly dynamic nature (
At the genome level, the presence of khipu satellite sequences at most terminal knobs combined with the fact that khipu is specific to the Phaseolus genus is evidence of intense interchromosomal shuffling between common bean subtelomeres (
In plants and other organisms such as Drosophila melanogaster and human, the presence of repeated sequences and proximity to heterochromatin is thought to favor unequal recombination as well as SDs (Puchta, 2005;
A significant finding, based on Ks analyses, was that numerous NL genes within the Co-2 cluster have been specifically removed or pseudogenized in a gene pool specific manner while new NL sequences have been duplicated at ectopic regions, leading to continued amplification and diversification of the Co-2 family. As a result, only three intact CNL genes were found in both G19833 and BAT93 while one and four CNLs were specifically intact in each genotype, indicating that Andean and Mesoamerican genotypes bear different NL repertoires. This suggests that most R specificities have diverged between these two gene pools. In agreement with this, previous cross-inoculation experiments between common bean plants and Colletotrichum lindemuthianum strains at the level of the centers of diversity of the plants indicated that ongoing processes of coevolution between the two protagonists have led to a differentiation for resistance (
With more than 52 NL sequences in G19833, the Co-2 cluster is one of the largest R gene clusters identified so far in plants, together with the maize Rp1 cluster (up to 52 NL sequences), the lettuce RGC2 cluster (up to 32 NL sequences) and the bean B4 cluster (more than 29 CNL sequences) (
Conclusion
Our results highlight the role of SDs in NL evolution at subtelomeric regions and suggest that common bean subtelomeres constitute a genomic environment where elevated intra- and inter-chromosomal recombination frequencies provide a perfect niche to fuel rapid NL diversification. This is in agreement with what has been observed in other organisms (human, yeast, Plasmodium…) where subtelomeric regions present several unusual characteristics, such as the presence of repetitive DNA, increased rate of evolution and enrichment for genes involved in adaptation. However, plasticity of subtelomeres has rarely been reported in plants. In fact, many papers report that large NL clusters are located at chromosome ends in plants, and here we propose that subtelomeres could act as R gene incubators in plant genomes. It’s interesting to note that in some eukaryotic pathogens, genes for host-translocated effectors are also concentrated at chromosome ends (
Accessions Numbers
Sequencing data generated for this study have been submitted to GenBank under accession numbers FO9255992 to FO926003, FO681292 and KP164990 (Supplementary Table S1).
Statements
Author contributions
NC and VG designed the study and wrote the manuscript. NC, VT, TR, GM, TA, RI, AP-H, and VG contributed to data analyses. All authors read and approved the final version of the manuscript.
Funding
This work was supported by grants from Institut National de la Recherche Agronomique, Centre National de la Recherche Scientifique, Saclay Plant Sciences LABEX (SPS), Institut Diversité Ecologie et Evolution du Vivant (IDEEV) and a grant from Genoscope/CEA-Centre National de Séquençage (VG), and partly by CNPq and CAPES (AP-H).
Acknowledgments
We thank Mireille Sévignac for excellent technical assistance. We thank Benoît Alunni and Perrine David for useful discussions. We thank James Kami and Paul Gepts (UC Davis, United States) for providing the BAT93 BAC library filters. We also thank Scott A. Jackson and Dongying Gao for helpful discussions on Phaseolus vulgaris repeated sequences.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Supplementary material
The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fpls.2018.01185/full#supplementary-material
References
1
AltschulS. F.MaddenT. L.SchäfferA. A.ZhangJ.ZhangZ.MillerW.et al (1997). Gapped BLAST and PSI-BLAST : a new generation of protein database search programs.Nucleic Acids Res.253389–3402. 10.1093/nar/25.17.3389
2
Ameline-TorregrosaC.WangB.-B.O’BlenessM. S.DeshpandeS.ZhuH.RoeB.et al (2008). Identification and characterization of nucleotide-binding site-leucine-rich repeat genes in the model plant Medicago truncatula.Plant Physiol.1465–21. 10.1104/pp.107.104588
3
AmmirajuJ. S. S.LuF.SanyalA.YuY.SongX.JiangN.et al (2008). Dynamic evolution of oryza genomes is revealed by comparative genomic analysis of a genus-wide vertical data set.Plant Cell203191–3209. 10.1105/tpc.108.063727
4
AshfieldT.EganA. N.PfeilB. E.ChenN. W. G.PodichetiR.RatnaparkheM. B.et al (2012). Evolution of a complex disease resistance gene cluster in diploid Phaseolus and tetraploid Glycine.Plant Physiol.159336–354. 10.1104/pp.112.195040
5
AshfieldT.OngL. E.NobutaK.SchneiderC. M.InnesR. W. (2004). Convergent evolution of disease resistance gene specificity in two flowering plant families.Plant Cell16309–318. 10.1105/tpc.016725.1
6
BaggsE.DagdasG.KrasilevaK. V. (2017). NLR diversity, helpers and integrated domains: making sense of the NLR IDentity.Curr. Opin. Plant Biol.3859–67. 10.1016/j.pbi.2017.04.012
7
BaiJ.PennillL. A.NingJ.LeeS. W.RamalingamJ.WebbC. A.et al (2002). Diversity in nucleotide binding site-leucine-rich repeat genes in cereals.Genome Res.121871–1884. 10.1101/gr.454902
8
BennetzenJ. L.MaJ.DevosK. M. (2005). Mechanisms of recent genome size variation in flowering plants.Ann. Bot.95127–132. 10.1093/aob/mci008
9
BressonA.JorgeV.DowkiwA.GuerinV.BourgaitI.TuskanG. A.et al (2011). Qualitative and quantitative resistances to leaf rust finely mapped within two nucleotide-binding site leucine-rich repeat (NBS-LRR)-rich genomic regions of chromosome 19 in poplar.New Phytol.192151–163. 10.1111/j.1469-8137.2011.03786.x
10
BrittenR. J. (1998). Precise sequence complementarity between yeast chromosome ends and two classes of just-subtelomeric sequences.Proc. Natl. Acad. Sci. U.S.A.955906–5912. 10.1073/pnas.95.11.5906
11
BroughtonW.HernandezG.BlairM. (2003). Beans (Phaseolus spp.)–model food legumes.Plant Soil25255–128. 10.1023/A:1024146710611
12
BrownC. A.MurrayA. W.VerstrepenK. J. (2010). Rapid expansion and functional divergence of subtelomeric gene families in yeasts.Curr. Biol.20895–903. 10.1016/j.cub.2010.04.027
13
BursetM.GuigóR. (1996). Evaluation of gene structure prediction programs.Genomics34353–367. 10.1006/geno.1996.0298
14
ChenN. W. G.SévignacM.ThareauV.MagdelenatG.DavidP.AshfieldT.et al (2010). Specific resistances against Pseudomonas syringae effectors AvrB and AvrRpm1 have evolved differently in common bean (Phaseolus vulgaris), soybean (Glycine max), and Arabidopsis thaliana.New Phytol.187941–956. 10.1111/j.1469-8137.2010.03337.x
15
ChengY.LiX.JiangH.MaW.MiaoW.YamadaT.et al (2012). Systematic analysis and comparison of nucleotide-binding site disease resistance genes in maize.FEBS J.2792431–2443. 10.1111/j.1742-4658.2012.08621.x
16
ChengZ.StuparR. M.GuM.JiangJ. (2001). A tandemly repeated DNA sequence is associated with both knob-like heterochromatin and a highly decondensed structure in the meiotic pachytene chromosomes of rice.Chromosoma11024–31. 10.1007/s004120000126
17
ChinD. B.Arroyo-GarciaR.OchoaO. E.KesseliR. V.LavelleD. O.MichelmoreR. W. (2001). Recombination and spontaneous mutation at the major cluster of resistance genes in lettuce (Lactuca sativa).Genetics157831–849.
18
ChumaI.HottaY.TosaY. (2011). Instability of subtelomeric regions during meiosis in Magnaporthe oryzae.J. Gen. Plant Pathol.77317–325. 10.1007/s10327-011-0338-6
19
CreusotF.MacadréC.Ferrier CanaE.RiouC.GeffroyV.SévignacM.et al (1999). Cloning and molecular characterization of three members of the NBS-LRR subfamily located in the vicinity of the Co-2 locus for anthracnose resistance in Phaseolus vulgaris.Genome42254–264. 10.1139/gen-42-2-254
20
CruteI. R.PinkD. (1996). Genetics and utilization of pathogen resistance in plants.Plant Cell81747–1755. 10.1105/tpc.8.10.1747
21
DavidP.ChenN. W. G.Pedrosa-HarandA.ThareauV.SevignacM.CannonS. B.et al (2009). A nomadic subtelomeric disease resistance gene cluster in common bean.Plant Physiol.1511048–1065. 10.1104/pp.109.142109
22
DavidP.SévignacM.ThareauV.CatillonY.KamiJ.GeptsP.et al (2008). BAC end sequences corresponding to the B4 resistance gene cluster in common bean: a resource for markers and synteny analyses.Mol. Genet. Genomics280521–533. 10.1007/s00438-008-0384-8
23
DevosK. M.BrownJ. K. M.BennetzenJ. L. (2002). Genome size reduction through illegitimate recombination counteracts genome expansion in Arabidopsis.Genome Res.121075–1079. 10.1101/gr.132102
24
EichlerE. E. (2001). Segmental duplications: what’s missing, misassigned, and misassembled-and should we care?Genome Res.11653–656. 10.1101/gr.188901
25
EitasT. K.NimchukZ. L.DanglJ. L. (2008). Arabidopsis TAO1 is a TIR-NB-LRR protein that contributes to disease resistance induced by the Pseudomonas syringae effector AvrB.Proc. Natl. Acad. Sci. U.S.A.1056475–6480. 10.1073/pnas.0802157105
26
El BaidouriM.PanaudO. (2013). Comparative genomic paleontology across plant kingdom reveals the dynamics of TE-driven genome evolution.Genome Biol. Evol.5954–965. 10.1093/gbe/evt025
27
EvtushenkoE. V.ElisafenkoE. A.VershininA. V. (2010). The relationship between two tandem repeat families in rye heterochromatin.Mol. Biol.441–7. 10.1134/s0026893310010012
28
Fiston-LavierA. S.AnxolabehereD.QuesnevilleH. (2007). A model of segmental duplication formation in Drosophila melanogaster.Genome Res.171458–1470. 10.1101/gr.6208307
29
FonsêcaA.FerreiraJ.Dos SantosT. R. B.MosiolekM.BellucciE.KamiJ.et al (2010). Cytogenetic map of common bean (Phaseolus vulgaris L.).Chromosom. Res.18487–502. 10.1007/s10577-010-9129-8
30
FranszP. F.ArmstrongS.de JongJ. H.ParnellL. D.van DrunenC.DeanC.et al (2000). Integrated cytogenetic map of chromosome arm 4S of A. thaliana: structural organization of heterochromatic knob and centromere region.Cell100367–376. 10.1016/S0092-8674(00)80672-8
31
GaoD. Y.AbernathyB.RohksarD.SchmutzJ.JacksonS. A. (2014). Annotation and sequence diversity of transposable elements in common bean (Phaseolus vulgaris).Front. Plant Sci.5:339. 10.3389/fpls.2014.00339
32
GeffroyV.CreusotF.FalquetJ.SévignacM.Adam-BlondonA. F.BannerotH.et al (1998). A family of LRR sequences in the vicinity of the Co-2 locus for anthracnose resistance in Phaseolus vulgaris and its potential use in marker-assisted selection.Theor. Appl. Genet.96494–502. 10.1007/s001220050766
33
GeffroyV.MacadréC.DavidP.Pedrosa-HarandA.SévignacM.DaugaC.et al (2009). Molecular analysis of a large subtelomeric nucleotide-binding-site-leucine- rich-repeat family in two representative genotypes of the major gene pools of Phaseolus vulgaris.Genetics181405–419. 10.1534/genetics.108.093583
34
GeffroyV.SévignacM.BillantP.DronM.LanginT. (2008). Resistance to Colletotrichum lindemuthianum in Phaseolus vulgaris: a case study for mapping two independent genes.Theor. Appl. Genet.116407–415. 10.1007/s00122-007-0678-y
35
GeffroyV.SévignacM.De OliveiraJ. C.FouillouxG.SkrochP.ThoquetP.et al (2000). Inheritance of partial resistance against Colletotrichum lindemuthianum in Phaseolus vulgaris and co-localization of quantitative trait loci with genes involved in specific resistance.Mol. Plant Microbe Interact.13287–296. 10.1094/MPMI.2000.13.3.287
36
GeffroyV.SicardD.de OliveiraJ. C.SévignacM.CohenS.GeptsP.et al (1999). Identification of an ancestral resistance gene cluster involved in the coevolution process between Phaseolus vulgaris and its fungal pathogen Colletotrichum lindemuthianum.Mol. Plant Microbe Interact.12774–784. 10.1094/MPMI.1999.12.9.774
37
González-GarcíaM.González-SánchezM.PuertasM. J. (2006). The high variability of subtelomeric heterochromatin and connections between nonhomologous chromosomes, suggest frequent ectopic recombination in rye meiocytes.Cytogenet. Genome Res.115179–185. 10.1159/000095240
38
GorbunovaV.LevyA. A. (1999). How plants make ends meet: DNA double-strand break repair.Trends Plant Sci.4263–269. 10.1016/S1360-1385(99)01430-2
39
GuindonS.GascuelO. (2003). A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood.Syst. Biol.52696–704. 10.1080/10635150390235520
40
HeacockM.SpanglerE.RihaK.PuizinaJ.ShippenD. E. (2004). Molecular analysis of telomere fusions in Arabidopsis: multiple pathways for chromosome end-joining.EMBO J.232304–2313. 10.1038/sj.emboj.7600236
41
HulbertS. H.WebbC. A.SmithS. M.SunQ. (2001). Resistance gene complexes: evolution and utilization.Annu. Rev. Phytopathol.39285–312. 10.1146/annurev.phyto.39.1.285
42
InnesR. W.Ameline-TorregrosaC.AshfieldT.CannonE.CannonS. B.ChackoB.et al (2008). Differential accumulation of retroelements and diversification of NB-LRR disease resistance genes in duplicated regions following polyploidy in the ancestor of soybean.Plant Physiol.1481740–1759. 10.1104/pp.108.127902
43
JonesJ. D. G.DanglL. (2006). The plant immune system.Nature444323–329. 10.1038/nature05286
44
JupeF.PritchardL.EtheringtonG. J.MacKenzieK.CockP. J.WrightF.et al (2012). Identification and localisation of the NB-LRR gene family within the potato genome.BMC Genomics13:75. 10.1186/1471-2164-13-75
45
KamiJ.PoncetV.GeffroyV.GeptsP. (2006). Development of four phylogenetically-arrayed BAC libraries and sequence of the APA locus in Phaseolus vulgaris.Theor. Appl. Genet.112987–998. 10.1007/s00122-005-0201-2
46
KamounS. (2007). Groovy times: filamentous pathogen effectors revealed.Curr. Opin. Plant Biol.10358–365. 10.1016/j.pbi.2007.04.017
47
KatohK.KumaK. I.TohH.MiyataT. (2005). MAFFT version 5: improvement in accuracy of multiple sequence alignment.Nucleic Acids Res.33511–518. 10.1093/nar/gki198
48
KellisM.PattersonN.EndrizziM.BirrenB.LanderE. S. (2003). Sequencing and comparison of yeast species to identify genes and regulatory elements.Nature423241–254. 10.1038/nature01644
49
KimS. B.KangW. H.HuyH. N.YeomS. I.AnJ. T.KimS.et al (2017). Divergent evolution of multiple virus-resistance genes from a progenitor in Capsicum spp.New Phytol.213886–899. 10.1111/nph.14177
50
KirschS.MünchC.JiangZ.ChengZ.ChenL.BatzC.et al (2008). Evolutionary dynamics of segmental duplications from human Y-chromosomal euchromatin/heterochromatin transition regions.Genome Res.181030–1042. 10.1101/gr.076711.108
51
KohlerA.RinaldiC.DuplessisS.BaucherM.GeelenD.DuchaussoyF.et al (2008). Genome-wide identification of NBS resistance genes in Populus trichocarpa.Plant Mol. Biol.66619–636. 10.1007/s11103-008-9293-9
52
KoukalovaB.MoraesA. P.Renny-ByfieldS.MatyasekR.LeitchA. R.KovarikA. (2010). Fall and rise of satellite repeats in allopolyploids of Nicotiana over c. 5 million years.New Phytol.186148–160. 10.1111/j.1469-8137.2009.03101.x
53
KourelisJ.van der HoornR. A. L. (2018). Defended to the nines: 25 years of resistance gene cloning identifies nine mechanisms for R protein function.Plant Cell30285–299. 10.1105/tpc.17.00579
54
KuangH. (2004). Multiple genetic processes result in heterogeneous rates of evolution within the major cluster disease resistance genes in lettuce.Plant Cell162870–2894. 10.1105/tpc.104.025502
55
KuoH. F.OlsenK. M.RichardsE. J. (2006). Natural variation in a subtelomeric region of arabidopsis: implications for the genomic dynamics of a chromosome end.Genetics173401–417. 10.1534/genetics.105.055202
56
LanderE. S.GreenP.AbrahamsonJ.BarlowA.DalyM. J.LincolnS. E.et al (1987). MAPMAKER: an interactive computer package for constructing primary genetic linkage maps of experimental and natural populations.Genomics1174–181. 10.1016/0888-7543(87)90010-3
57
LavinM.HerendeenP. S.WojciechowskiM. F. (2005). Evolutionary rates analysis of leguminosae implicates a rapid diversification of lineages during the tertiary.Syst. Biol.54575–594. 10.1080/10635150590947131
58
LibradoP.RozasJ. (2009). DnaSP v5: a software for comprehensive analysis of dna polymorphism data.Bioinformatics251451–1452. 10.1093/bioinformatics/btp187
59
LieberM. R.MaY. M.PannickeU.SchwarzK. (2003). Mechanism and regulation of human non-homologous DNA end-joining.Nat. Rev. Mol. Cell Biol.4712–720. 10.1038/nrm1202
60
LinJ. Y.StuparR. M.HansC.HytenD. L.JacksonS. A. (2010). Structural and functional divergence of a 1-mb duplicated region in the soybean (Glycine max) genome and comparison to an orthologous region from Phaseolus vulgaris.Plant Cell222545–2561. 10.1105/tpc.110.074229
61
LinardopoulouE. V.WilliamsE. M.FanY.FriedmanC.YoungJ. M.TraskB. J. (2005). Human subtelomeres are hot spots of interchromosomal recombination and segmental duplication.Nature43794–100. 10.1038/nature04029
62
LopezC. E.AcostaI. F.JaraC.PedrazaF.Gaitan-SolisE.GallegoG.et al (2003). Identifying resistance gene analogs associated with resistances to different pathogens in common bean.Phytopathology9388–95. 10.1094/PHYTO.2003.93.1.88
63
LukashinA. V.BorodovskyM. (1998). GeneMark.hmm: new solutions for gene finding.Nucleic Acids Res.261107–1115. 10.1093/nar/26.4.1107
64
LuoS.ZhangY.HuQ.ChenJ.LiK.LuC.et al (2012). Dynamic nucleotide-binding site and leucine-rich repeat-encoding genes in the grass family.Plant Physiol.159197–210. 10.1104/pp.111.192062
65
MaJ.BennetzenJ. L. (2004). Rapid recent growth and divergence of rice nuclear genomes.Proc. Natl. Acad. Sci. U.S.A.10112404–12410. 10.1073/pnas.0403715101
66
MamidiS.RossiM.MoghaddamS. M.AnnamD.LeeR.PapaR.et al (2013). Demographic factors shaped diversity in the two gene pools of wild common bean Phaseolus vulgaris L.Heredity110267–276. 10.1038/hdy.2012.82
67
MartinD.RybickiE. (2000). RDP: detection of recombination amongst aligned sequences.Bioinformatics16562–563. 10.1093/bioinformatics/16.6.562
68
MartinD. P.LemeyP.LottM.MoultonV.PosadaD.LefeuvreP. (2010). RDP3: a flexible and fast computer program for analyzing recombination.Bioinformatics262462–2463. 10.1093/bioinformatics/btq467
69
MartinD. P.PosadaD.CrandallK. A.WilliamsonC. (2005). A modified bootscan algorithm for automated identification of recombinant sequences and recombination breakpoints.AIDS Res. Hum. Retroviruses2198–102. 10.1089/aid.2005.21.98
70
MayorC.BrudnoM.SchwartzJ. R.RubinE. M.FrazerK. A.PachterL. S.et al (2000). VISTA : visualizing global DNA sequence alignments of arbitrary length.Bioinformatics161046–1047. 10.1093/bioinformatics/16.11.1046
71
McCarthyE. M.McDonaldJ. F. (2003). LTR STRUC: a novel search and identification program for LTR retrotransposons.Bioinformatics19362–367. 10.1093/bioinformatics/btf878
72
McclintockB. (1929). Chromosome morphology in zea mays.Science69:629. 10.1126/science.69.1798.629
73
McDowellJ. M. (1998). Intragenic recombination and diversifying selection contribute to the evolution of downy mildew resistance at the RPP8 locus of Arabidopsis.Plant Cell101861–1874. 10.1105/tpc.10.11.1861
74
McDowellJ. M.SimonS. A. (2006). Recent insights into R gene evolution.Mol. Plant Pathol.7437–448. 10.1111/J.1364-3703.2006.00342.X
75
McDowellJ. M.SimonS. A. (2008). Molecular diversity at the plant-pathogen interface.Dev. Comp. Immunol.32736–744. 10.1016/j.dci.2007.11.005
76
McHaleL. K.TrucoM. J.KozikA.WroblewskiT.OchoaO. E.LahreK. A.et al (2009). The genomic architecture of disease resistance in lettuce.Theor. Appl. Genet.118565–580. 10.1007/s00122-008-0921-921
77
MeffordH. C.TraskB. J. (2002). the complex structure and dynamic evolution of human subtelomeres.Nat. Rev. Genet.391–102. 10.1038/nrg727
78
MeyersB. C.KozikA.GriegoA.KuangH.MichelmoreR. W. (2003). Genome-wide analysis of NBS-LRR – encoding genes in Arabidopsis.Plant Cell15809–834.
79
MeyersB. C.ShenK. A.RohaniP.GautB. S.MichelmoreR. W. (1998). Receptor-like genes in the major resistance locus of lettuce are subject to divergent selection.Plant Cell101833–1846. 10.1105/tpc.10.11.1833
80
MeziadiC.RichardM. M. S.DerquennesA.ThareauV.BlanchetS.GratiasA.et al (2016). Development of molecular markers linked to disease resistance genes in common bean based on whole genome sequence.Plant Sci.242351–357. 10.1016/j.plantsci.2015.09.006
81
MiklasP. N.KellyJ. D.BeebeS. E.BlairM. W. (2006). Common bean breeding for resistance against biotic and abiotic stresses: from classical to MAS breeding.Euphytica147105–131. 10.1007/s10681-006-4600-5
82
NeedlemanS. B.WunschC. D. (1970). A general method applicable to the search for similarities in the amino acid sequence of two proteins.J. Mol. Biol.48443–453. 10.1016/0022-2836(70)90057-4
83
PacherM.Schmidt-PuchtaW.PuchtaH. (2007). Two unlinked double-strand breaks can induce reciprocal exchanges in plant genomes via homologous recombination and nonhomologous end joining.Genetics17521–29. 10.1534/genetics.106.065185
84
PadidamM.SawyerS.FauquetC. M. (1999). Possible emergence of new geminiviruses by frequent recombination.Virology265218–225. 10.1006/viro.1999.0056
85
Pedrosa-HarandA.De AlmeidaC. C. S.MosiolekM.BlairM. W.SchweizerD.GuerraM. (2006). Extensive ribosomal DNA amplification during Andean common bean (Phaseolus vulgaris L.) evolution.Theor. Appl. Genet.112924–933. 10.1007/s00122-005-0196-8
86
PosadaD.CrandallK. A. (2001). Evaluation of methods for detecting recombination from DNA sequences: computer simulations.Proc. Natl. Acad. Sci. U.S.A.9813757–13762. 10.1073/pnas.241370698
87
PryorT. P.EllisJ. (1993). The genetic complexity of fungal resistance genes in plants.Adv. Plant Pathol.10281–307.
88
PuchtaH. (2005). The repair of double-strand breaks in plants: mechanisms and consequences for genome evolution.J. Exp. Bot.561–14. 10.1093/jxb/eri025
89
RamirezM.GrahamM. A.Blanco-LopezL.SilventeS.Medrano-SotoA.BlairM. W.et al (2005). Sequencing and analysis of common bean ESTs. Building a foundation for functional genomics.Plant Physiol.1371211–1227. 10.1104/pp.104.054999.gumes
90
RatnaparkheM. B.WangX.LiJ.ComptonR. O.RainvilleL. K.LemkeC.et al (2011). Comparative analysis of peanut NBS-LRR gene clusters suggests evolutionary innovation among duplicated domains and erosion of gene microsynteny.New Phytol.192164–178. 10.1111/j.1469-8137.2011.03800.x
91
RichardM. M.ChenN. W.ThareauV.PfliegerS.BlanchetS.Pedrosa-HarandA.et al (2013). The subtelomeric khipu satellite repeat from Phaseolus vulgaris: lessons learned from the genome analysis of the andean genotype G19833.Front. Plant Sci.4:109. 10.3389/fpls.2013.00109
92
RichardM. M. S.GratiasA.ThareauV.KimK. D.BalzergueS.JoetsJ.et al (2017). Genomic and epigenomic immunity in common bean: the unusual features of NB-LRR gene family.DNA Res.25161–172. 10.1093/dnares/dsx046
93
RichterT. E.PryorT. J.BennetzenJ. L.HulbertS. H. (1995). New rust resistance specificities associated with recombination in the Rp1 complex in maize.Genetics141373–381.
94
RutherfordK.ParkhillJ.CrookJ.HorsnellT.RiceP.RajandreamM. A.et al (2000). Artemis: sequence visualization and annotation.Bioinformatics16944–945. 10.1093/bioinformatics/16.10.944
95
SamsonF.BrunaudV.DuchêneS.De OliveiraY.CabocheM.LecharnyA.et al (2004). FLAGdb++: a database for the functional analysis of the Arabidopsis genome.Nucleic Acids Res.32D347–D350. 10.1093/nar/gkh134
96
SasakiM.LangeJ.KeeneyS. (2010). Genome destabilization by homologous recombination in the germ line.Nat. Rev. Mol. Cell Biol.11182–195. 10.1038/nrm2849
97
SchlueterJ.GoicoecheaJ. L.ColluraK.GillN.LinJ.-Y.YuY.et al (2008). BAC-end sequence analysis and a draft physical map of the common bean (Phaseolus vulgaris L.) Genome.Trop. Plant Biol.140–48. 10.1007/s12042-007-9003-9
98
SchmutzJ.McCleanP. E.MamidiS.WuG. A.CannonS. B.GrimwoodJ.et al (2014). A reference genome for common bean and genome-wide analysis of dual domestications.Nat. Genet.46707–713. 10.1038/ng.3008
99
ShenK. A.MeyersB. C.Islam-FaridiM. N.ChinD. B.StellyD. M.MichelmoreR. W. (1998). Resistance gene candidates identified by PCR with degenerate oligonucleotide primers map to clusters of resistance genes in lettuce.Mol. Plant Microbe Interact.11815–823. 10.1094/MPMI.1998.11.8.815
100
SmithS. M.PryorA. J.HulbertS. H. (2004). Allelic and haplotypic diversity at the Rp1 rust resistance locus of maize.Genetics1671939–1947. 10.1534/genetics.104.029371
101
SudupakM. A.BennetzenJ. L.HulbertS. H. (1993). Unequal exchange and meiotic instability of disease-resistance genes in the Rp1 region of maize.Genetics133119–125.
102
TamuraK.DudleyJ.NeiM.KumarS. (2007). MEGA4: molecular evolutionary genetics analysis (MEGA) software version 4.0.Mol. Biol. Evol.241596–1599. 10.1093/molbev/msm092
103
UpsonJ. L.ZessE. K.BiałasA.WuC.KamounS. (2018). The coming of age of EvoMPMI: evolutionary molecular plant–microbe interactions across multiple timescales.Curr. Opin. Plant Biol.44108–116. 10.1016/j.pbi.2018.03.003
104
VelascoR.ZharkikhA.TroggioM.CartwrightD. A.CestaroA.PrussD.et al (2007). A high quality draft consensus sequence of the genome of a heterozygous grapevine variety.PLoS One2:e1326. 10.1371/journal.pone.0001326
105
VitteC.BennetzenJ. L. (2006). Analysis of retrotransposon structural diversity uncovers properties and propensities in angiosperm genome evolution.Proc. Natl. Acad. Sci. U.S.A.10317638–17643. 10.1073/pnas.0605618103
106
WanzenböckE. M.SehöferC.SchweizerD.BachmairA. (1997). Ribosomal transcription units integrated via T-DNA transformation associate with the nucleolus and do not require upstream repeat sequences for activity in Arabidopsis thaliana.Plant J.111007–1016. 10.1046/j.1365-313X.1997.11051007.x
107
WawrzynskiA.AshfieldT.ChenN. W. G.MammadovJ.NguyenA.PodichetiR.et al (2008). Replication of nonautonomous retroelements in soybean appears to be both recent and common.Plant Physiol.1481760–1771. 10.1104/pp.108.127910
108
WeiC.ChenJ.KuangH. (2016). Dramatic number variation of r genes in solanaceae species accounted for by a few r gene subfamilies.PLoS One11:e0148708. 10.1371/journal.pone.0148708
109
YoungN. D.DebelléF.OldroydG. E. D.GeurtsR.CannonS. B.UdvardiM. K.et al (2011). The Medicago genome provides insight into the evolution of rhizobial symbioses.Nature480520–524. 10.1038/nature10625
110
ZhouB.DolanM.SakaiH.WangG.-L. (2007). The genomic dynamics and evolutionary mechanism of the Pi2/9 locus in rice.Mol. Plant Microbe Interact.2063–71. 10.1094/MPMI-20-0063
111
ZhouT.WangY.ChenJ. Q.ArakiH.JingZ.JiangK.et al (2004). Genome-wide identification of NBS genes in japonica rice reveals significant expansion of divergent non-TIR NBS-LRR genes.Mol. Genet. Genomics271402–415. 10.1007/s00438-004-0990-z
Summary
Keywords
Phaseolus vulgaris (common bean), subtelomere, resistance gene, segmental duplication, evolution, satellite, homologous recombination (HR), non-homologous end-joining (NHEJ)
Citation
Chen NWG, Thareau V, Ribeiro T, Magdelenat G, Ashfield T, Innes RW, Pedrosa-Harand A and Geffroy V (2018) Common Bean Subtelomeres Are Hot Spots of Recombination and Favor Resistance Gene Evolution. Front. Plant Sci. 9:1185. doi: 10.3389/fpls.2018.01185
Received
25 April 2018
Accepted
24 July 2018
Published
14 August 2018
Volume
9 - 2018
Edited by
Anne Bagg Britt, University of California, Davis, United States
Reviewed by
Umesh K. Reddy, West Virginia State University, United States; Jer-Young Lin, University of California, Los Angeles, United States
Updates

Check for updates
Copyright
© 2018 Chen, Thareau, Ribeiro, Magdelenat, Ashfield, Innes, Pedrosa-Harand and Geffroy.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Valérie Geffroy, valerie.geffroy@ips2.universite-paris-saclay.fr
This article was submitted to Plant Genetics and Genomics, a section of the journal Frontiers in Plant Science
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.