Abstract
Supercentenarians (age 110+ years old) generally delay or escape age-related diseases and disability well beyond the age of 100 and this exceptional survival is likely to be influenced by a genetic predisposition that includes both common and rare genetic variants. In this report, we describe the complete genomic sequences of male and female supercentenarians, both age >114 years old. We show that: (1) the sequence variant spectrum of these two individuals’ DNA sequences is largely comparable to existing non-supercentenarian genomes; (2) the two individuals do not appear to carry most of the well-established human longevity enabling variants already reported in the literature; (3) they have a comparable number of known disease-associated variants relative to most human genomes sequenced to-date; (4) approximately 1% of the variants these individuals possess are novel and may point to new genes involved in exceptional longevity; and (5) both individuals are enriched for coding variants near longevity-associated variants that we discovered through a large genome-wide association study. These analyses suggest that there are both common and rare longevity-associated variants that may counter the effects of disease-predisposing variants and extend lifespan. The continued analysis of the genomes of these and other rare individuals who have survived to extremely old ages should provide insight into the processes that contribute to the maintenance of health during extreme aging.
Introduction
Human aging is affected by genes, life style, and environmental factors. The genetic contribution to average human aging can be modest with genes explaining ∼20–25% of the variability of human survival to the mid-eighties (Herskind et al., ; Fraser and Shavlik, ). By contrast, genetic factors may have greater impact on survival to the ninth through eleventh decades (Tan et al., ). Notably, exceptional longevity is rare and may involve biological mechanisms that differ from those implicated in usual human aging.
The nature and contribution of genetic variation to exceptional longevity remains unclear, particularly the role for undiscovered rare genetic variants with large effects and/or the presence of many common genetic variants with small effects (Bloss et al., ). Exceptional longevity is typically characterized by strong familiality (Perls et al., , ; Atzmon et al., ; Schoenmaker et al., ) as well as a marked delay in disability (Terry et al., ) and, as human lifespan is approached at about age 110 years, many such individuals compress not only disability but also age-related diseases (Andersen et al., ). Studies of centenarians have provided strong evidence to support the hypothesis that a genetic contribution to human exceptional longevity is decisive, although only a small number of genetic variants with modest effects have been irrefutably linked to this phenotype (Schachter et al., ; Barzilai et al., ; Christensen et al., ; Wheeler and Kim, ). The technology of next generation sequencing provides a tool to generate data that may eventually provide an answer (Metzker, ).
In this report, we describe the complete DNA sequences of two supercentenarians, a male and a female, both ages >114 years old. Although these data cannot provide conclusive evidence about the genetic determination of human exceptional longevity, they are the first step toward the generation of a comprehensive reference panel of exceptionally long-lived individuals. The data also provide interesting insights into genetic backgrounds that are conducive to exceptional longevity and allow us to test different models of exceptional longevity.
Results
Subjects’ characteristics
Figure 1 shows some of the characteristics of the two subjects, PG17 (female) and PG26 (male). We are purposefully vague about the exact ages of these individuals to maintain their confidentiality and decrease the risk of their being identified. Both individuals were Caucasians enrolled in the New England Centenarian Study (NECS), and were reassessed annually until their deaths (Andersen et al., ). Their European ancestry was verified by genome-wide principal component analysis. They were selected for our sequencing study because of their exceptional lifespans and health histories, and the extreme delays in the ages of onset of disability that they exhibited. The bottom panel of Figure 1 shows the average scores of cognitive and physical functions of the female (PG17) and the male (PG26) relative to trajectories of cognitive and functional declines that we generated using longitudinal data of more than 1,300 centenarians and their nonagenarian siblings (Andersen et al., ). While the decline of physical function of the female matches the average trend of supercentenarians enrolled in the NECS (n = 104), the male maintained functional independence until at least the last 6 months of his life. The average Blessed Information–Memory–Concentration Test scores (measures of cognitive function) of both subjects were higher than the average of the NECS supercentenarian sample, until their deaths. The man had substantial longevity in his family; 25% of his siblings lived past the age of 100 and 50% lived past the age of 90 years. Unfortunately we did not have complete information about the family history of the woman.
Figure 1
Description of the DNA sequencing
DNA from both individuals was sequenced using the Illumina Genome Analyzer II by Illumina’s Clinical Laboratory Service, using paired-end reads of 100 bp producing 1,650,463,996 reads for the woman and 1,804,595,182 reads for the man. Reads were mapped to the genome reference NCBI36 and NCBI37 using the procedures described in Figure A1 in Appendix. Both reference genomes were used because some databases have not yet completed the transition to NCBI37. The Eland aligner (Bentley et al.,
Variant calling (SNPs and Indels)
Single nucleotide polymorphisms (SNPs) were called using the Illumina CASAVA (Bentley et al.,
Figure 2

(A) Summary of SNPs and small insertions and deletions (Indels) detected in the two sequences, their union and intersection. The January 2011 release of the 1000 Genomes was used to compare the SNPs detected in the two sequences that were not found in dbSNP. (B) Venn diagram that illustrates the proportion of known and novel SNPs in the woman (PG17) and man (PG26) combined. (C) Distribution of effective coverage (reads used to call genotypes in CASAVA) for novel SNPs. (D) Distribution of Phred scores for the novel SNPs. The Phred scores were computed with BWA/SAMTools.
The two subjects shared 1,997,897 SNPs, and 66% of these SNPs had the same genotypes (Figure 2A). Comparison with dbSNP 132 and the 1000 Genomes database2 showed that 97.6 and 97.5% of SNPs in the female and the male subjects were in dbSNP, and up to 99.7% of the SNPs in common between the two subjects were in dbSNP. Only 1% of the SNPs in both individuals’ sequences were reported in the 1000 Genomes database but not in dbSNP, and 2,590 of these SNPs were shared between the two subjects, 81% of which with the same genotypes. 44,713 SNPs in the woman and 48,725 SNPs in the man were neither in dbSNP nor in the 1000 Genomes database, and were therefore labeled “novel.” Only 2,254 of these novel SNPs were shared between the two subjects, and 73% of these were called with the same genotypes. Figure 2 summarizes these findings. The decreasing trend of shared SNPs that were either in the 1000 Genomes database or novel is consistent with these SNPs being rare in the population.
The number of called SNPs in the man and the fraction of novel SNPs in both subjects were consistent with the projections made by The 1000 Genomes Project Consortium (
For additional assessment of the quality of the data, we computed the concordance between genotype calls in the sequences of the two subjects and their SNP array data. For both subjects, the concordance between genotype calls of SNPs in the array and SNPs in the sequences was >99.7%. In addition, more than 98% of the SNPs in the array not included in the list of called SNPs from the sequencing were homozygous for the referent allele. Transition to transversion ratios (2.11 in the man and 2.07 in the woman) were consistent with the expected number in Caucasians (Ebersberger et al.,
We used both SAMTools and Dindel (Albers et al.,
Functional annotation
We used a suite of bioinformatics tools assembled and built by researchers at The Scripps Research Institute (Torkamani et al.,
Figure 3

(A) functional annotation of SNPs in PG17 (female), PG26 (male), their union and intersection. Multiple transcripts of the same genes were counted individually. (B) Rate of non-synonymous SNPs in coding SNPs in PG17 and PG26 (brown) and other whole genome sequences generated at Complete Genomics (orange), Sanger (green), and Illumina (blue). CG (9 CEPH) is the average rate of non-synonymous SNPs in nine whole genome sequences generated at Complete Genomics. The rates of non-synonymous SNPs in Venter (green), AK, CH, and YRI (blue) are from publications (see Table A2 in Appendix for references). (C,D) Compare the distribution of SNPs locations and role in PG17 and PG26 (brown) and nine whole genome sequences generated at Complete Genomics (orange). (E) Provides a summary of the predicted impact of coding SNPs in PG17, PG26, their union and intersection, and the summary in nine Caucasians genotyped at complete genomics and annotated in the same way as PG17 and PG26.
Different genetic models of exceptional longevity
We used the whole genome sequences of these two subjects to test different hypotheses about the genetics of exceptional longevity. These non-exclusive hypotheses and the results of the analyses are described in the sections that follow.
Hypothesis 1: The “metabolic” hypothesis
Genes involved in the insulin pathway (Guarente and Kenyon,
Table 1
| A. Coding SNPs linked to exceptional longevity in published literature | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| SNP | HG18 | Alleles | Ref Allele | EL Allele | p (EL allele) | Pubmed | PG17 | PG26 | |
| FOXO3A | rs12206094 | chr6:109012893 | C/T | C | T | 0.27 | CC | TT | |
| rs2764264 | chr6:109041154 | C/T | C | C | 0.29 | TT | CC | ||
| rs7762395 | chr6:109051800 | A/G | G | A | 0.17 | 20849522 | GG | AA | |
| rs9400239 | chr6:109084356 | C/T | T | T | 0.24 | CC | TT | ||
| rs479744 | chr6:109126725 | G/T | G | T | 0.20 | GG | TT | ||
| IGF1R | rs2229765 | chr15:97295748 | G/A | G | A | 0.42 | 12843179 | AG | AA |
| rs34516635 | chr15:97269499 | G/A | G | A | 0.02 | GG | GG | ||
| chr15:97068418 | G/A | G | A | 18316725 | GG | GG | |||
| chr15:97272104 | G/A | G | A | GG | GG | ||||
| HSP70 | rs2227956 | chr6:31886251 | G/A | G | A | 0.76 | PMC1576475 | AA | AA |
| CETP | rs5882 | chr16:55573593 | A/G | G | GG | 0.36 | 20068209 | AA | AA |
| PON1 | rs662 | chr7:94775382 | T/C | T | C | 0.33 | 15050299 | CT | TT |
| MINPP1 | rs9664222 | chr10:89338633 | A/C | A | C | 0.76 | 20304771 | CC | AC |
| SIRT1 | Rs3758391 | chr10:69313348 | C/T | T | T | 0.27 | 17895433 | CC | CT |
| Klotho | rs9536314 | chr13:32526138 | T/G | T | G | 0.14 | 11792841 | GT | GT |
| rs9527025 | chr13:32526193 | G/C | G | C | 0.19 | CG | CG | ||
| B. Additional coding SNPs in genes linked to exceptional longevity | |||||||||
| Gene | SNP | HG18 | Alleles | Ref | EL Allele | CEU | Impact | PG17 | PG26 |
| SIRT1 | rs2273773 | chr10:69336604 | C/T | T | T | 0.96 | Syn | TT | CT |
| SIRT3 | rs28365927 | chr11:226091 | A/G | G | G | 0.85 | Non-syn | AG | GG |
| klotho | rs2772364 | chr13:32488851 | C/T | T | C | 1.00 | Syn | CC | TT |
| rs9527026 | chr13:32526239 | A/G | G | G | 0.83 | Syn | AG | AG | |
| rs564481 | chr13:32532983 | C/T | C | C | 0.63 | Syn | CT | CT | |
| rs648202 | chr13:32533463 | C/T | T | C | 0.85 | Syn | CC | CC | |
| rs649964 | chr13:32533835 | C/T | T | C | 1.00 | Syn | CC | CC | |
| IGF1R | rs35812156 | chr15:97252339 | A/C | C | C | 0.98 | Syn | AC | CC |
| SIRT6 | rs352493 | chr19:4131836 | C/T | C | T | 0.96 | Non-syn | CT | TT |
| SIRT5 | rs3757261 | chr6:13707282 | C/T | C | C | 0.73 | Syn | CT | CT |
| P ON1 | rs854560 | chr7:94784020 | A/T | A | A | 0.59 | Non-syn | AA | TT |
(A) List of SNPs that were linked to aging and exceptional longevity in the literature based on candidate gene studies and/or animal models. Columns 5–7 define the referent allele (Ref Allele), the allele that was associated with longevity (EL Allele), and the frequency of the longevity allele in Caucasians from dbSNP. Boldface highlights genotypes with one or more longevity variants. (B) All additional coding SNPs in the list of candidate genes in which PG17 or PG26 carry non-referent alleles. Column 7 (CEU) reports the frequency of the referent alleles in Caucasians from dbSNP.
The woman was homozygous for only one of the 16 longevity variants (allele A of rs2227956 in HSP70), and heterozygous for 4 others. The man was homozygous for seven longevity variants and heterozygous for four. Neither of the two sequences carried the longevity-associated variant of CETP that was found in Ashkenazi Jewish centenarians (Barzilai et al.,
Hypothesis 2: The lack of disease-associated variants hypothesis
As noted earlier, both subjects markedly delayed both disability and age-related diseases until very late in their lives. We tested the hypothesis that these two whole genome sequences did not include disease-predisposing variants or, if they did, the number was significantly lower compared to currently available genomes. We compiled a list of 62,339 disease-annotated variants from the Human Genome Mutation Database (HGMD@; Stenson et al.,
Figure 4A shows that while the two sequences include only 1% of mutations from the HGMD, they include approximately 50% of the mutations that were linked to common diseases in genome-wide association studies. More than 50% of all the noted mutations were heterozygous, but this number was smaller when we only considered coding mutations from the HGMD. Figure 4B shows the breakdown of these variants by disease group and by role: (1) either damaging, if they are associated with increased risk for disease, or (2) protective if the mutations are known to decrease disease risk relative to the general population. Only 1% of known disease-annotated mutations in the woman were protective, while 2% of known disease-annotated mutations in the man were protective. The woman carried at least 30 mutations that were linked to Alzheimer’s disease and amyotrophic lateral sclerosis and one mutation linked to decreased risk for Alzheimer’s disease (CT genotype for rs2736911 in ARMS2, Gatta et al.,
Figure 4

(A) Summary of SNPs associated with disease in the HGMD and the GWAS catalog, and number of these SNPs and rates found in PG17 and PG26. (B) Number of SNPs in PG17 and PG26 that have either a known protective or deleterious role in major age-related diseases. (C) The bar plot shows the rate of disease-annotated variants in PG17 and PG26 and 11 other whole genome sequences. We did not include in this analysis the genomes of NA12878 and NA18507 that were sequenced with SOLID. Blue: all SNPs from HGMD and GWAS; red: only coding SNPs from HGMD; green: only SNPs from GWAS. Rates are per 100,000 SNPs. (D). Rate of protective variants in PG17 and PG26 and 11 other whole genome sequences. The major difference seems to be race related rather than subject related.
The bar plot in Figure 4C shows the rate of disease-associated variants in both subjects and 11 additional genomes described in Moore et al. (
Overall, the two subjects carried 403 disease-associated variants with the same genotype (Table S1 in Supplementary Material), but only 209 of these were coding variants, and in only 76 of these positions the two sequences were homozygous for the risk allele. For example, both subjects were CC homozygotes for rs222859 (YBX2, chromosome 17). This gene is associated with male infertility, and the male subject did not have children (but we have no additional information). Both subjects were homozygous for the A allele of rs4880 in SOD2. The alternative allele G is a missense mutation that changes the amino acid valine to alanine. The common allele A has been associated with increased oxidative stress that should act negatively on lifespan, and increased risk for cardiovascular disease and susceptibility to cancer (Bastaki et al.,
This analysis shows that the two supercentenarians did not carry a smaller rate of disease-associated variants compared to other genomes.
Hypothesis 3: The rare variants hypothesis
One of the explanations for the small number of genetic variants irrefutably linked to exceptional longevity (Schachter et al.,
Eleven of the genes with novel mutations in the woman were linked to the class of diseases “infection” in the genetic association database (LTBP2, IL16, SELL, CABIN1, IRF4, KIR2DL3, DDX5, KIR2DL2, IFNGR2, TLR9, GJB2, p-value 0.066 that however does not remain significant after Bonferroni correction), but no additional disease categories were found to be significantly enriched. Thirty six genes with novel coding mutations in the man were linked to the disease classes “infection” and “immune.” These genes collectively point to the innate immune response and extracellular matrix remodeling genes collectively arguing for immune and tissue homeostasis as a nodal point for these subjects.
Figure 5A shows the distribution of genes with novel coding mutations in the woman (PG17) and the man (PG26) by specific diseases. The woman had a higher rate of novel mutations in genes linked to immune diseases and infections, as well as hematological diseases and development compared to the man. The man had a higher rate of novel mutations in genes linked to cancer, cardiovascular disease, neurological disease, and metabolism. To test whether the distributions of genes with novel mutations were significantly enriched for these disease categories, we generated reference distributions for the disease-annotated genes by randomly selecting genes from the whole list of genes with coding SNPs in the two subjects. The genes were annotated by disease groups and average rates and SD were computed in 1,000 resampled sets. The analysis showed that the rate of genes with novel coding SNPs annotated as “development,” “immune,” “infection,” and “renal” differed from the distributions expected if the genes were selected at random in the woman (Figure 5B). Interestingly, while the rate of genes with novel coding mutations annotated as “immune” and “infection” in the man was higher than the mean rate, it was not significantly different from what is expected by chance (Figure 5C). The significant enrichment of these disease categories in the woman but not the man suggests that the woman may be enriched for private mutations that promote exceptional longevity while the man may be enriched for more common longevity-associated variants. Figure 5D shows some examples of novel variants.
Figure 5

(A) Rates of genes with novel coding mutations annotated by disease. Disease groups are as in the labels of (B,C). Note that the woman did not carry any novel mutations in known aging genes, and neither of them carried novel coding SNPs in mitochondrial genes. (B) Resampling-based distributions of the rates of disease-annotated genes in PG17 when ∼400 genes were selected at random from the list of genes with coding SNPs. Red asterisks display the rate of disease-annotated genes with novel coding mutations in PG17. Distributions are based on 1,000 simulations. Error bars represent mean rate ± 2 SD. (C) Resampling-based distributions of the rates of disease-annotated genes in PG26 when ∼500 genes were selected at random in the list of genes with coding SNPs. Distributions are based on 1,000 simulations. (D) A sample of novel mutations in the female (PG17) and male (PG26). The complete list is in Table S2 in Supplementary Material.
Hypothesis 4: The enrichment of longevity variants hypothesis
In a genome-wide association study of exceptional longevity with 801 centenarians (median age at death 104 years) and 914 genetically matched controls, we identified 281 SNPs that were significantly associated with exceptional longevity and could be used to predict the phenotype with 78% sensitivity for a replication set with a mean age of 108 years (Sebastiani et al.,
Figure 6

(A) Distance in log 10 (bp) of 51 SNPs that are predictive of exceptional longevity and nearest non-referent coding SNPs in PG17 and PG26. (B) Details of the 17 SNPs that are within 10 kb from coding SNPs. (C) The box plot in red shows the distribution of log 10 (bp) distance between the longevity-associated variants and the closest coding SNPs in PG17 and PG26. The 100 box plots in white show the distributions of the distance between SNPs chosen at random from the SNPs included in the genome-wide association study in Sebastiani et al. (
Discussion
We generated the whole genome sequences of a male and a female supercentenarian who were selected for this study because of both their exceptional lifespan and healthspan. Although two sequences do not provide sufficient data for general inference on the genetics of exceptional longevity, they are a first step toward the generation of a reference panel of exceptionally long-lived individuals and provide some interesting insights about genetic backgrounds that might be conducive to exceptional longevity. The $10 million Archon Genomics X Prize6 will greatly expand this reference panel with extremely accurate, 100 “medical grade” and economically feasible whole genome sequences of centenarians sequenced by multiple competing teams. It was recently announced that over 1,000 individuals who have lived to at least age 80 without suffering from any common chronic diseases will have their entire genomes sequenced and made available to the scientific community for use as a reference panel7.
The analysis of next generation sequence data is still very challenging, and no single mapping and variant calling algorithm has emerged as the standard tool to use. Therefore, we used two algorithms to select a set of robust variants to examine further. The specifically selected data show that the genetic architectures of the two subjects are comparable to published sequenced genomes, particularly in terms of rates of coding variants and predicted damaging variants, and it is likely that overall differences seen between genomes are due to different platforms, algorithms, and annotation methods rather than major structural changes that can be linked to exceptional longevity, at least in the case of germ line mutations.
We also used the genome sequences of these two subjects to test different genetic models of exceptional longevity. The insulin pathway, caloric restriction, and lipid metabolism significantly influence lifespan in other organisms including the mouse, fly, and worm (Christensen et al.,
One of the hypothetical genetic models of exceptional longevity is that, in order for centenarians to achieve their exceptional survival, they must lack disease-predisposing variants. Phenotypically, there is evidence that this is not necessarily the case, since approximately 40% of centenarians have had onset of age-related diseases before the age of 80 (Evert et al.,
An alternative, complementary hypothesis is that variants, possibly rare, not hitherto associated with health maintenance, compensate for disease-causing variants among the healthy elderly. In this light, the genomes of these two supercentenarians were not particularly enriched for novel variants, although the stringency of our approach to variant calling may have increased the false negative rate to reduce the false-positive rate. Nevertheless, the 1% novel variants discovered through this analysis may lead to the discovery of novel genes involved with exceptional longevity. The observation that the novel variants were in genes implicated in “alternative splicing” highlights the importance of pursuing follow-up studies involving transcriptome sequencing from different cells and tissues to identify RNA isoforms or expression profiles that may be important for exceptional longevity.
In summary, the two supercentenarian genomes we studied have different features including, for example, a number of private mutations in specific categories of genes in the case of the woman but not the man. Other ongoing studies of centenarians are showing that different centenarian genomes have different characteristics (Cirulli et al.,
It is also likely that environmental factors and possibly the genetic ancestry may influence the likelihood of an individual to live long ages directly or by interacting with the genetic background. The NECS has shown that the chance of male and female siblings of centenarians to live past 100 can be 8 and 17 times higher than the risk in the general population (Perls et al.,
Materials and Methods
Ethics statement
Subjects provided informed consent and this project was approved by the Boston University Medical Campus Institutional Review Board.
Subjects
The man and woman were both age 114+ years at the time of blood collection for DNA extraction. Their age was confirmed by birth/baptism certificates and early U.S. Census entries. For both subjects we examined the European ancestry by principal component analysis of genome-wide genotype data, as described in Sebastiani et al. (
Sequencing
Peripheral blood was obtained from each subject and used for DNA extraction. DNA samples were sequenced using a GAII sequencer at the Illumina Clinical Service Laboratory (San Diego, CA) using 100 bp paired-end reads. Fastq files were generated from the image files using the Illumina pipeline software (Firecrest for image analysis and Bustard for base calling).
Mapping
Reads in Fastq files were mapped to the NCBI36 reference genome including all chromosomes and Mt DNA using the Eland aligner (Bentley et al.,
aln -n 4 -o 1 -e 2 -k 2 -l 35 -R 5
to limit the maximum edit distance to 4, the maximum number of gap opens to 1, the maximum number of gap extensions to 2, the maximum number of edit distance in the seed to 2, the seed to the first 35 bases, and to proceed with suboptimal alignment if there are no more than 5 best hits. The more relaxed thresholds produced a larger number of aligned reads [1,559,315,652 aligned reads for PG26 (86.41%) and 1,334,406,235 aligned reads for PG17 (80.85%)].
SNP calling
Single nucleotide polymorphisms were called using the CASAVA algorithm (see http://www.illumina.com/software/genomestudio_software.ilmn and Bentley et al.,
Identification of novel variants
We annotated the single polymorphic bases that differ from the reference genome (hg18) against dbSNP build 132 and the summary data distributed by the 1000 Genomes project9 that were published using coordinates of the UCSC release hg19 of the human genome. We used the liftover conversion tool in the UCSC Genome Browser to convert genomic coordinates from hg18 to hg19, and then calculated the number of polymorphic bases in the sequences of PG17 and PG26 that are not in dbSNP 132 but were reported in the December 2010 data release from the 1000 Genomes project. With the exception of 65 positions that map to 22 unique regions in hg18 and could not be remapped to hg19, all others were successfully translated, and more than 65% of the variants not in dbSNP were found in the 1000 Genomes data release. We also checked novel variants to detect mapping errors and tried to remove novel variants that were within few bases as those could be due to mapping errors. After this annotation, approximately 1.5% of the SNPs from the original set that were not found in dbSNP or 1000 Genomes were labeled as “novel” (Figure 2A). These steps were conducted with Genephony10, an online tool for the manipulation of large datasets of genomic information (Nuzzo and Riva,
Additional genomes
We identified 15 additional whole genome sequences for comparison. The samples and references are in Table A2 in Appendix.
SNP quality
We assessed SNP call quality using several measures.
Runs of homozygosity
We used SNP array data genotyped with the Illumina array to describe the genome-wide structure of chromosomal alterations of PG17 and PG26. The two samples were genotyped with the Illumina 610 array (PG17) and 1M array (PG26) and data were processed as described in Sebastiani et al. (
Concordance with genotype calls
We compared genotype calls detected from next generation sequence analysis against the genotype calls determined with SNP array analysis. Missing genotypes were ignored.
Transition to transversion ratios and rate of heterozygosity
We calculated the number of transitions (A ↔ G or C ↔ T) and transversions (A ↔ C,T; C ↔ G; and T ↔ G) and their rate in all called SNPs, in SNPs reported only in the 1000 Genomes project, and in novel SNPs. Rates of heterozygosity were computed by the ratio of the number of heterozygous calls versus homozygous calls only for the SNPs with alleles that differ from the reference genome.
Indel calling
For short indel finding, the reads were mapped to the human reference genome (NCBI build 37) using BWA in paired-end mode with default settings except for the following differences in BWA aln module: to control false-positive rate, the maximum edit distance in mapped end (-n) was lowered from 5 (default for 100 bp reads) to 4, and, to increase pairing accuracy, the seed size (-l) was decreased to 20 and the number of best hits to search in suboptimal alignments (-R) was increased to 100. In BWA sampe module, the number of occurrences for one end (-o) was increased to 10 million.
We found that using BWA with the last three parameters reduces the number of discordant pairs and pairs with ends mapped to different chromosomes without sacrificing overall alignment accuracy. Since Dindel uses both ends of discordant pairs for realignment and for calculation of variant qualities, we wanted to keep the number of discordant pairs low. The BAM alignment files produced by BWA were merged and split by chromosomes, library and read group information was added, and duplicates were removed on per-library basis using Picard MarkDuplicates. Short indel calls were made using Dindel and SAMTools mpileup. Since our data were relatively high coverage, before running Dindel, candidate indels were selected with – minCount parameter set to 3 using the script in the Dindel package. SAMTools mpileup was run with mapping quality coefficient – C50. For the mitochondrial chromosomes where coverage exceeded 3,000 due to high sample copy number, we increased maximum depth in mpileup to one million. Only indels that were called by both algorithms and had a Phred score >30 were selected for further annotation.
Functional annotation
For bioinformatics analysis, we used the complementary genome-wide variant annotation tools embedded in a suite of tools developed by researchers at The Scripps Research Institute (Torkamani et al.,
List of disease-annotated variants
We compiled this list by merging disease-annotated variants from the HGMD11 with the catalog of genome-wide association studies12. We completed the list of variants in the GWAS catalog by searching for the unreported SNP alleles, and then compared the full sequences of PG17 and PG26 to detect the SNPs with risk alleles. Note that for this analysis we used the assembled sequences and not only the SNPs that were called by the SNP caller algorithms, because several of the risk alleles are actually the reference allele in the hg18 version of the human genome. All SNPs in the catalog were recoded according to the forward strand to make the results comparable to the whole genome sequences. We analyzed the number of risk alleles stratified by whether they belong to a coding SNP. Protective variants were identified by searching for the key-words “reduced risk” in disease annotation.
Enrichment analysis
We used the David functional tool (Huang da et al.,
Longevity-associated variants
We detected the longevity-associated variants in PG17 and PG26 as those genotypes that increase the posterior probability of exceptional longevity in the two subjects in the list of 281 SNPs reported in Sebastiani et al. (
Supplementary Material
The Supplementary Material for this article can be found online at http://www.frontiersin.org/genetics_of_aging/10.3389/fgene.2011.00090/abstract
Statements
Ethics statement
Subjects provided informed consent and this project was approved by the Boston University Medical Campus Institutional Review Board.
Acknowledgments
This work was funded by grants from the National Institute on Aging R56-AG027216 (Clinton T. Baldwin), R01HL087681 (Martin H. Steinberg), K24AG025727 and a Glenn Foundation for Medical Research Breakthroughs in Gerontology Award (Thomas T. Perls), National Science Foundation IIS-1017621 (Gary Benson), NIH/NCRR UL1 RR025774; NIH/NHLBI R01 HL089655-03; and NIH/NIDA R01 DA030976, and the Price Foundation and Scripps Genomic Medicine (Ali Torkamani, Nicholas J. Schork, Phillip Pham). We are beholden to the very special and wonderful subjects of this study for their interest and participation.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Footnotes
1.^http://www.illumina.com/Documents/products/technotes/technote_snp_caller_sequencing.pdf
2.^http://www.1000genomes.org/
3.^http://www.ncbi.nlm.nih.gov/gap
4.^http://www.snpedia.com/index.php/Rs4880
5.^http://david.abcc.ncifcrf.gov/
6.^http://genomics.xprize.org/
7.^http://ir.completegenomics.com/releasedetail.cfm?ReleaseID=610178
8.^http://www.illumina.com/software/genomestudio_software.ilmn
9.^ftp://ftp-trace.ncbi.nih.gov/1000genomes/ftp/release/2010_11/
References
1
AlbersC. A.LunterG.MacArthurD. G.McVeanG.OuwehandW. H.DurbinR. (2011). Dindel: accurate indel calls from short-read data. Genome Res.21, 961–973.10.1101/gr.112326.110
2
Al-RegaieyK. A.MasternakM. M.BonkowskiM.SunL.BartkeA. (2005). Long-lived growth hormone receptor knockout mice: interaction of reduced insulin-like growth factor i/insulin signaling and caloric restriction. Endocrinology146, 851–860.10.1210/en.2004-1120
3
AndersenS.SebastianiP.DworkisD. A.FeldmanL.PerlsT. T. (2011). Health span approximates life span amongst many supercentenarians. J. Gerontol. A Biol. Sci. Med. Sci. [Epub ahead of print].10.1093/gerona/glr223
4
AraiY.TakayamaM.GondoY.InagakiH.YamamuraK.NakazawaS.KojimaT.EbiharaY.ShimizuK.MasuiY.KitagawaK.TakebayashiT.HiroseN. (2008). Adipose endocrine function, insulin-like growth factor-1 axis, and exceptional survival beyond 100 years of age. J. Gerontol. A Biol. Sci. Med. Sci.63, 1209–1218.10.1093/gerona/63.11.1209
5
ArkingD. E.KrebsovaA.MacekM. Sr.MacekM.Jr.ArkingA.MianI. S.FriedL.HamoshA.DeyS.McIntoshI.DietzH. C. (2002). Association of human aging with a functional variant of klotho. Proc. Natl. Acad. Sci. U.S.A.99, 856–861.10.1073/pnas.022484299
6
AtzmonG.RinconM.RabizadehP.BarzilaiN. (2005). Biological evidence for inheritance of exceptional longevity. Mech. Ageing Dev.126, 341–345.10.1016/j.mad.2004.08.026
7
BarzilaiN.AtzmonG.DerbyC. A.BaumanJ. M.LiptonR. B. (2006). A genotype of exceptional longevity is associated with preservation of cognitive function. Neurology67, 2170–2175.10.1212/01.wnl.0000249116.50854.65
8
BarzilaiN.AtzmonG.SchechterC.SchaeferE. J.CupplesA. L.LiptonR.ChengS.ShuldinerA. R. (2003). Unique lipoprotein phenotype and genotype associated with exceptional longevity. JAMA290, 2030–2040.10.1001/jama.290.15.2030
9
BastakiM.HuenK.ManzanilloP.ChandeN.ChenC.BalmesJ. R.TagerI. B.HollandN. (2006). Genotype-activity relationship for Mn-superoxide dismutase, glutathione peroxidase 1 and catalase in humans. Pharmacogenet. Genomics16, 279–286.10.1097/01.fpc.0000199498.08725.9c
10
BentleyD. R.BalasubramanianS.SwerdlowH. P.SmithG. P.MiltonJ.BrownC. G.HallK. P.EversD. J.BarnesC. L.BignellH. R.BoutellJ. M.BryantJ.CarterR. J.Keira CheethamR.CoxA. J.EllisD. J.FlatbushM. R.GormleyN. A.HumphrayS. J.IrvingL. J.KarbelashviliM. S.KirkS. M.LiH.LiuX.MaisingerK. S.MurrayL. J.ObradovicB.OstT.ParkinsonM. L.PrattM. R.RasolonjatovoI. M.ReedM. T.RigattiR.RodighieroC.RossM. T.SabotA.SankarS. V.ScallyA.SchrothG. P.SmithM. E.SmithV. P.SpiridouA.TorranceP. E.TzonevS. S.VermaasE. H.WalterK.WuX.ZhangL.AlamM. D.AnastasiC.AnieboI. C.BaileyD. M.BancarzI. R.BanerjeeS.BarbourS. G.BaybayanP. A.BenoitV. A.BensonK. F.BevisC.BlackP. J.BoodhunA.BrennanJ. S.BridghamJ. A.BrownR. C.BrownA. A.BuermannD. H.BunduA. A.BurrowsJ. C.CarterN. P.CastilloN.Chiara E. CatenazziM.ChangS.Neil CooleyR.CrakeN. R.DadaO. O.DiakoumakosK. D.Dominguez-FernandezB.EarnshawD. J.EgbujorU. C.ElmoreD. W.EtchinS. S.EwanM. R.FedurcoM.FraserL. J.Fuentes FajardoK. V.Scott FureyW.GeorgeD.GietzenK. J.GoddardC. P.GoldaG. S.GranieriP. A.GreenD. E.GustafsonD. L.HansenN. F.HarnishK.HaudenschildC. D.HeyerN. I.HimsM. M.HoJ. T.HorganA. M.HoschlerK.HurwitzS.IvanovD. V.JohnsonM. Q.JamesT.Huw JonesT. A.KangG. D.KerelskaT. H.KerseyA. D.KhrebtukovaI.KindwallA. P.KingsburyZ.Kokko-GonzalesP. I.KumarA.LaurentM. A.LawleyC. T.LeeS. E.LeeX.LiaoA. K.LochJ. A.LokM.LuoS.MammenR. M.MartinJ. W.McCauleyP. G.McNittP.MehtaP.MoonK. W.MullensJ. W.NewingtonT.NingZ.Ling NgB.NovoS. M.O’NeillM. J.OsborneM. A.OsnowskiA.OstadanO.ParaschosL. L.PickeringL.PikeA. C.PikeA. C.Chris PinkardD.PliskinD. P.PodhaskyJ.QuijanoV. J.RaczyC.RaeV. H.RawlingsS. R.Chiva RodriguezA.RoeP. M.RogersJ.Rogert BacigalupoM. C.RomanovN.RomieuA.RothR. K.RourkeN. J.RuedigerS. T.RusmanE.Sanches-KuiperR. M.SchenkerM. R.SeoaneJ. M.ShawR. J.ShiverM. K.ShortS. W.SiztoN. L.SluisJ. P.SmithM. A.Ernest Sohna SohnaJ.SpenceE. J.StevensK.SuttonN.SzajkowskiL.TregidgoC. L.TurcattiG.VandevondeleS.VerhovskyY.VirkS. M.WakelinS.WalcottG. C.WangJ.WorsleyG. J.YanJ.YauL.ZuerleinM.RogersJ.MullikinJ. C.HurlesM. E.McCookeN. J.WestJ. S.OaksF. L.LundbergP. L.KlenermanD.DurbinR.SmithA. J. (2008). Accurate whole human genome sequencing using reversible terminator chemistry. Nature456, 53–59.10.1038/nature07517
11
BlessedG.TomlinsonB. E.RothM. (1968). The association between quantitative measures of dementia and of senile change in the cerebral grey matter of elderly subjects. Br. J. Psychiatry114, 797–811.10.1192/bjp.114.512.797
12
BlossC. S.PawlikowskaL.SchorkN. J. (2010). Contemporary human genetic strategies in aging research. Ageing Res. Rev.10, 191–200.10.1016/j.arr.2010.07.005
13
BonafeM.BarbieriM.MarchegianiF.OlivieriF.RagnoE.GiampieriC.MugianesiE.CenturelliM.FranceschiC.PaolissoG. (2003). Polymorphic variants of insulin-like growth factor I (IGF-I) receptor and phosphoinositide 3-kinase genes affect IGF-I plasma levels and human longevity: cues for an evolutionarily conserved mechanism of life span control. J. Clin. Endocrinol. Metab.88, 3299–3304.10.1210/jc.2002-021810
14
ChristensenK.JohnsonT. E.VaupelJ. W. (2006). The quest for genetic determinants of human longevity: challenges and insights. Nat. Rev. Genet.7, 436–448.10.1038/nrg1871
15
ChuaK. F.MostoslavskyR.LombardD. B.PangW. W.SaitoS.FrancoS.KaushalD.ChengH. L.FischerM. R.StokesN.MurphyM. M.AppellaE.AltF. W. (2005). Mammalian SIRT1 limits replicative life span in response to chronic genotoxic stress. Cell Metab.2, 67–76.10.1016/j.cmet.2005.06.007
16
CirulliE. T.ZhuM.ShiannaK. V.GeD.GoldsteinD. B. (2011). Next Generation Sequencing of Centenarian and Control Genomes. Montreal: ACHG.
17
DrmanacR.SparksA. B.CallowM. J.HalpernA. L.BurnsN. L.KermaniB. G.CarnevaliP.NazarenkoI.NilsenG. B.YeungG.DahlF.FernandezA.StakerB.PantK. P.BaccashJ.BorcherdingA. P.BrownleyA.CedenoR.ChenL.ChernikoffD.CheungA.ChiritaR.CursonB.EbertJ. C.HackerC. R.HartlageR.HauserB.HuangS.JiangY.KarpinchykV.KoenigM.KongC.LandersT.LeC.LiuJ.McBrideC. E.MorenzoniM.MoreyR. E.MutchK.PerazichH.PerryK.PetersB. A.PetersonJ.PethiyagodaC. L.PothurajuK.RichterC.RosenbaumA. M.RoyS.ShaftoJ.SharanhovichU.ShannonK. W.SheppyC. G.SunM.ThakuriaJ. V.TranA.VuD.ZaranekA. W.WuX.DrmanacS.OliphantA. R.BanyaiW. C.MartinB.BallingerD. G.ChurchG. M.ReidC. A. (2009). Human genome sequencing using unchained base reads on self-assembling DNA nanoarrays. Science327, 78–81.10.1126/science.1181498
18
EbersbergerI.MetzlerD.SchwarzC.PaaboS. (2002). Genomewide comparison of DNA sequences between humans and chimpanzees. Am. J. Hum. Genet.70, 1490–1497.10.1086/340787
19
ErikssonM.BrownW. T.GordonL. B.GlynnM. W.SingerJ.ScottL.ErdosM. R.RobbinsC. M.MosesT. Y.BerglundP.DutraA.PakE.DurkinS.CsokaA. B.BoehnkeM.GloverT. W.CollinsF. S. (2003). Recurrent de novo point mutations in lamin A cause Hutchinson-Gilford progeria syndrome. Nature423, 293–298.10.1038/nature01629
20
EvertJ.LawlerE.BoganH.PerlsT. (2003). Morbidity profiles of centenarians: survivors, delayers, and escapers. J. Gerontol. A Biol. Sci. Med. Sci.58, 232–237.10.1093/gerona/58.3.M232
21
FraserG. E.ShavlikD. J. (2001). Ten years of life: is it a matter of choice?Arch. Intern. Med.161, 1645–1652.10.1001/archinte.161.13.1645
22
GattaL. B.VitaliM.ZanolaA.VenturelliE.FenoglioC.GalimbertiD.ScarpiniE.FinazziD. (2008). Polymorphisms in the LOC387715/ARMS2 putative gene and the risk for Alzheimer’s disease. Dement. Geriatr. Cogn. Disord.26, 169–174.10.1159/000151050
23
GrayM. D.ShenJ. C.Kamath-LoebA. S.BlankA.SopherB. L.MartinG. M.OshimaJ.LoebL. A. (1997). The Werner syndrome protein is a DNA helicase. Nat. Genet.17, 100–103.10.1038/ng0997-100
24
GuarenteL.KenyonC. (2000). Genetic pathways that regulate ageing in model organisms. Nature408, 255–262.10.1038/35041700
25
HerskindA. M.McGueM.HolmN. V.SorensenT. I.HarvaldB.VaupelJ. W. (1996). The heritability of human longevity: a population-based study of 2872 Danish twin pairs born 1870-1900. Hum. Genet.97, 319–323.10.1007/BF02185763
26
HindorffL. A.JunkinsH. A.MehtaJ. P.ManolioT. A. (2011). A Catalog of Published Genome-Wide Association Studies. Available at: http://www.genome.gov/gwastudies/ [accessed April 30, 2011].
27
HolstegeH.SieD.HarkinsT.LeeC.RossT.McLaughlinS.ShahM.YlstraB.MeijerG.Meijers-HeijboerH.HeutinkP.Shaw MurrayS.ReindersM.HolstegeG.SistermansE.LevyS. (2011). A Longevity Reference Genome Generated From the World’s Oldest Woman. Montreal: ACHG.
28
HolzenbergerM.DupontJ.DucosB.LeneuveP.GeloenA.EvenP. C.CerveraP.Le BoucY. (2003). IGF-1 receptor regulates lifespan and resistance to oxidative stress in mice. Nature421, 182–187.10.1038/nature01298
29
Huang daW.ShermanB. T.LempickiR. A. (2009). Systematic and integrative analysis of large gene lists using DAVID bioinformatics resources. Nat. Protoc.4, 44–57.10.1038/nprot.2008.211
30
KawasC.KaragiozisH.ResauL.CorradaM.BrookmeyerR. (1995). Reliability of the blessed telephone information-memory-concentration test. J. Geriatr. Psychiatry Neurol.8, 238–242.
31
KimJ. I.JuY. S.ParkH.KimS.LeeS.YiJ. H.MudgeJ.MillerN. A.HongD.BellC. J.KimH. S.ChungI. S.LeeW. C.LeeJ. S.SeoS. H.YunJ. Y.WooH. N.LeeH.SuhD.LeeS.KimH. J.YavartanooM.KwakM.ZhengY.LeeM. K.ParkH.KimJ. Y.GokcumenO.MillsR. E.ZaranekA. W.ThakuriaJ.WuX.KimR. W.HuntleyJ. J.LuoS.SchrothG. P.WuT. D.KimH.YangK. S.ParkW. Y.KimH.ChurchG. M.LeeC.KingsmoreS. F.SeoJ. S. (2009). A highly annotated whole-genome sequence of a Korean individual. Nature460, 1011–1015.
32
KojimaT.KameiH.AizuT.AraiY.TakayamaM.NakazawaS.EbiharaY.InagakiH.MasuiY.GondoY.SakakiY.HiroseN. (2004). Association analysis between longevity in the Japanese population and polymorphic variants of genes involved in insulin and insulin-like growth factor 1 signaling pathways. Exp. Gerontol.39, 1595–1598.10.1016/j.exger.2004.05.007
33
KopsG. J.DansenT. B.PoldermanP. E.SaarloosI.WirtzK. W.CofferP. J.HuangT. T.BosJ. L.MedemaR. H.BurgeringB. M. (2002). Forkhead transcription factor FOXO3a protects quiescent cells from oxidative stress. Nature419, 316–321.10.1038/nature01036
34
KuningasM.PuttersM.WestendorpR. G.SlagboomP. E.van HeemstD. (2007). SIRT1 gene, age-related diseases, and mortality: the Leiden 85-plus study. J. Gerontol. A Biol. Sci. Med. Sci.62, 960–965.10.1093/gerona/62.9.960
35
Kuro-oM.MatsumuraY.AizawaH.KawaguchiH.SugaT.UtsugiT.OhyamaY.KurabayashiM.KanameT.KumeE.IwasakiH.IidaA.Shiraki-IidaT.NishikawaS.NagaiR.NabeshimaY. I. (1997). Mutation of the mouse klotho gene leads to a syndrome resembling ageing. Nature390, 45–51.10.1038/36285
36
KurosuH.YamamotoM.ClarkJ. D.PastorJ. V.NandiA.GurnaniP.McGuinnessO. P.ChikudaH.YamaguchiM.KawaguchiH.ShimomuraI.TakayamaY.HerzJ.KahnC. R.RosenblattK. P.Kuro-oM. (2005). Suppression of aging in mice by the hormone Klotho. Science309, 1829–1833.10.1126/science.1112766
37
LeeS. S.KennedyS.TolonenA. C.RuvkunG. (2003). DAF-16 target genes that control C. elegans life-span and metabolism. Science300, 644–647.10.1126/science.1083614
38
LevyS.SuttonG.NgP. C.FeukL.HalpernA. L.WalenzB. P.AxelrodN.HuangJ.KirknessE. F.DenisovG.LinY.MacDonaldJ. R.PangA. W.ShagoM.StockwellT. B.TsiamouriA.BafnaV.BansalV.KravitzS. A.BusamD. A.BeesonK. Y.McIntoshT. C.RemingtonK. A.AbrilJ. F.GillJ.BormanJ.RogersY. H.FrazierM. E.SchererS. W.StrausbergR. L.VenterJ. C. (2007). The diploid genome sequence of an individual human. PLoS Biol.5, e254.10.1371/journal.pbio.0050254
39
LiH.DurbinR. (2009). Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics25, 1754–1760.10.1093/bioinformatics/btp100
40
LiH.HandsakerB.WysokerA.FennellT.RuanJ.HomerN.MarthG.AbecasisG.DurbinR.1000 Genome Project Data Processing Subgroup. (2009). The Sequence Alignment/Map format and SAMtools. Bioinformatics25, 2078–2079.10.1093/bioinformatics/btp100
41
LohmuellerK. E.IndapA. R.SchmidtS.BoykoA. R.HernandezR. D.HubiszM. J.SninskyJ. J.WhiteT. J.SunyaevS. R.NielsenR.ClarkA. G.BustamanteC. D. (2008). Proportionally more deleterious genetic variation in European than in African populations. Nature451, 994–997.10.1038/nature06611
42
MahoneyF. I.BarthelD. W. (1965). Functional evaluation: the Barthel Index. Md. State Med. J.14, 61–65.
43
MetzkerM. L. (2009). Sequencing technologies – the next generation. Nat. Rev. Genet.11, 31–46.10.1038/nrg2626
44
MooreB.HuH.SingletonM.ReeseM. G.De La VegaF. M.YandellM. (2011). Global analysis of disease-related DNA sequence variation in 10 healthy individuals: implications for whole genome-based clinical diagnostics. Genet. Med.13, 210–217.10.1097/GIM.0b013e31820ed321
45
NallsM. A.Simon-SanchezJ.GibbsJ. R.Paisan-RuizC.BrasJ. T.TanakaT.MatarinM.ScholzS.WeitzC.HarrisT. B.FerrucciL.HardyJ.SingletonA. B. (2009). Measures of autozygosity in decline: globalization, urbanization, and its implications for medical genetics. PLoS Genet.5, e1000415.10.1371/journal.pgen.1000415
46
NuzzoA.RivaA. (2009). Genephony: a knowledge management tool for genome-wide research. BMC Bioinformatics10, 278.10.1186/1471-2105-10-278
47
PaynterN. P.ChasmanD. I.PareG.BuringJ. E.CookN. R.MiletichJ. P.RidkerP. M. (2010). Association between a literature-based genetic risk score and cardiovascular events in women. JAMA303, 631–637.10.1001/jama.2010.119
48
PawlikowskaL.HuD.HuntsmanS.SungA.ChuC.ChenJ.JoynerA. H.SchorkN. J.HsuehW. C.ReinerA. P.PsatyB. M.AtzmonG.BarzilaiN.CummingsS. R.BrownerW. S.KwokP. Y.ZivE.Study of Osteoporotic Fractures. (2009). Association of common genetic variation in the insulin/IGF1 signaling pathway with human longevity. Aging Cell8, 460–472.10.1111/j.1474-9726.2009.00493.x
49
PerlsT.Shea-DrinkwaterM.Bowen-FlynnJ.RidgeS. B.KangS.JoyceE.DalyM.BrewsterS. J.KunkelL.PucaA. A. (2000). Exceptional familial clustering for extreme longevity in humans. J. Am. Geriatr. Soc.48, 1483–1485.
50
PerlsT. T.WilmothJ.LevensonR.DrinkwaterM.CohenM.BoganH.JoyceE.BrewsterS.KunkelL.PucaA. (2002). Life-long sustained mortality advantage of siblings of centenarians. Proc. Natl. Acad. Sci. U.S.A.99, 8442–8447.10.1073/pnas.122587599
51
SchachterF.Faure-DelanefL.GuenotF.RougerH.FroguelP.Lesueur-GinotL.CohenD. (1994). Genetic associations with human longevity at the APOE and ACE loci. Nat. Genet.6, 29–32.10.1038/ng0194-29
52
SchoenmakerM.de CraenA. J.de MeijerP. H.BeekmanM.BlauwG. J.SlagboomP. E.WestendorpR. G. (2006). Evidence of genetic enrichment for exceptional survival using a family approach: the Leiden Longevity Study. Eur. J. Hum. Genet.14, 79–84.
53
SebastianiP.SolovieffN.DeWanA.WalshK.PucaA.HartleyS. W.MelistaE.AndersenS.DworkisD. A.WilkJ. B.MyersR. H.SteinbergM. H.MontanoM.BaldwinC. T.HohJ.PerlsT. T. (2012). Genetic signatures of exceptional longevity in humans. PLoS ONE.10.1371/journal.pone.0029848
54
SinoffG.OreL. (1997). The Barthel activities of daily living index: self-reporting versus actual performance in the old-old (> or = 75 years). J. Am. Geriatr. Soc.45, 832–836.
55
SmithE. N.KollerD. L.PanganibanC.SzelingerS.ZhangP.BadnerJ. A.BarrettT. B.BerrettiniW. H.BlossC. S.ByerleyW.CoryellW.EdenbergH. J.ForoudT.GershonE. S.GreenwoodT. A.GuoY.HipolitoM.KeatingB. J.LawsonW. B.LiuC.MahonP. B.McInnisM. G.McMahonF. J.McKinneyR.MurrayS. S.NievergeltC. M.NurnbergerJ. I.Jr.NwuliaE. A.PotashJ. B.RiceJ.SchulzeT. G.ScheftnerW. A.ShillingP. D.ZandiP. P.ZöllnerS.CraigD. W.SchorkN. J.KelsoeJ. R. (2011). Genome-wide association of bipolar disorder suggests an enrichment of replicable associations in regions near genes. PLoS Genet.7, e1002134.10.1371/journal.pgen.1002134
56
SolovieffN.HartleyS. W.BaldwinC. T.PerlsT. T.SteinbergM. H.SebastianiP. (2010). Clustering by genetic ancestry using genome-wide data. BMC Genet.11, 108.10.1186/1471-2164-11-108
57
SongF.JiP.ZhengH.WangY.HaoX.WeiQ.ZhangW.ChenK. (2010). Definition of a functional single nucleotide polymorphism in the cell migration inhibitory gene MIIP that affects the risk of breast cancer. Cancer Res.70, 1024–1032.10.1158/1538-7445.AM10-1024
58
StensonP. D.MortM.BallE. V.HowellsK.PhillipsA. D.ThomasN. S.CooperD. N. (2009). The Human Gene Mutation Database: 2008 update. Genome Med.1, 13.10.1186/gm13
59
SuhY.AtzmonG.ChoM. O.HwangD.LiuB.LeahyD. J.BarzilaiN.CohenP. (2008). Functionally significant insulin-like growth factor I receptor mutations in centenarians. Proc. Natl. Acad. Sci. U.S.A.105, 3438–3442.10.1073/pnas.0705467105
60
TanQ.ZhaoJ. H.ZhangD.KruseT. A.ChristensenK. (2008). Power for genetic association study of human longevity using the case-control design. Am. J. Epidemiol.168, 890–896.10.1093/aje/kwn205
61
TerryD. F.SebastianiP.AndersenS. L.PerlsT. T. (2008). Disentangling the roles of disability and morbidity in survival to exceptional old age. Arch. Intern. Med.168, 277–283.10.1001/archinternmed.2007.75
62
The 1000 Genomes Project Consortium. (2010). A map of human genome variation from population-scale sequencing. Nature467, 1061–1073.10.1038/nature09534
63
TorkamaniA.Scott-Van ZeelandA. A.TopolE. J.SchorkN. J. (2011). Annotating individual human genomes. Genomics98, 233–241.10.1016/j.ygeno.2011.07.006
64
VaziriH.DessainS.EatonE. N.ImaiiS. I.FryeA. R.PanditaT. K.GuarenteL.WeinbergR. A. (2001). hSir2 (SirT1) functions as an NAD-dependent p53 deacetylase. Cell107, 149–159.10.1016/S0092-8674(01)00527-X
65
WangJ.WangW.LiR.LiY.TianG.GoodmanL.FanW.ZhangJ.LiJ.ZhangJ.GuoY.FengB.LiH.LuY.FangX.LiangH.DuZ.LiD.ZhaoY.HuY.YangZ.ZhengH.HellmannI.InouyeM.PoolJ.YiX.ZhaoJ.DuanJ.ZhouY.QinJ.MaL.LiG.YangZ.ZhangG.YangB.YuC.LiangF.LiW.LiS.LiD.NiP.RuanJ.LiQ.ZhuH.LiuD.LuZ.LiN.GuoG.ZhangJ.YeJ.FangL.HaoQ.ChenQ.LiangY.SuY.SanA.PingC.YangS.ChenF.LiL.ZhouK.ZhengH.RenY.YangL.GaoY.YangG.LiZ.FengX.KristiansenK.WongG. K.NielsenR.DurbinR.BolundL.ZhangX.LiS.YangH.WangJ. (2008). The diploid genome sequence of an Asian individual. Nature456, 60–65.10.1038/nature07509
66
WangJ. H.PappasD.De JagerP. L.PelletierD.de BakkerP. I.KapposL.PolmanC. H.Australian and New Zealand Multiple Sclerosis Genetics Consortium (ANZgene)ChibnikL. B.HaflerD. A.MatthewsP. M.HauserS. L.BaranziniS. E.OksenbergJ. R. (2011). Modeling the cumulative genetic risk for multiple sclerosis from genome-wide association data. Genome Med.3, 3.10.1186/gm225
67
WegnerL.AnthonsenS.Bork-JensenJ.DalgaardL.HansenT.PedersenO.PoulsenP.VaagA. (2010). LMNA rs4641 and the muscle lamin A and C isoforms in twins – metabolic implications and transcriptional regulation. J. Clin. Endocrinol. Metab.95, 3884–3892.10.1210/jc.2009-2675
68
WheelerH. E.KimS. K. (2011). Genetics and genomics of human ageing. Philos. Trans. R. Soc. Lond. B Biol. Sci.366, 43–50.10.1098/rstb.2010.0259
69
WillcoxB. J.DonlonT. A.HeQ.ChenR.GroveJ. S.YanoK.MasakiK. H.WillcoxD. C.RodriguezB.CurbJ. D. (2008). FOXO3A genotype is strongly associated with human longevity. Proc. Natl. Acad. Sci. U.S.A.105, 13987–13992.10.1073/pnas.0801030105
70
ZhangY.DeS.GarnerJ. R.SmithK.WangS. A.BeckerK. G. (2010). Systematic analysis, comparison, and integration of disease based human genetic association data and mouse genetic phenotypic information. BMC Med. Genomics3, 1.10.1186/1755-8794-3-1
Appendix
Figure A1

Schematic of the mapping/alignment and variant calling steps. We used the Eland and BWA aligners to map reads to the reference genome, and CASAVA and SAMTools to call SNPs as explained in Section “Materials and Methods.” Short insertion and deletions were called using SAMTools and Dindel. Only variants that were called by both algorithms and passed quality control filters were included in the follow-up analyses.
Figure A2

(A) Number of SNPs called by both SNP callers, the CASAVA algorithm and SAMTools, with >99% concordant genotype calls, by chromosome. (B) Distribution of Phred scores in SNPs called in PG17 and PG26 (From SAMTools). (C) Distribution of reads used for SNPs calling in PG17 and PG26 (From CASAVA). The median number of used reads was 36 in PG17 and 43 in PG26.
Figure A3

Plot of the B allele frequency (a measure of heterozygosity/homozygosity) and the log R ratio (a measure of signal intensity that relates to copy number variations) using Illumina BeadStudio. The plot of B allele frequency shows that there are significant stretches of homozygosity across many chromosomes (see red arrows). However, this was not associated with a shift in log R ratio at these locations suggesting that this was not due to big deletions but probably to inbreeding of her ancestors with distant relatives that resulted in higher homozygosity.
Table A1
| Illumina report (hg18) | BWA (hg18) | |||
|---|---|---|---|---|
| Male >110 | Female >110 | Male >110 | Female >110 | |
| No of reads | 1,804,595482 | 1,650,463,996 | 1,804,595,182 | 1,650,463,996 |
| Read length | 100 | 100 | 100 | 100 |
| Number of bases used to generate consensus | 114,850,993,633 | 98,120,678,830 | 156,017,244,672 | 133,347,354,219 |
| Base pairs reported (consensus) | 2,736,567,883 | 2,699,489,763 | 2,855,361,624 | 2,828,277,701 |
| Number aligned reads | 1,430,677,828 (79.3%) | 1,246,883,285 (75.5%) | 1,559,315,652 (86.41%) | 1,334,406,235 (80.85%) |
| Rate of genome covered | 0.957 | 0.952 | 0.999 | 0.998 |
| Coverage depth | 43.14 | 36.40 | 54.60 | 47.04 |
Summary of sequencing and mapping using the Illumina aligner and the BWA aligner.
The parameters used with the aligner algorithms are in the Section “Materials and Methods.”
Figure A4

Plot of the B allele frequency and the log R ratio in PG26. The plot of B allele frequency shows that there are no large structural variations.
Figure A5

Trends in QC parameters. We evaluated the effect of tighter thresholds on the minimum coverage (=number of reads) on the transition to transversion ratio (A), the heterozygous to homozygous ratios (B) and the mean depths (C). The quality parameters appear to be stable for different thresholds on minimum coverage.
Figure A6

Cumulative distributions of depth. The plots show the cumulative distributions of SNPs with coverage < values in the x-axes. For example, 80% of the reads in PG17 have more than 20-fold coverage.
Figure A7

Distribution of coverage and Phred scores of indels.
Figure A8

Functional analysis of the insertions detected in PG17 and PG26, their union and intersection. The chart in maroon shows the breakdown of insertions and the rate of coding insertions in PG17 and PG26. The chart in orange shows the summary of insertions in nine whole genomes generated from complete genomics and annotated in the same way as PG17 and PG26.
Figure A9

Functional analysis of the deletions detected in PG17 and PG26, their union and intersection. The chart in maroon shows the breakdown of deletions and the rate of coding deletions in PG17 and PG26. The chart in orange shows the summary of deletions in nine whole genomes generated from complete genomics and annotated in the same way as PG17 and PG26.
Figure A10

Population structure of the two supercentenarians. The two scatter plots display the principal components 1 and 2 (PCI and 2 PC2, top panels), and principal components 3 and 4 (PC3 and PC4, bottom panels) in 801 subjects from the NECS that were estimated using genome-wide data. We colored the points by one of 16 ancestral groups that were inferred using an algorithm described in Solovieff et al. (
Figure A11

Genomic Coverage from ELAND. Barplots show the % of genomic coverage per chromosome when the reads were mapped to HG18 using the Illumina Elander aligner.
Table A2
List of genome sequences that were used for comparisons with PG17 and PG26.
Table A3
| NECS | Watson | Venter | Complete genomics | AK1 | YH | Illumina | 1000 Genomes | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Male >111 | Female >113 | NA07022 (CEU/M) | NA20431 (PG/M) | NA19240 (YRI/F) | NA 18507 (YRI/m) | CEU trio | YRI trio | |||||
| Ti/Tv (in dbSNP) | 2.11 | 2.07 | 2.01 | 2.04 | 2.17 | 2.19 | 2.13 | 2.07 | 2.01 | 2.06 | 2.07 | 2.09 |
| Ti/Tv (only 1000 genomes) | 2.33 | 2.27 | ||||||||||
| Ti/Tv (Novel) | 2.07 | 2.13 | 1.94 | 2.02 | ||||||||
| het/homo (in dbSNP) | 1.41 | 1.24 | 1.30 | 1.2 | 1.64 | 1.72 | 2.01 | 1.57 | 1.27 | 1.75 | ||
| het/homo (only 1000 genomes) | 14.57 | 7.12 | ||||||||||
| het/homo (Novel) | 6.28 | 9.16 | 12.3 | 27.48 | 13.73 | |||||||
| Mean mapped depth | 41.64 | 35.05 | 7.40 | 7.5 | 87 | 45 | 63 | 29 | 36 | 41 | 43 | 40 |
| Mean mapped depth (only 1000 genomes) | 35.51 | 30.02 | ||||||||||
| Mean mapped depth (Novel) | 37.93 | 34.08 | ||||||||||
| CGI | http://www.completegenomics.com/sequence-data/data-summary/ | |||||||||||
| Watson/Venter/AK1/YH | http://www.nature.com/nature/journal/v460/n7258/full/nature08211.html | |||||||||||
| Illumina | http://www.nature.com/nature/journal/v456/n7218/full/nature07517.html | |||||||||||
| 1000 Genomes | http://www.nature.com/nature/journal/v467/n7319/full/nature09534.html | |||||||||||
| Venter | http://www.plosbiology.org/article/info:doi/10.1371/journal.pbio.0050254 | |||||||||||
Comparison of SNPs summaries in different genomes.
The comparative values were taken from published references noted in the table. These references did not report the summaries stratified by whether they were novel and/or reported only in 1000 genomes database.
Table A4
| Source | NECS | Complete genomic (2009) | Watson | CJ Venter | AK1 | YH (Chinese) | Illumina | 1000 Genomes | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SNPs | ||||||||||||
| Variation type | PG17 (Fern CEU) | PG26 (Male CEU) | NA19240 (Fem YRI) | NA07022 (Male CEU) | NA20431 (Male PG) | Male (CEU) | Male (CEU) | Male (Korean) | YH (Chinese) | NA18507 (Male YRI) | CEU trio | YRI trio |
| All | 3,084,838 (1.4) | 3,253,896 (1.5) | 4,042,801 (19) | 3,076,869 (10) | 2,905,517 (10) | 3,322,093 (18) | 3,075,281 (18) | 3,453,653 (17) | 3,074,097 (14) | 4,139,196 (16) | 3,646,764 (11) | 4,502,439 (23) |
| Homozygous | 1,378,735 (0.37) | 1,352,646 (0.52) | 1,297,601 (4) | 1,097,899 (2) | 965,029 (1) | 1,464,414 | 1,450,860 | 1,343,250 (4) | 1,351,824 | 1,503,420 | ||
| Heterozygous | 1,706,103 (2.54) | 1,901,250 (2.18) | 2,639,864 (27) | 1,800,287 (15) | 1,657,540 (16) | 1,857,679 | 1,624,421 | 2,110,403 (25) | 1,722,273 | 2,635,776 | ||
| Transitions | 2,080,095 (1.46) | 2,208,983 (1.48) | 3,635,882 | 2,858,818 | 2,658,112 | |||||||
| Transversions | 1,004,743 (1.43) | 1,044,913 (1.53) | 1,706,195 | 1,316,837 | 1,213,232 | |||||||
| Coding | 19,581 (1.9) | 21,377 (2.1) | 23,000 (16) | 18,723 (9) | 16,532 (10) | 21,152 | 27,762 | 15,686 | 26,140 | |||
| INDELS | ||||||||||||
| Non-synonymous | 9,152 (2.5) | 10,022 (2.8) | 11,400 (19) | 9,286 (U) | 8,215 (12) | 10,569 | 6,889 | 10,162 | 7,062 | 10,875 | 8,299–11,122 | 10,349–11,122 |
| All | 338,455 | 370,266 | 496,194 | 337,635 | 269,794 | 227,718 | 214,691 | 170,202 (62) | 135,262 (59) | 404,416 (50) | 411,611 (25) | 502,462 (37) |
| Short insertions | 169,833 (63.9) | 187,058 (65.2) | 242,391 (40) | 168,909 (37) | 136,786 (37) | 188,388 (20) | 226,361 (30) | |||||
| Short deletions | 168,622 (53.3) | 183,208 (53.5) | 253,803 (44) | 168,726 (37) | 133,008 (36) | 225,189 (29) | 277,036 (43) | |||||
| Coding short indels | 314 (56.4) | 420 (58.3) | 549 (56) | 556 (58) | 435 (59) | 345 | 863 | 212 | 65 | 476 (21) | 658 (33) | |
| Complete genomics | http://www.completegenomics.com/sequence-data/data-summary/ | |||||||||||
| Watson/Venter/AK1/Yh | http://www.nature.com/nature/journal/v460/n7258/full/nature08211.html | |||||||||||
| Illumina | http://www.nature.com/nature/journal/v456/n7218/full/nature07517.html | |||||||||||
| 1000 Genomes | http://www.nature.com/nature/journal/v467/n7319/full/nature09534.html | |||||||||||
| Venter | http://www.nature.com/nature/journal/v467/n7319/full/nature09534.html | |||||||||||
Comparison of SNPs summaries in different genomes.
The comparative values were taken from published references noted in the table. Cells report the counts and rates of novel variants when available.
Table A5
| PG17 (all SNPs) | PG26 (all SNPs) | PG17 (all ins) | PG26 (all ins) | PG17 (all del) | PG26 (all del) | |
|---|---|---|---|---|---|---|
| miRNA | ||||||
| Total number of variants creating miRNA binding sites | 5,965 | 6,150 | 208 | 241 | 265 | 298 |
| Total number of variants destroying miRNA binding sites | 6,007 | 6,267 | 298 | 320 | 197 | 192 |
| Total number of variants perturbing miRNA binding sites (diff deltaG > 0) | 13,190 | 13,655 | 880 | 938 | 808 | 877 |
| TFB | ||||||
| Total number of TFBS disrupting variants | 587,761 | 638,857 | 28,896 | 33,234 | 50,997 | 58,709 |
| Total number of major TFBS disrupting variants (dScore < −7) | 4,037 | 4,619 | 7,560 | 9,005 | 30,123 | 35,117 |
| Total number of TFBS deleting variants: | 52 | 62 | 0 | 0 | 416 | 524 |
| SPLICE | ||||||
| Total number of splice site variants | 3,385 | 4,284 | 306 | 352 | 269 | 308 |
| Total number of “splicing change” variants | 23 | 28 | 85 | 104 | 102 | 129 |
| Total number of variants creating exonic splicing enhancer binding sites | 6,628 | 7,240 | 39 | 66 | 189 | 200 |
| SITES | ||||||
| Total number of variants destroying exonic splicing enhancer binding sites | 6,646 | 7,300 | 50 | 65 | 193 | 211 |
| Total number of variants creating exonic splicing silencer binding sites | 3,738 | 4,038 | 38 | 43 | 176 | 179 |
| Total number of variants destroying exonic splicing silencer binding sites | 3,607 | 3,982 | 14 | 23 | 158 | 153 |
Predicted functions (TFBS and miRs).
The functional annotation was done with the complementary genome-wide variant annotation tools embedded in a suite of tools developed by researchers at The Scripps Research Institute (see reference Torkamani et al.,
Table A6
| PG17 (all SNPs, %) | PG26 (all SNPs, %) | PG17 (all ins, %) | PG26 (all ins, %) | PG17 (all del, %) | PG26 (all del, %) | |
|---|---|---|---|---|---|---|
| miRNA | ||||||
| Total number of variants creating miRNA binding sites | 0.19 | 0.19 | 0.12 | 0.13 | 0.16 | 0.16 |
| Total number of variants destroying miRNA binding sites | 0.19 | 0.19 | 0.18 | 0.17 | 0.12 | 0.10 |
| Total number of variants perturbing miRNA binding sites (diff deltaG > 0) | 0.43 | 0.42 | 0.52 | 0.50 | 0.48 | 0.48 |
| TFB | ||||||
| Total number of TFBS disrupting variants | 19.05 | 19.63 | 17.01 | 17.77 | 30.24 | 32.04 |
| Total number of major TFBS disrupting variants (dScore < −7) | 0.13 | 0.14 | 4.45 | 4.81 | 17.86 | 19.17 |
| Total number of TFBS deleting variants | 0.00 | 0.00 | 0.00 | 0.00 | 0.25 | 0.29 |
Rates of SNPs and indels that affect microRNAs and transcription factor binding sites (TFBS).
The functional annotation was done with the complementary genome-wide variant annotation tools embedded in a suite of tools developed by researchers at The Scripps Research Institute (see reference Torkamani et al.,
Summary
Keywords
whole genome sequence, genetics, longevity, centenarian, supercentenarian, aging
Citation
Sebastiani P, Riva A, Montano M, Pham P, Torkamani A, Scherba E, Benson G, Milton JN, Baldwin CT, Andersen S, Schork NJ, Steinberg MH and Perls TT (2012) Whole Genome Sequences of a Male and Female Supercentenarian, Ages Greater than 114 Years. Front. Gene. 2:90. doi: 10.3389/fgene.2011.00090
Received
20 October 2011
Accepted
04 December 2011
Published
03 January 2012
Volume
2 - 2011
Edited by
Thomas Flatt, Vetmeduni Vienna, Austria
Reviewed by
Josephine Hoh, Yale University, USA; Stuart Kim, Stanford University, USA
Copyright
© 2012 Sebastiani, Riva, Montano, Pham, Torkamani, Scherba, Benson, Milton, Baldwin, Andersen, Schork, Steinberg and Perls.
This is an open-access article distributed under the terms of the Creative Commons Attribution Non Commercial License, which permits non-commercial use, distribution, and reproduction in other forums, provided the original authors and source are credited.
*Correspondence: Paola Sebastiani, Department of Biostatistics, Boston University School of Public Health, 801 Massachusetts Avenue, Room 317, Boston, MA 02118, USA. e-mail: sebas@bu.edu; Thomas T. Perls, Geriatrics Section, Boston University Medical Campus, 88 E Newton St, Boston, MA 02118, USA. e-mail: thperls@bu.edu
This article was submitted to Frontiers in Genetics of Aging, a specialty of Frontiers in Genetics.
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.