Abstract
Mutations in genes potentially lead to a number of genetic diseases with differing severity. These disease genes have been the focus of research in recent years showing that the disease gene population as a whole is not homogeneous, and can be categorized according to their interactions. Locus heterogeneity describes a single disorder caused by mutations in different genes each acting individually to cause the same disease. Using datasets of experimentally derived human disease genes and protein interactions, we created a protein interaction network to investigate the relationships between the products of genes associated with a disease displaying locus heterogeneity, and use network parameters to suggest properties that distinguish these disease genes from the overall disease gene population. Through the manual curation of known causative genes of 100 diseases displaying locus heterogeneity and 397 single-gene Mendelian disorders, we use network parameters to show that our locus heterogeneity network displays distinct properties from the global disease network and a Mendelian network. Using the global human proteome, through random simulation of the network we show that heterogeneous genes display significant interconnectivity. Further topological analysis of this network revealed clustering of locus heterogeneity genes that cause identical disorders, indicating that these disease genes are involved in similar biological processes. We then use this information to suggest additional genes that may contribute to diseases with locus heterogeneity.
INTRODUCTION
The characterization of mutations in genes that cause human genetic disease is vitally important. Once identified, these mutant genes (termed disease genes) provide an opportunity to study the origins of genetic disorders and develop potential therapeutics to mitigate symptoms or deliver curative strategies targeting these genes. In recent years, the discovery and classification of disease genes within the human genome has received increasing attention. As databases of disease gene associations, such as the Online Mendelian Inheritance in Man (OMIM; ), continue to increase in size and accuracy, we can use these data to further understand disease pathogenesis. In a previous study () we found that disease genes do not form a homogeneous group of genes with shared characteristics – but instead cluster into distinct groups each with shared characteristics. Isolating genes displaying similar attributes may therefore lead to the discovery of further associated gene groups, allowing us to examine their relationship with disease.
Is it now appreciated that human disease is characterized by genetic heterogeneity, for which two different types exist. Allelic heterogeneity refers to instances where mutations in different alleles at the same locus produce the same disease. By contrast, locus heterogeneity describes mutations in different genes whereby any one mutation generates the same disorder (Figure 1; ). Many genetic diseases display locus heterogeneity, with affected genes being associated with almost all disease categories and cell types. Perhaps the most striking example of locus heterogeneity is the disorder retinitis pigmentosa, a retinal dystrophy resulting from the loss of photoreceptors in the retina for which more than 45 genes have been identified (). A number of recent studies into the mechanisms by which these genes cause identical disorders suggest that protein products of affected genes are likely to be functionally similar, interacting with one another and displaying an involvement in the same biological pathways and processes (Wang et al., 2012; ). With this is mind, an appropriate method to study the associations between genes involved in these disorders is to investigate the complex interconnections between cellular components.
FIGURE 1
The advent of high-throughput, ‘omic’ technologies in the last decade has resulted in rapid growth in the number of identified and mapped protein interactions available within interaction databases. For example, BioGRID (Stark et al., 2006) provides genetic and biological interaction data for a range of species and the Human Protein Reference Database (
Existing network analysis based studies have utilized the analytical advantages of interaction networks to reveal the highly interconnected relationships between genes expressing locus heterogeneity. A study by
In this study we tested the hypothesis that within protein interaction networks, locus heterogeneity genes are more highly interconnected to other genes causing the same disorder than genes associated with Mendelian diseases or non-disease genes. Throughout, locus heterogeneity disorders were classed as those caused by mutations in a number of genes, but inherited in a monogenic/simple fashion. Complex heterogeneous disorders caused by mutations in multiple alleles acting together were not considered here. To complete our investigations we manually curated a number of locus heterogeneous disorders and their associated genes. We generated a global human protein interaction network from various human interaction databases. By considering the local neighborhood of heterogeneous genes, we were able to identify potential novel locus heterogeneity genes involved in specific disorders. A comparison of the locus heterogeneity curated genes with those that cause single-gene Mendelian disorders served as a method to isolate and identify properties of locus heterogeneity genes. The results of this study demonstrate that locus heterogeneity genes display distinct network properties, forming clusters of disorder specific genes. These network clusters can be utilized to suggest novel disease genes for further experimental studies.
MATERIALS AND METHODS
DATA RETRIEVAL
Disease genes were parsed from the OMIM database genemap (03/02/2014 update;
Disease gene data relating to heterogeneous and Mendelian disorders were obtained from a combination of ResNet (10/02/2014 update;
Human protein–protein interaction data was retrieved using ConsensusPathDB (CPDB, release 28;
DISEASE CATEGORIZATION
Genes were classified into appropriate disease categories using the Medical Subject Headings controlled vocabulary (MeSH;
NETWORK VISUALIZATION AND TOPOLOGICAL ANALYSIS
Protein–protein interaction networks were visualized and analyzed using Cytoscape (version 2.8.3 and version 3.1.0; Shannon et al., 2003). All networks presented here are undirected and use the edge-weighted spring embedded layout, unless otherwise stated, and have had self-loops and duplicated edges removed. The Cytoscape plugin AllegroLayout (
Topological analysis of the network was achieved within Cytoscape using the clustering tool AllegroMCODE 2.1 (
Additional methods were utilized to validate selected clusters. The Louvain method for network community analysis attempt to reveal a hierarchical structure for larger networks, discussed in
GENE FUNCTIONAL AND PATHWAY ANALYSIS
Identifying key properties of unannotated genes found with disease enriched clusters was achieved using Ingenuity Pathway Analysis (IPA;
The Cytoscape plugin BiNGO 3.0.2 (
STATISTICAL ANALYSIS
Statistical analyses were performed using R-Development-Core-Team (2009). Pearson’s Chi-squared test was used to assess whether disease classifications was significantly different between heterogeneity and Mendelian datasets. The Benjamini and Hochberg False Discovery Rate was used to calculate corrected p-values for GO functional classification testing to minimize multiple comparison errors.
A Perl script utilizing the Graph module (
RESULTS
LOCUS HETEROGENEITY AND MENDELIAN DISORDER CLASSIFICATION
Using a combination of OMIM’s genemap (
FIGURE 2

Proportional display of diseases by MESH classification. The proportion of locus heterogeneity (left) and Mendelian (right) disease genes characterized in our study that affect different physiological systems. Colors correspond to specific physiological systems affected by these disease genes (key at far right).
In order to prevent any potential bias, we chose Mendelian disease genes to include in our dataset because they shared the same disease classification proportions as our locus heterogeneity genes. It was not possible to eliminate all variation between the two datasets, however, these differences have been minimized by the selection of Mendelian disorders affecting the same physiological systems as those affected in diseases showing locus heterogeneity. A Pearson’s Chi-squared test confirmed that the two datasets were not significantly different in the systems affected (p = 0.372).
LOCUS HETEROGENEITY NETWORKS SHOW DISTINCT PROPERTIES COMPARED TO OTHER DISEASE-ASSOCIATED NETWORKS
The full human protein–protein interaction network was retrieved from CPDB, consisting of 16363 nodes and 179685 edges (Figure 3). Since this interaction data is sourced from a number of interaction databases and experimental studies, the resulting collection of data contains protein interactions from multiple sources, such as co-immunoprecipitation and yeast two-hybrid studies. To extract and analyze specific networks in isolation, the proteins encoded by disease genes, locus heterogeneity genes and Mendelian genes were mapped onto the network. Although a total of 674 locus heterogeneity genes and 397 Mendelian genes were identified from ResNet (
FIGURE 3

Full CPDB protein interaction network. The network displays the full set of interactions available from CPDB used in this study. Circles (nodes) represent proteins, whereas the lines (edges) connecting two circles signify an interaction between two proteins. Locus heterogeneity genes relating to our 100 selected disorders are highlighted red, with gray nodes symbolizing other genes in the dataset.
Table 1
| Full networks | Largest connected component | |||||
|---|---|---|---|---|---|---|
| Full disease | Heterogeneity | Mendelian | Full disease | Heterogeneity | Mendelian | |
| Number of nodes | 2485 | 535 | 301 | 2040 (82.1%) | 323 (60.4%) | 134 (44.5%) |
| Average degree | 7.305 | 2.931 | 1.362 | 8.881 | 4.669 | 2.866 |
| Isolated nodes | 415 (16.7%) | 163 (30.5%) | 148 (49.2%) | N/A | N/A | N/A |
| Network centralization | 0.113 | 0.128 | 0.049 | 0.137 | 0.207 | 0.100 |
| Clustering coefficient | 0.119 | 0.141 | 0.049 | 0.145 | 0.233 | 0.094 |
Disease network parameters.
Network properties were calculated for each of the three full disease networks, and the largest connected component of these networks. The LCC calculations ignore isolated nodes and clusters. For specific parameters, the percentage of nodes within the full network displaying each property is listed in parentheses.
Analysis was performed on both the full network and the largest connected component (the largest interconnected group of nodes within the network, LCC) to exclude disconnected nodes. Initial parameter calculations revealed a large percentage of isolated nodes (nodes with a degree value of 0) within the three networks. As detailed in previous studies (
In both the full networks and the LCC networks, average degree (the average number of interactions across all nodes) is largest in the disease network and lowest in the Mendelian network. Although the full disease network has a larger average degree, we would expect to observe clustering in the heterogeneous network due to the perturbation of different genes causing identical disorders as a result of their functional pathway similarities (
Additionally, we analyzed clustering in the various disease networks. The average clustering coefficient characterizes the tendency of nodes to form highly connected clusters, used previously by Ravasz et al. (2002) to study the modular organization of metabolic networks. Our data show that the locus heterogeneity network has the largest average clustering coefficient of the three disease networks for both the full network and the LCC. This suggests that locus heterogeneity genes form groups of highly interconnected clusters, confirming the prediction that gene-products causing the same disorder interact with each other.
LOCUS HETEROGENEITY GENES SHOW SIGNIFICANT INTERCONNECTIVITY WITHIN THE GLOBAL PROTEIN INTERACTION NETWORK
To investigate the connectivity of locus heterogeneity associated proteins within the full CPDB interaction network, we utilized the Perl module package Graph (
The connectivity of actual locus heterogeneity proteins within the network was 79.7%, which was significantly higher than the connectivity in any of our random simulations, which displayed a mean connectivity value of 41.9% (p < 0.0001; Figure 4). Whilst showing that the connectivity of locus heterogeneity genes is higher than expected by chance, this test also further confirms the high degree of connectivity of heterogeneity genes within our interaction network.
FIGURE 4

Locus heterogeneity gene interconnectivity within the full CPDB network compared to random simulations. There is a normal distribution of random simulations (black bars), with a mean value of 41.9%. The red arrow indicates the actual percentage connectivity of locus heterogeneity genes (79.7%), showing a significant difference from 10,000 random simulations.
Using this same method to examine the connectivity of proteins associated with single-gene Mendelian disorders produced a significant result, although in this case the initial connectivity percentage was 64%. This interconnectivity between Mendelian genes may be due to the large number of Mendelian disease genes in our dataset affecting the same physiological systems (Figure 2). However, proteins associated with locus heterogeneity are more connected than proteins associated with Mendelian disease (79.7% compared with 64%), despite both sets of disorders showing an equal distribution of physiological pathologies. This further emphasizes the greater interconnectivity of disease-associated locus heterogeneity genes compared to disease-associated Mendelian genes.
CLUSTERING ANALYSIS OF THE HUMAN PROTEOME REVEALS HIGHLY INTERCONNECTED MODULES OF LOCUS HETEROGENEITY GENES
Clustering analysis was performed on protein interaction networks in an attempt to find protein complexes and functional clusters, which can be identified as highly interconnected subgraphs. Topological modules signify areas of dense local connectivity within a network, and with the use of experimental data, can be validated as functional modules of proteins defining an aggregation of proteins with similar or related biological function (Vidal et al., 2011). Here, a pre-existing algorithmic approach was used to identify densely interconnected groups within the locus heterogeneity disease network, and through the application of IPA and GO, we were able to confirm the functional relatedness of these genes.
As suggested by the average clustering coefficient of the locus heterogeneity disease network, we found that locus heterogeneity genes responsible for the same disease tended to be highly interconnected, and were present in the same topological modules. This result provides additional evidence for the highly interconnected nature of locus heterogeneity proteins. We further predict that a number of proteins within these modules positioned in close proximity to a group of locus heterogeneity proteins may be involved in the pathology of similar disorders, or may in fact be an undiscovered cause for locus heterogeneity disorders. The following examples [Bardet–Biedl syndrome, Leigh syndrome (LS), and Kabuki syndrome (KS)] demonstrate how genes within the local modular neighborhood of a locus heterogeneity disease gene may be possible disease gene candidates.
These functional modules displayed an MCODE complex score higher than 3, which indicates a greater accuracy and reliability of predictions. To determine if the clustering algorithm altered the modules produced from the network, modules were validated using alternative clustering algorithms. The Louvain method (
Bardet–Biedl syndrome
Bardet–Biedl syndrome (BBS) is a genetically and clinically heterogeneous disorder of developmental origin caused by mutations in a number of loci, with primary features including retinal dystrophy, hypogenitalism, renal malformations, and obesity (
Clustering analysis of our network using the MCODE algorithm (
FIGURE 5

Interconnectivity of Bardet–Biedl syndrome genes. Circular nodes represent proteins, with the lines between them signifying an interaction between the two proteins.
The most highly connected of these ‘healthy’ proteins is CCDC28B. Literature searches confirmed that this gene-product has known involvement in an alternative form of BBS. BBS is usually inherited in a monogenic autosomal recessive manner; in rare cases three mutations across two loci modify the onset and severity of the phenotype. Along with genes already annotated within our dataset, studies have shown that CCDC28B is one of these modifier genes (
The proteins in our network currently lacking in BBS annotations preferentially connect with BBS1, BBS 2, BBS4 and BBS7, with the exception of PCM1, which also interacts with BBIP1. According to GO analysis, a number of these proteins are involved in cilium assembly (p = 7.37e-18) and epithelial neoplasia (p = 1.21e-22), similar to known BBS causing genes. The molecular chaperone HscB only has three characterized protein interactions, all of which are with BBS causing proteins. The HscB protein displays similar cellular localization and interactions with BBS proteins, and previous studies have shown the HscB mutations have the ability to cause protein folding malformations (Vickery and Cupp-Vickery, 2007). Therefore, these data suggest that HscB may be a potential BBS candidate. Another protein with no current disease annotations is RAB3IP. A number of studies have shown that core BBS proteins form a complex that cooperate with GTPases, including RAB3IP, to promote ciliary membrane biogenesis (
Leigh syndrome
Leigh syndrome is characterized by severe neurodegeneration arising typically within the first year of life, manifesting clinically through rapid deterioration of cognitive and motor functions due to lesions in the basal ganglia and brain stem of affected patients, with clinical and genetic heterogeneity (
The module shown in Figure 6 involves four LS affected proteins surrounded by a number of proteins without disease annotations. Compared to the previous example, these non-disease proteins show a more varied connection to locus heterogeneity proteins. Two proteins, NDUFA6 and NDUFB10, both connect to three LS genes and, according to IPA, belong to the identical canonical pathways as these three LS affected proteins (mitochondrial dysfunction and oxidative phosphorylation). Further inspection using GO analysis confirmed that the two unmarked proteins are involved in the same biological processes as our LS causing genes, for example the respiratory electron transport chain (p = 6.94e-15). Previous studies analyzing these mitochondrial enzymes have suggested that they have an involvement in neurodegeneration, and that their perturbation may play a role in neurodegenerative disorders (
FIGURE 6

Leigh syndrome gene clustering. Each circular node denotes a protein and a line illustrates an interaction between two proteins.
In contrast, other surrounding genes only connect to one LS protein and show less connectivity to the disorder, therefore making them less likely to be disease candidates. Although these proteins localize to the mitochondria, they are not found in the same canonical pathways (mitochondrial dysfunction and oxidative phosphorylation) as known LS genes. This result suggests that genes with a higher degree of connectivity to multiple heterogeneous genes increases the likelihood of that gene’s involvement in the same biological processes, and therefore increases a gene’s potential to be a disease candidate.
Kabuki syndrome
Whilst the two disorders discussed previously show severe locus heterogeneity, KS is only known to occur through mutations in the histone methyltransferases KMT2D (also known as MLL2) or KDM6A, causing a breakdown in the epigenetic control of active chromatin states (
Clustering analysis revealed a module whereby five proteins interconnect with one another, including the two KS associated proteins (Figure 7). Additionally, according to GO analysis all proteins within this submodule localize within the nucleus, specifically within histone methyltransferase complexes (p = 3.56e-7), and have identical biological processes in chromatin modification (p = 2.19e-7). Perhaps the most interesting of these connected genes is KMT2C (MLL3), a lysine-specific methyltransferase that acts in a similar manner to KMT2D, and has recently shown strong associations to other neoplasmic disorders (
FIGURE 7

Heterogeneity of genes causing Kabuki syndrome. Lines signify interactions between two proteins, represented by circular nodes.
Current knowledge of the roles of the genes KMT2C, KMT2D, and PAXI1, along with their interconnectivity with the two known KS proteins and evidence in the literature that KS may be caused by mutations in additional genes, suggests that these genes should be the target of genetic screening in patients where KMT2D and KDM6A mutations have not been detected.
DISCUSSION
Our study demonstrates that disease genes expressing locus heterogeneity display properties that allow them to be distinguished from disease genes causing simple Mendelian disorders, such as sickle cell anemia, and disease genes as a whole. Analysis of the human proteome revealed that proteins encoded by locus heterogeneity genes are highly interconnected with those involved in the same disorder, grouping together in the clustering analysis of the network (Figures 5–7). In agreement with a study by
The techniques employed here have been used in recent studies concerning ASD and ID, a group of disorders that display considerable locus heterogeneity (
The use of protein interaction networks in this study allowed for large-scale comparisons of 1000s of protein interactions curated from a number of experimental sources. Despite the ability to easily identify relationships between genes, and the extent to which proteins interconnect, these networks, and the methods used to analyze them, have important limitations which must be considered. Firstly, even though interactions within the network have been experimentally verified from a number of sources, protein interactions are often difficult to assay on a proteomic scale, leading to false negative and false positive results. As well as an inability to distinguish between transient and obligate interactions within the network, data concerning the spatial and temporal nature of interactions is often limited or ignored for network reconstructions such as this. Finally, the importance of particular interactions can vary between nodes, even within clusters, meaning that experimental validation of candidate predictions is vital (
Although the complete landscape of heterogeneous disease is larger and more diverse than explored here, our results imply that locus heterogeneity genes show distinct properties allowing the identification of novel disease genes in the local network neighborhood, providing a pathway for further experimental study and candidate gene identification. Our finding that proteins encoded by locus heterogeneity disease genes are more highly interconnected than other types of disease genes indicates that clustering analysis will have particular value in identifying additional as yet unknown causative genes for diseases displaying locus heterogeneity. Increasing our understanding of specific gene classifications is essential to improve our knowledge of human disorders. As shown here, focusing on specific subsets of disease genes allows us to provide novel insights on a systems level to direct future research. As proteomic research continues, delivering a greater depth and reliability to human protein interaction data, we believe that studies such as this will become essential in providing novel advances to aid the identification of disease genes.
Statements
Acknowledgments
We thank Ryan Ames for helpful suggestions regarding graph analysis using Perl. We thank Jean-Marc Schwartz and Ruth Stoney for their help in implementing the Louvain community clustering method. This research was supported by BBSRC grant BB/L018276/1 to Kathryn E. Hentges.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
REFERENCES
1
AllegroViva Inc. (2014a). AllegroLayout Plugin.Available at: http://allegroviva.com/allegrolayout2/ [accessed April 6, 2014].
2
AllegroViva Inc. (2014b). AllegroMCODE Plugin.Available at: http://allegroviva.com/allegromcode/ [accessed April 6, 2014].
3
AshburnerM.BallC. A.BlakeJ. A.BotsteinD.ButlerH.CherryJ. M.et al (2000). Gene ontology: tool for the unification of biology.Nat. Genet. 25. 25–29. 10.1038/75556
4
BadanoJ. L.AnsleyS. J.LeitchC. C.LewisR. A.LupskiJ. R.KatsanisN. (2003). Identification of a novel Bardet-Biedl syndrome protein, BBS7, that shares structural features with BBS1 and BBS2.Am. J. Hum. Genet.72650–658. 10.1086/368204
5
BadanoJ. L.LeitchC. C.AnsleyS. J.May-SimeraH.LawsonS.LewisR. A.et al (2006). Dissection of epistasis in oligogenic Bardet-Biedl syndrome.Nature439326–330. 10.1038/nature04370
6
BaderG. D.HogueC. W. (2003). An automated method for finding molecular complexes in large protein interaction networks.BMC Bioinformatics4:2. 10.1186/1471-2105-4-2
7
BaertlingF.RodenburgR. J.SchaperJ.SmeitinkJ. A.KoopmanW. J. H.MayatepekE.et al (2014). A guide to diagnosis and treatment of Leigh syndrome.J. Neurol. Neurosurg. Psychiatry85257–265. 10.1136/jnnp-2012-304426
8
BakerK.BealesP. L. (2009). Making sense of cilia in disease: the human Cilloplathies.Am. J. Med. Genet. C Semin. Med. Genet. 151C, 281–295. 10.1002/ajmg.c.30231
9
BarabasiA.-L.GulbahceN.LoscalzoJ. (2011). Network medicine: a network-based approach to human disease.Nat. Rev. Genet.1256–68. 10.1038/nrg2918
10
BarabasiA. L.OltvaiZ. N. (2004). Network biology: understanding the cell’s functional organization.Nat. Rev. Genet.5101–113. 10.1038/nrg1272
11
Bauer-MehrenA.BundschusM.RautschkaM.MayerM. A.SanzF.FurlongL. I. (2011). Gene-disease network analysis reveals functional modules in Mendelian, complex and environmental diseases.PLoS ONE6:e20284. 10.1371/journal.pone.0020284
12
BealesP. L.BadanoJ. L.RossA. J.AnsleyS. J.HoskinsB. E.KirstenB.et al (2003). Genetic interaction of BBS1 mutations with alleles at other BBS loci can result in non-Mendelian Bardet-Biedl syndrome.Am. J. Hum. Genet.721187–1199. 10.1086/375178
13
BlondelV. D.GuillaumeJ.-L.LambiotteR.LefebvreE. (2008). Fast unfolding of communities in large networks.J. Stat. Mech.2008:P10008. 10.1088/1742-5468/2008/10/p10008
14
BokinniY. (2012). Kabuki syndrome revisited.J. Hum. Genet.57223–227. 10.1038/jhg.2012.28
15
CallenE.FaryabiR. B.LuckeyM.HaoB.DanielJ. A.YangW.et al (2012). The DNA damage- and transcription-associated protein paxip1 controls thymocyte development and emigration.Immunity37971–985. 10.1016/j.immuni.2012.10.007
16
ChialH. (2008). Rare genetic disorders: learning about genetic disease through gene mapping, SNPs, and microarray data.Nat. Educ.1192.
17
DaigerS. P.RossiterB. F.GreenbergJ.ChristoffelsA.HideW. (1998). Data services and software for identifying genes and mutations causing retinal degeneration.Annu. Meet. Assoc. Res. Vis. Ophthalmol.39:S295.
18
DickersonJ. E.RobertsonD. L. (2012). On the origins of Mendelian disease genes in man: the impact of gene duplication.Mol. Biol. Evol.2961–69. 10.1093/molbev/msr111
19
DickersonJ. E.ZhuA.RobertsonD. L.HentgesK. E. (2011). Defining the role of essential genes in human disease.PLoS ONE6:e27368. 10.1371/journal.pone.0027368
20
DongJ.HorvathS. (2007). Understanding network concepts in modules.BMC Syst. Biol.1:24. 10.1186/1752-0509-1-24
21
FinstererJ. (2008). Leigh and Leigh-like syndrome in children and adults.Pediatr. Neurol.39223–235. 10.1016/j.pediatrneurol.2008.07.013
22
FomousC.MitchellJ. A.MccrayA. (2006). ‘Genetics home reference’: helping patients understand the role of genetics in health and disease.Community Genet.9274–278. 10.1159/000094477
23
FurlongL. I. (2013). Human diseases through the lens of network biology.Trends Genet.29150–159. 10.1016/j.tig.2012.11.004
24
GerraldsM. (2014). Leigh syndrome: the genetic heterogeneity story continues.Brain1372872–2873. 10.1093/brain/awu264
25
GohK. I.CusickM. E.ValleD.ChildsB.VidalM.BarabasiA. L. (2007). The human disease network.Proc. Natl. Acad. Sci. U.S.A.1048685–8690. 10.1073/pnas.0701361104
26
GuoY.WeiX.DasJ.GrimsonA.LipkinS. M.ClarkA. G.et al (2013). Dissecting disease inheritance modes in a three-dimensional protein network challenges the “guilt-by-association” principle.Am. J. Hum. Genet.9378–89. 10.1016/j.ajhg.2013.05.022
27
HamoshA.ScottA. F.AmbergerJ. S.BocchiniC. A.MckusickV. A. (2005). Online Mendelian Inheritance in Man (OMIM), a knowledgebase of human genes and genetic disorders.Nucleic Acids Res.33D514–D517. 10.1093/nar/gki033
28
HannibalM. C.BuckinghamK. J.NgS. B.MingJ. E.BeckA. E.McmillinM. J.et al (2011). Spectrum of MLL2 (ALR) mutations in 110 cases of Kabuki syndrome.Am. J. Med. Genet. A 155A, 1511–1516. 10.1002/ajmg.a.34074
29
HarrisS. E.FoxH.WrightA. F.HaywardC.StarrJ. M.WhalleyL. J.et al (2007). A genetic association analysis of cognitive ability and cognitive ageing using 325 markers for 109 genes associated with oxidative stress or cognition.BMC Genet.8:43. 10.1186/1471-2156-8-43
30
HartongD. T.BersonE. L.DryjaT. P. (2006). Retinitis pigmentosa.Lancet3681795–1809. 10.1016/s0140-6736(06)69740-7
31
HietaniemiJ. (2014). Graph-0.96.Available at: http://search.cpan.org/~jhi/Graph-0.96/lib/Graph.pod [accessed April 14, 2014].
32
HirschhornJ. N.DalyM. J. (2005). Genome-wide association studies for common diseases and complex traits.Nat. Rev. Genet.695–108. 10.1038/nrg1521
33
HuangD. W.ShermanB. T.LempickiR. A. (2009a). Bioinformatics enrichment tools: paths toward the comprehensive functional analysis of large gene lists.Nucleic Acids Res.371–13. 10.1093/nar/gkn923
34
HuangD. W.ShermanB. T.LempickiR. A. (2009b). Systematic and integrative analysis of large gene lists using DAVID bioinformatics resources.Nat. Protoc.444–57. 10.1038/nprot.2008.211
35
Ingenuity® Systems. (2014). Ingenuity Pathway Analysis.Available at: http://www.ingenuity.com/ [accessed April 17, 2014].
36
IossifovI.RonemusM.LevyD.WangZ.HakkerI.RosenbaumJ.et al (2012). De novo gene disruptions in children on the autistic spectrum.Neuron74285–299. 10.1016/j.neuron.2012.04.009
37
JiangH.ShuklaA.WangX.ChenW.-Y.BernsteinB. E.RoederR. G. (2011). Role for Dpy-30 in ES cell-fate specification by regulation of H3K4 methylation within bivalent domains.Cell144513–525. 10.1016/j.cell.2011.01.020
38
JonesS.ZhangX.ParsonsD. W.LinJ. C.-H.LearyR. J.AngenendtP.et al (2008). Core signaling pathways in human pancreatic cancers revealed by global genomic analyses.Science3211801–1806. 10.1126/science.1164368
39
KaltenbachL. S.RomeroE.BecklinR. R.ChettierR.BellR.PhansalkarA.et al (2007). Huntingtin interacting proteins are genetic modifiers of neurodegeneration.PLoS Genet.3:e82. 10.1371/journal.pgen.0030082
40
KamburovA.StelzlU.LehrachH.HerwigR. (2013). The consensusPathDB interaction database: 2013 update.Nucleic Acids Res.41D793–D800. 10.1093/nar/gks1055
41
KrummN.O’RoakB. J.ShendureJ.EichlerE. E. (2014). A de novo convergence of autism genetics and molecular neuroscience.Trends Neurosci.3795–105. 10.1016/j.tins.2013.11.005
42
LiB.LiuH.-Y.GuoS.-H.SunP.GongF.-M.JiaB.-Q. (2013a). Mll3 genetic variants affect risk of gastric cancer in the chinese han population.Asian Pac. J. Cancer Prev.144239–4242. 10.7314/apjcp.2013.14.7.4239
43
LiW.-D.LiQ.-R.XuS.-N.WeiF.-J.YeZ.-J.ChengJ.-K.et al (2013b). Exome sequencing identifies an MLL3 gene germ line mutation in a pedigree of colorectal cancer and acute myeloid leukemia.Blood1211478–1479. 10.1182/blood-2012-12-470559
44
LiM.WangJ.ChenJ.PanY. (2009). Hierarchical organization of functional modules in weighted protein interaction networks using clustering coefficient.Bioinformatics Res. Appl.554275–86. 10.1007/978-3-642-01551-9_8
45
LiuJ. L.BaynamG. (2010). “Cornelia de Lange Syndrome,” inDiseases of DNA Repaired.AhmadS. I. (Berlin: Springer-Verlag) 111–123.
46
LoweH. J.BarnettG. O. (1994). Understanding and using the medical subject-headings (Mesh) vocabulary to perform literature searches.JAMA2711103–1108. 10.1001/jama.271.14.1103
47
MaereS.HeymansK.KuiperM. (2005). BiNGO: a cytoscape plugin to assess overrepresentation of gene ontology categories in biological networks.Bioinformatics213448–3449. 10.1093/bioinformatics/bti551
48
McClellanJ.KingM.-C. (2010). Genetic heterogeneity in human disease.Cell141210–217. 10.1016/j.cell.2010.03.032
49
McKusickV. A. (1991). Current trends in mapping human genes.FASEB J.512–20.
50
MiyakeN.KoshimizuE.OkamotoN.MizunoS.OgataT.NagaiT.et al (2013). MLL2 and KDM6A mutations in patients with Kabuki syndrome.Am. J. Med. Genet. A1612234–2243. 10.1002/ajmg.a.36072
51
NachuryM. V.LoktevA. V.ZhangQ.WestlakeC. J.PeranenJ.MerdesA.et al (2007). A core complex of BBS proteins cooperates with the GTPase Rab8 to promote ciliary membrane biogenesis.Cell1291201–1213. 10.1016/j.cell.2007.03.053
52
NealeB. M.KouY.LiuL.Ma’ayanA.SamochaK. E.SaboA.et al (2012). Patterns and rates of exonic de novo mutations in autism spectrum disorders.Nature485242–245. 10.1038/nature11011
53
O’RoakB. J.VivesL.GirirajanS.KarakocE.KrummN.CoeB. P.et al (2012). Sporadic autism exomes reveal a highly interconnected protein network of de novo mutations.Nature485246–250. 10.1038/nature10989
54
O’SullivanB. P.FreedmanS. D. (2009). Cystic fibrosis.Lancet3731891–1904. 10.1016/S0140-6736(09)60327-5
55
PapatriantafyllouM. (2013). Lymphocyte development PAXIP1-a gatekeeper of thymocyte development.Nat. Rev. Immunol.132–3. 10.1038/nri3367
56
PeriS.NavarroJ. D.KristiansenT. Z.AmanchyR.SurendranathV.MuthusamyB.et al (2004). Human protein reference database as a discovery resource for proteomics.Nucleic Acids Res.32D497–D501. 10.1093/nar/gkh070
57
RavaszE.SomeraA. L.MongruD. A.OltvaiZ. N.BarabasiA. L. (2002). Hierarchical organization of modularity in metabolic networks.Science2971551–1555. 10.1126/science.1073374
58
R-Development-Core-Team. (2009). R: A Language and Environment for Statistical Computing.Vienna: R Foundation for Statistical Computing.
59
SatohJ.-I.KawanaN.YamamotoY. (2013). Pathway analysis of ChIP-Seq-based NRF1 target genes suggests a logical hypothesis of their involvement in the pathogenesis of neurodegenerative diseases.Gene Regul. Syst. Biol.7139–152. 10.4137/grsb.s13204
60
ShannonP.MarkielA.OzierO.BaligaN. S.WangJ. T.RamageD.et al (2003). Cytoscape: a software environment for integrated models of biomolecular interaction networks.Genome Res.132498–2504. 10.1101/gr.1239303
61
ShenH.ChengX.CaiK.HuM.-B. (2009). Detect overlapping and hierarchical community structure in networks.Physica A Stat. Mech. Appl.3881706–1712. 10.1016/j.physa.2008.12.021
62
StarkC.BreitkreutzB.-J.RegulyT.BoucherL.BreitkreutzA.TyersM. (2006). BioGRID: a general repository for interaction datasets.Nucleic Acids Res.34D535–D539. 10.1093/nar/gkj109
63
Thanh-PhuongN.Tu-BaoH. (2012). Detecting disease genes based on semi-supervised learning and protein-protein interaction networks.Artif. Intell. Med.5463–71. 10.1016/j.artmed.2011.09.003
64
TobinJ. L.BealesP. L. (2007). Bardet-Biedl syndrome: beyond the cilium.Pediatr. Nephrol.22926–936. 10.1007/s00467-007-0435-0
65
VickeryL. E.Cupp-VickeryJ. R. (2007). Molecular chaperones HscA/Ssq1 and HscB/Jac1 and their roles in iron-sulfur protein maturation.Crit. Rev. Biochem. Mol. Biol.4295–111. 10.1080/10409230701322298
66
VidalM.CusickM. E.BarabasiA.-L. (2011). Interactome networks and human disease.Cell144986–998. 10.1016/j.cell.2011.02.016
67
WalshT.KingM.-C. (2007). Ten genes for inherited breast cancer.Cancer Cell11103–105. 10.1016/j.ccr.2007.01.010
68
WangJ.ZhongJ.ChenG.LiM.WuF. X.PanY. (2014). ClusterViz: a cytoscape APP for clustering analysis of biological networks.IEEE/ACM Trans. Comput. Biol. Bioinform.1:1. 10.1109/TCBB.2014.2361348
69
WangX.WeiX.ThijssenB.DasJ.LipkinS. M.YuH. (2012). Three-dimensional reconstruction of protein networks provides insight into human genetic disease.Nat. Biotechnol.30159–164. 10.1038/nbt.2106
70
WestlakeC. J.BayeL. M.NachuryM. V.WrightK. J.ErvinK. E.PhuL.et al (2011). Primary cilia membrane assembly is initiated by Rab11 and transport protein particle II (TRAPPII) complex-dependent trafficking of Rabin8 to the centrosome.Proc. Natl. Acad. Sci. U.S.A.1082759–2764. 10.1073/pnas.1018823108
71
ZaghloulN. A.KatsanisN. (2009). Mechanistic insights into Bardet-Biedl syndrome, a model ciliopathy.J. Clin. Invest.119428–437. 10.1172/jci37041
Summary
Keywords
locus heterogeneity, protein interaction network, systems biology, Bardet–Biedl syndrome, Leigh syndrome, Kabuki syndrome
Citation
Keith BP, Robertson DL and Hentges KE (2014) Locus heterogeneity disease genes encode proteins with high interconnectivity in the human protein interaction network. Front. Genet. 5:434. doi: 10.3389/fgene.2014.00434
Received
17 September 2014
Accepted
24 November 2014
Published
09 December 2014
Volume
5 - 2014
Edited by
Firas H. Kobeissy, University of Florida, USA
Reviewed by
Cheng Zhu, Genzyme, USA; Tarek H. Mouhieddine, American University of Beirut Medical Center, Lebanon
Copyright
© 2014 Keith, Robertson and Hentges.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Kathryn E. Hentges, Faculty of Life Sciences, University of Manchester, Michael Smith Building, Oxford Road, Manchester M13 9PT, UK e-mail: kathryn.hentges@manchester.ac.uk
This article was submitted to Systems Biology, a section of the journal Frontiers in Genetics.
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.