Abstract
Hypothetical proteins (HPs) are the proteins predicted to be expressed from an open reading frame, making a substantial fraction of proteomes in both prokaryotes and eukaryotes. Genome projects have led to the identification of many therapeutic targets, the putative function of the protein, and their interactions. In this review we enlist various methods linking annotation to structural and functional prediction of HPs that assist in the discovery of new structures and functions serving as markers and pharmacological targets for drug designing, discovery, and screening. Further we give an overview of how mass spectrometry as an analytical technique is used to validate protein characterisation. We discuss how microarrays and protein expression profiles help understanding the biological systems through a systems-wide study of proteins and their interactions with other proteins and non-proteinaceous molecules to control complex processes in cells. Finally, we articulate challenges on how next generation sequencing methods have accelerated multiple areas of genomics with special focus on uncharacterized proteins.
Introduction
Proteins are biological macromolecules translated from DNA to perform myriad functions. As a structural entity, they participate in the regulation of genes to perform function as enzymes or catalysts further playing a role in immune system or as a transporter. With the phenotype of an organism depending on the proteins expressed from the genotype, there are different types and classes of proteins based on their composition, configuration, property, and function. Added to this big diversified world of proteins, a new-fangled race called ‘Hypothetical proteins’ (HPs), involved to describe as functional candidates cannot be ignored. The HPs are proteins that are predicted to be expressed from an open reading frame (ORF), but for which there is no experimental evidence of translation. They constitute a substantial fraction of proteomes in both prokaryotes and eukaryotes with a majority of them included in humans and bacteria (). Many HPs show as ‘hypothetical’ when the genome is just sequenced; this is because of lack of annotation. Comparative genomics shows that a substantial fraction of the genes in sequenced genomes encodes ‘conserved hypothetical proteins’ (CHPs). CHPs are the proteins that are conserved among organisms from several phylogenetic lineages but for which there is no functional validation. Genome sequencing has flooded our information base with novel genes of unpredictable functions. Though genome projects have led to the identification of many therapeutic targets, the putative function of the protein, and their interactions could be predicted for only fewer than half of them. In the recent past, an effort has been made to define CHPs as a large fraction of genes in sequenced genomes encoding phylogenetic lineages but those that have no functional characterization for these ‘therapeutic’ targets ().
The Current Status
As on October 08, 2014, the GenBank labels about 48591211 HPs sequences in NCBI on which 7234262 are in eukaryotes and 34064553 are in bacteria. As on date, humans have an approximately 1040 HPs with conserved domains. Come next generation sequencing (NGS), there has been a huge interest in deciphering the function of these HPs not just limited to the sequences generated from traditional sequencing but just to check whether or not any new sequences are generated from NGS. HPs turn up during the genome analysis by bioinformatic tools in the process of identifying new genes. As these tools are pre-fed with instructions for finding the ORF in the genome, they return all possible sequences including those without any protein analog in the protein database or showing less identity to known, annotated protein. There are several in silico methods available for the functional predictions of HP, however, no single tool is sufficient enough to perform the annotation all by itself. With fallacy of using several predictors, we firmly reason that using different combination of prediction tools would help reach consensus and validate them to have a significant role which can further be proven by experimental analysis (; ; see Figure 1).
FIGURE 1
Annotation Linked to Structural and Functional Prediction
Annotation of HPs from a particular genome helps in the discovery of new structures and new function which further allows them to be classified into additional protein pathways and cascades. They also serve as markers and pharmacological targets for drug design, discovery, and screening (
Table 1
| List of bioinformatics tools and databases used for sequence based function annotation | ||
| S.no | Software | Function |
| A | Sequence similarity search | |
| 1 | Basic local alignment tool (BLAST) | Used for finding similar sequences in protein databases |
| B | Physiochemical characterization | |
| 2 | ExPASy -- Protparam tool | Used for computation of various physical and chemical parameters like molecular weight, isoelectric point (Pi), amino acid composition, atomic composition, extinction co-efficient, instability index, aliphatic index, and grand average of hydropathy (GRAVY) |
| C | Sub-cellular localization | |
| 3 | signalP | Predicts signal peptide cleavage sites. |
| 4 | secretomeP | Used for identifying proteins involved in non-classical secretory pathway. |
| 5 | PSORT B | Predicts subcellular localization of bacterial proteins. |
| 6 | PSLpred | Predicts subcellular localization of proteins from Gram-negative bacteria. |
| 7 | CELLO | Assign localization to both prokaryotic and eukaryotic proteins |
| 8 | TMHMM | used to authenticate whether the protein is a membrane protein or not. |
| 9 | HMMTOP | Predict transmembrane topology. |
| D | Domain analysis and protein | |
| 10 | Pfam | Collection of multiple protein sequence alignments |
| 11 | SVMprot | SVM (Support vector machine based classification of proteins |
| 12 | SYSTERS | For grouping of proteins on the basis of their functions. |
| 13 | SUPERFAMILY | Hierarchical domain classification of PDB structures. NCBI Entrez protein database search of domain architecture |
| 14 | CATH (Class, Architecture, Topology, Homology) | Used for finding protein similarities across evolutionary distances based on domain architecture. Classification based on HMM--HMM search. PANTHER is a |
| I5 | CDART (The conserved domain architecture | comprehensively organized database of protein families and |
| retrieval tool) | sub-families, their evolutionary relationships in the form of | |
| phylogenetic trees | ||
| 16 | PANTHER (Protein analysis through evolutionary relationships) | Identification and annotation of protein domains. |
| 17 | SMART | Automatic hierarchical clustering of the protein sequences |
| 18 | ProtoNet | |
| E | Motif Analysis | |
| 19 | InterProScan | Searches interPro for motif discovery. It is the integration of |
| several large protein signature databases. | ||
| 20 | MOTIF | used for Motif discovery. |
| 21 | MEME suite | Database searching for assigning function to the discovered motifs. |
| F | Protein--Protein interaction | |
| 22 | STRING | Used for predicting protein--protein interactions. |
| List of some wet lab experiments for protein characterization | ||
| Method | Function | |
| A | Chromatographic separations | |
| 1 | Gel filtration chromatography | Separates proteins based on their size (which is closely related to their molecular weight) |
| 2 | Ion- exchange chromatography | Purify proteins according to their overall charge |
| 3 | Affinity chromatography | Separates proteins based on their affinity to bind to a known ligand. |
| B | Electrophoresis | |
| 4 | SDS-PAGE | Separates protein according to molecular weight and allows the measurement of the molecular weight in comparison with marker proteins. |
| 5 | Isoelectric focusing | Separates proteins based on their PI on a polyacryl-amide gel with a PH gradient. |
| 6 | 2D-Electrophoresis | Isoelectric focussing is often used in conjunction with SDS-PAGE to give a very powerful method of protein characterization by separating the sample of protein first by isoelectric point and then by molecular weight. |
| C | Spectroscopic analysis | |
| 7 | NMR spectroscopy | For determining three dimensional structure of proteins |
| 8 | Mass spectrometry | For protein identification and characterization. |
| D | Others | |
| 9 | Yeast two hybrid assay | For studying protein--protein interactions. |
| 10 | Phage display method | For studying protein--protein interactions |
| 11 | Microarray analysis | For systems-oriented study of proteins |
| 12 | Next generation sequencing | For high-throughput sequencing of genome and proteome analysis. |
Methods used for protein characterization and annotation.
Wet- Lab Experiments are Used to Confirm the Candidate Hypothetical Protein
Although gene prediction programs using various bioinformatics tools have become more accurate and sensitive, analysis of HPs, there is a want of more reliable evidence for existence and function of predicted proteins (
While biochemical characterization of proteins provide insight of gene function, physiochemical properties of the protein such as molecular weight, stability, proper folding, etc. have to be determined. Conventional technologies of protein separation and characterization such as chromatographic separation, protein and DNA electrophoresis, cell sorting, affinity assays (e.g., immunoassays), spectroscopic analysis have been miniatured by microfluidic technologies. Technologies such as microfluidics and other lab-on-a-chip methods rely on assays that are rapid and inexpensive (
Mass spectrometry is a powerful analytical technique for validating protein coding genes. It analyses and quantifies 1000s of proteins from complex samples and thus permits the characterisation of putative gene products at the level of translation (
Microarrays and Protein Expression Profiles
Current technologies limit our analysis to only one or two of the parameters to be studied and to only fraction of proteins. Systems-oriented proteomics provide us with integrated understanding of biological systems by studying many components simultaneously. Furthermore it helps us to understand how proteins interact with other proteins and non-proteinaceous molecules to control complex processes in cells and tissues and even whole organism. In systems-oriented proteomics the subset of proteins to be analyzed is well defined such that sequences or collection of proteins are related by function. Microarray technology is well suited to systems-oriented studies. Two features that make microarray technology so well suited to systems-oriented proteomics are 1000s of proteins can be interrogated simultaneously by spotting them on a single slide or similar support and similar proteins can be probed repeatedly with many different molecules under many different conditions by fabricating 100s–1000s of copies of an array in parallel (
Next generation sequencing technologies are way ahead of microarrays and fundamentally altered the genomics research. Experiments that were not technically feasible or affordable previously are now made possible with the advent of NGS technology thus accelerating multiple areas of genomics research. Thanks to many NGS platforms that are available sharing a common technological feature of massively parallel sequencing of clonally amplified or single DNA molecules that are spatially separated in a flow cell (
Conclusion
A comprehensive identification of the HPs is needed for the functional interpretation of fully sequenced genomes and further understanding of the diverse functions of its unique structures, which in turn facilitates search for potential proteins of interest for researchers. Development of computational approaches and programs on elucidation of the functions of CHPs create an opportunity for biologists to produce a complete record of their biological functions and the genes involved. Protein science on the other hand have taken a new look with advances in the chemical synthesis of peptides and site-directed mutagenesis as standard research tools. This creates way for the construction of new proteins with customized structural and functional properties. However, the most important step in this process understands the complex folding patterns of these synthetic polypeptides to form a functional protein. We have tabulated and discussed several in silico methods available for the functional predictions of HP from sequence to structural levels like homology search, identification of domains and motifs, comparative analysis, phylogenetic profiling, and so on. Interpreting the physiological function of the HPs could establish greater interest in understanding evolutionary relationship of genes and organisms and would as well assist in drug discovery. We believe with the increase in the amount of sequence data with respect to HPs, there is a pressing need to organize this data and network their function to the existing known sequences. This process would allow us to identify HPs localized to different organelles involved with crucial prime functions, linked to various diseases. A permutation and combination of bioinformatics methods followed by wet lab experiments as listed in the figure above would be very useful for rapid functional annotation of novel proteins and will be useful for design of novel peptides and will have immediate impact on drug design research. Though few databases exist for analyzing HPs, a large public repository exclusive for HPs for ready reference to biologists and researchers around the world, would bring a greater impact and solution to many on-going projects.
Statements
Acknowledgments
We thank Dr. Prashanth Suravajhala, Bioclues.org for his valuable comments and suggestions on this manuscript. We would also like to thank Bioclues Organization and CFRD (Osmania University) for their support and guidance throughout this review preparation.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
References
1
AxtellM. J.SnyderJ. A.BarberD. P. (2007). Common functions for diverse small RNAs of land plants.Plant cell191750–1769. 10.1105/tpc.107.051706
2
BensoA.Di CarloS.ur RehmanH.PolitanoG.SavinoA.SuravajhalaP. (2013). A combined approach for genome wide protein function annotation/prediction.Proteome Science 11(Suppl.1) S1.10.1186/1477-5956-11-S1-S1
3
BrenneckeJ.AravinA. A.StarkA.DusM.KellisM.SachidanandamR.et al (2007). Discrete small RNA-generating loci as master regulators of transposon activity in drosophila.Cell1281089–1103. 10.1016/j.cell.2007.01.043
4
DeslerC.DurhuusJ. A.RasmussenL. J. (2012). Genome-wide screens for expressed hypothetical proteins.Methods Mol. Biol.81525–38. 10.1007/978-1-61779-424-7_3
5
EspinosaC. E. S.SlackF. J. (2006). The role of MicroRNAs in Cancer.Yale J. Biol. Med.79131–140.
6
FountoulakisM.LangenH. (1997). Identification of proteins by matrix assisted laser desorption ionization mass spectrometry following in-gel digestion in low salt, non-volatile buffer and simplified peptide recovery.Anal. Biochem.250153–156. 10.1006/abio.1997.2213
7
GalperinM. Y.KooninE. V. (2004). Conserved hypothetical’ proteins: prioritization of targets for experimental study.Nucleic Acids Res.325452–5463. 10.1093/nar/gkh885
8
GorgA.WeissW.DunnM. J. (2004). Current two-dimensional electrophoresis technology for proteomics.Proteomics43665–3685. 10.1002/pmic.200401031
9
HenzelW. J.BilleciT. M.StultsJ. T.WongS. C.GrimleyC.WatanabeC. (1993). Identifying proteins from two-dimensional gels by molecular mass searching of peptide fragments in protein sequence databases.Proc. Natl. Acad. Sci. U.S.A.905011–5015. 10.1073/pnas.90.11.5011
10
HouwingS.KammingaL. M.BerezikovE.CronemboldD.GirardA.van den ElstH.et al (2007). A role of piwi and pi RNAs in germ cell maintenance and transposon silencing in zebra fish.Cell12969–82. 10.1016/j.cell.2007.03.026
11
HurdP. J.NelsonC. J. (2009). Advantages of next-generation sequencing versus the microarray in epigenetic research.Brief. Funct. Genomic Proteomic.8174–183. 10.1093/bfgp/elp013
12
KodadekT. (2001). Protein microarrays-prospects and problems.Chem. Biol.8105–115. 10.1016/S1074-5521(00)90067-X
13
LubecG.Afjehi-SadatL.YangJ.-W.JohnJ. P. P. (2005). Searching for hypothetical proteins: theory and practice based upon original data and literature.Progr. Neurobiol.7790–127. 10.1016/j.pneurobio.2005.10.001
14
MacBeathG. (2002). Protein microarrays and proteomics.Nat. Genet. 32(Suppl.), 526–532. 10.1038/ng1037
15
MardisE. R. (2008). Next-generation DNA sequencing methods.Annu. Rev. Genomics Hum. Genet.9387–402. 10.1146/annurev.genom.9.081307.164359
16
MarvinL. F.RubertsM. A.FayL. B. (2003). Matrix-assisted laser desorption/ionisation time –of-flight mass spectrometry in clinical chemistry.Clin. Chim. Acta33711–21. 10.1016/j.cccn.2003.08.008
17
MeierM. 1.SitR. V.QuakeS. R. (2013). Proteome-wide protein interaction measurements of bacterial 17.proteins of unknown function.Proc. Natl. Acad. Sci. U.S.A.110477–482. 10.1073/pnas.1210634110
18
MelinJ.QuakeS. R. (2007). Microfluidic large-scale integration: the evolution of design rules for biological automation.Annu. Rev. Biophys. Biomol. Struct.36213–231. 10.1146/annurev.biophys.36.040306.132646
19
MohanR.VenugopalS. (2012). Computational structures and functional analysis of hypothetical proteins of Staphylococcus aureus.Bioinformation8722–728. 10.6026/97320630008722
20
MolloyM. P.WitzmannF. A. (2002). Proteomics: technologies and applications.Brief. Funct. Genomic Proteomic123–29. 10.1093/bfgp/1.1.23
21
ShahbaazM.HassanM. I.AhmadF. (2013). Functional annotation of conserved hypothetical proteins from Haemophilus influenzae Rd KW20.PLoS ONE8:e84263. 10.1371/journal.pone.0084263
22
ShinJ.-H.YangJ.-W.JuranvilleJ.-F.FountoulakisM.LubecG. (2004). Evidence for existence of thirty hypothetical proteins in rat brain.Proteome Sci.21. 10.1186/1477-5956-2-1
23
SivashankariS.ShanmughavelP. (2006). Functional annotation of hypothetical proteins – A review.Bioinformation1335–338. 10.6026/97320630001335
24
SuravajhalaP.BizzaroJ. W. (2015). A conceptual outline for ‘omics experiments using bioinformatics analogies. BioProtocol5a1963.
25
SuravajhalaP.SundararajanV. S. (2012). A classification scoring schema to validate protein interactors.Bioinformation.834–39. 10.6026/97320630008034
26
TannerS.ShenZ.NgJ.FloreaL.GuigóR.BriggsS. P.et al (2007). Improving gene annotation using peptide mass spectrometry.Genome Res.17231–239. 10.1101/gr.5646507
27
ThiedeB.HohenwarterW.KrahA.MattowJ.SchmidM.SchmidtF.et al (2005). Peptide mass fingerprinting.Methods35237–247. 10.1016/j.ymeth.2004.08.015
28
VoelkerdingK. V.DamesS. A.DurtschiJ. D. (2009). Next generation sequencing: from basic research to diagnostics.Clin. Chem.55641–658. 10.1373/clinchem.2008.112789
29
WhitesidesG. M. (2006). The origins and the future of microfluidics.Nature442368–373. 10.1038/nature05058
30
YouS. A.WangQ. K. (2007). Proteomics with two-dimensional gel electrophoresis and mass spectrometry analysis in cardiovascular research.Methods Mol. Med.12915–26.
31
ZhaoT.LiG.MiS.LiS.HannonG. J.WangX. J.et al (2007). A complex system of small RNAs in the unicellular green algae Chlamydomonas reinhardtii.Genes Dev.151190–1203. 10.1101/gad.1543507
Summary
Keywords
hypothetical proteins, annotation, functional prediction, protein–protein interactions, drug design research, public repository
Citation
Ijaq J, Chandrasekharan M, Poddar R, Bethi N and Sundararajan VS (2015) Annotation and curation of uncharacterized proteins- challenges. Front. Genet. 6:119. doi: 10.3389/fgene.2015.00119
Received
24 September 2014
Accepted
11 March 2015
Published
31 March 2015
Volume
6 - 2015
Edited by
Alfredo Benso, Politecnico di Torino, Italy
Reviewed by
Athanasios Demiris, Micro2Gen Ltd., Greece Raul Isea; Fundaciòn Instituto de Estudios Avanzados, Venezuela
Copyright
© 2015 Ijaq, Chandrasekharan, Poddar, Bethi and Sundararajan.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Johny Ijaq and Vijayaraghava S. Sundararajan, Bioclues Organization, Kukatpally, Hyderabad 500072, Telangana, India johny@bioclues.org; sundar@bioclues.org
This article was submitted to Bioinformatics and Computational Biology, a section of the journal Frontiers in Genetics
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.