Abstract
The ubiquitous aquatic bacterium, Vibrio cholerae, is the causative agent of the human diarrheal disease cholera, which occurs through the consumption of contaminated food or water and can lead to potentially fatal outcomes. The V. cholerae genome contains two nonhomologous chromosomes that are ~2.96 Mbp and ~1.07 Mbp, respectively. Here we report a complete, updated comparative analysis of the functional annotations for the V. cholerae genome. Available genomic sequencing data from the original analysis in 2000 was compared to the updated genomic sequence in 2026 for the O1 El Tor str. N16961. The classification of genes as ‘hypothetical protein’ or ‘conserved hypothetical protein’ has dropped from over 40% in 2000 to ~5.7% as of 2026. Additional protein-protein blast comparisons performed for this study provide potential roles for ~38.5% of the remaining hypothetical proteins as of 2026. In combination, these changes highlight the progress made over the past quarter century. We present this analysis as an easily usable resource that allows the user to search for and compare the changes, if any, to the functional annotations of the over 3500 protein-coding genes present in the V. cholerae O1 El Tor str. N16961 genome.
Introduction
Vibrio cholerae is the etiological agent of the human diarrheal disease cholera. Expression of human virulence factors present on pathogenicity islands in the V. cholerae genome directly contributes to successful intestinal colonization and the voluminous efflux of watery diarrhea characteristic of the disease (; ). The O1 El Tor strain N16961 was the first complete V. cholerae genome to be sequenced in 2000 (; ). This genome was revised in 2019 and further refined in 2026 (Wellcome Sanger Institute, Hinxton, Cambridgeshire, UK) with considerable improvements of the functional annotations of genes and is publicly available on NCBI (). These revisions bring the genomic functional annotations of the V. cholerae O1 El Tor str. N16961 genome in line with the contemporary knowledge available, highlighted through the significant increase in annotated genes in 2026 in comparison to 2000.
Hypothetical and uncharacterized proteins
Proteins that are encoded by genes whose functional characterizations are not clear and evident at the time of analysis have been conventionally designated ‘unknown protein’, ‘hypothetical protein’, or ‘conserved hypothetical protein’. For this work, only the latter two designations are observed in the genomic annotation. As the functions of new proteins are uncovered, the reference genome does not automatically update the annotations for such genes. As a result, the genomic annotation, which was once considered groundbreaking, becomes obscure over time. The revision and refinement of the reference genome () significantly reduced the number of genes that were assigned to one of the aforementioned traditional designations. While this improvement is beneficial and should be performed over time to maintain the quality of the genome, we still lack a defined manner in which anyone can compare the changes to the gene annotations over time between the reference genomes for V. cholerae O1 El Tor str. N16961.
Methods
Two sources of data were used for this work [1]: the first complete V. cholerae genome sequence, published by Heidelberg et al. (2000) (), and [2] the latest genome sequence available for the same strain, published by the Wellcome Sanger Institute (2026) (). For ease of discussion, the two sequences will be referred to as Vc (2000) and Vc (2026), respectively, hereafter. For Vc (2000), all genes are designated a ‘VC’ number for genes on chromosome 1 and a ‘VC_A’ number for genes on chromosome 2. For Vc (2026), all genes are either designated by a four-letter gene name or by a ‘FY484_RS*****’ number where the asterisks represent a five-digit number; this classification applies to both chromosomes. By hovering over the gene, critical information such as the gene name, gene annotation, gene length in nucleotides, position on the chromosome, and any comments can be ascertained. We have meticulously evaluated Vc (2000) and Vc (2026) by going gene by gene to assess the changes or updates that may have occurred, utilizing the gene annotation, gene size, and position on the chromosome as comparative benchmarks between the two reference genomes. For genes with uncharacterized functions, we have utilized the standard protein blast search available through NCBI; the ClusteredNR database provides highly specific matches at a minimum of 90% sequence identity and length for the hypothetical proteins present in the genome. Here we present a resource that accurately and efficiently allows the user to search for any V. cholerae gene using the aforementioned designations, while additionally providing a direct comparison for every gene annotation between Vc (2000) and Vc (2026).
Results
The genome annotation for Vc (2000) contains 3890 genes (). Of these, 2307 genes are annotated with a gene name, 1577 genes are designated as ‘hypothetical protein’ or ‘conserved hypothetical protein,’ and 6 genes annotated in the 2026 genome are not listed. The genome annotation for Vc (2026) contains 3524 genes (). Of these, 3323 genes are annotated with a gene name, 201 genes are designated as ‘hypothetical protein’, and 366 genes annotated in the 2000 genome are not listed. A significant reduction in the number of proteins encoded by a gene labeled as ‘hypothetical’ or ‘conserved hypothetical’ is evident, dropping from over 40% of the total genes in 2000 to ~5.7% of the total genes in 2026. This change can be attributed to the discovery of numerous proteins whose functions were unknown at the time of the first annotation. To provide a current perspective on the reference genome, we conducted a protein-protein blast comparison search (BLASTP 2.17.0+) () to predict functions for the genes that were still designated as ‘hypothetical protein’ in Vc (2026) and to illustrate any possible changes that may have occurred. We present these results based on the matching species and highest percentage of sequence identity for each ‘hypothetical protein’ in Vc (2026). Of the 201 genes designated as ‘hypothetical protein’ in Vc (2026), we can confirm that BLASTP analysis provides significant sequence identity for 72 genes, facilitating annotation of their putative functions, a positive return of ~38.5%. By acknowledging the genes with predicted functions alongside the current annotated genome, we are able to report that over 96% of the V. cholerae genome has an attributed functional annotation in 2026. Additionally, a considerable number of genes described in the original sequence annotation () seemed to be simply present as ‘placeholders’, many of which share characteristics such as consisting of a fairly minute number of nucleotides (nt) (typically <100 nt). It is for these reasons that we highlight the overall reduction in the total number of genes with a functional annotation from 3890 to 3524 courtesy of the updated annotation along with the removal of those ‘placeholders’ which are not found in the updated sequence () (Table 1). The complete list of genes with the comparison of functional annotations between Vc (2000) and Vc (2026) and the BLASTP prediction analysis for the genes of unknown function in Vc (2026) are provided as Excel files in the supplementary data.
Table 1
| Year | Annotated | Hypothetical | Not listed |
|---|---|---|---|
| Vc (2000) | 2307 (59.4%) | 1577 (40.6%) | 6 |
| Vc (2026) | 3323 (94.3%) | 201 (5.7%) | 366 |
Quantification of annotation values by group of the Vibrio cholerae genome.
Percentages are representative of the total characterized genes for the respective year. “Not listed” indicates the number of genes that were counted in one genomic annotation but not the other, as described in the text.
Conclusion
Reference genomes provide a wealth of information for bacterial species and serve as a primary resource for genomic data. Advancements in genomic sequencing techniques have made this task significantly easier for any bacterial species or strain in question. The first complete V. cholerae genome sequence for the O1 El Tor str. N16961 in 2000 was an immense feat that has served and will continue to serve as a cornerstone for V. cholerae research. Through continued efforts, critical updates and discoveries have allowed the necessary changes to the genomic sequence and annotation to be made, providing a new resource that further allows for the characterization and discovery of new protein functions. We would like to thank the V. cholerae research community for their efforts in uncovering the unknown and establishing a well-annotated genome, while emphasizing the necessary importance of constant maintenance of the reference genome annotation.
Statements
Data availability statement
The original contributions presented in the study are included in the article/Supplementary Material. Further inquiries can be directed to the corresponding author.
Author contributions
RS: Formal Analysis, Writing – review & editing, Writing – original draft, Investigation, Visualization. JHW: Supervision, Resources, Writing – review & editing, Funding acquisition.
Funding
The author(s) declared that financial support was received for this work and/or its publication. This work is supported by the National Institute of Allergy and Infectious Diseases (NIAID) grant R21AI171072 (to JW).
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
The author JHW declared that they were an editorial board member of Frontiers, at the time of submission. This had no impact on the peer review process and the final decision.
Generative AI statement
The author(s) declared that generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Supplementary material
The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fbrio.2026.1833038/full#supplementary-material
References
1
AltschulS. F.MaddenT. L.SchäfferA. A.ZhangJ.ZhangZ.MillerW.et al. (1997). Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res.25, 3389–3402. doi: 10.1093/nar/25.17.3389. PMID:
2
HeidelbergJ. F.EisenJ. A.NelsonW. C.ClaytonR. A.GwinnM. L.DodsonR. J.et al. (2000). DNA sequence of both chromosomes of the cholera pathogen Vibrio cholerae. Nature406, 477–483. doi: 10.1038/35020000. PMID:
3
KaraolisD. K.JohnsonJ. A.BaileyC. C.BoedekerE. C.KaperJ. B.ReevesP. R. (1998). A Vibrio cholerae pathogenicity island associated with epidemic and pandemic strains. PNAS95, 3134–3139. doi: 10.1073/pnas.95.6.3134. PMID:
4
U.S. National Library of Medicine (2003). Vibrio cholerae O1 biovar El Tor Str. N16961 (ID 36) ( Washington, DC: National Center for Biotechnology Information). Available online at: https://www.ncbi.nlm.nih.gov/bioproject/PRJNA36.
5
U.S. National Library of Medicine (2019). Updated VC N16961 reference genome (ID 561735) ( Washington, DC: National Center for Biotechnology Information). Available online at: https://www.ncbi.nlm.nih.gov/bioproject/561735.
6
WaldorM. K.MekalanosJ. J. (1996). Lysogenic conversion by a filamentous phage encoding cholera toxin. Sci. (New York N.Y.)272, 1910–1914. doi: 10.1126/science.272.5270.1910. PMID:
Summary
Keywords
annotation, cholera, genome, updated, Vibrio cholerae
Citation
Singh R and Withey JH (2026) Updated functional annotation of the Vibrio cholerae O1 El Tor str. N16961 reference genome. Front. Bacteriol. 5:1833038. doi: 10.3389/fbrio.2026.1833038
Received
17 March 2026
Revised
23 April 2026
Accepted
29 April 2026
Published
19 May 2026
Volume
5 - 2026
Edited by
Ansel Hsiao, University of California, Riverside, United States
Reviewed by
Nitin Kamble, University of Cincinnati Medical Center, United States
Gustavo Espinoza-Vergara, University of Technology Sydney, Australia
Updates
Copyright
© 2026 Singh and Withey.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Jeffrey H. Withey, jwithey@med.wayne.edu
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.