<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="brief-report" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">821715</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2021.821715</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Brief Research Report</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Multiple Cases of Bacterial Sequence Erroneously Incorporated Into Publicly Available Chloroplast Genomes</article-title>
<alt-title alt-title-type="left-running-head">Robinson et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Errors in Public Chloroplast References</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Robinson</surname>
<given-names>Aaron J.</given-names>
</name>
<xref ref-type="fn" rid="fn1">
<sup>&#x2020;</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1570467/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Daligault</surname>
<given-names>Hajnalka E.</given-names>
</name>
<xref ref-type="fn" rid="fn1">
<sup>&#x2020;</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1580201/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Kelliher</surname>
<given-names>Julia M.</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/1620963/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>LeBrun</surname>
<given-names>Erick S.</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/1577750/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Chain</surname>
<given-names>Patrick S. G.</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/18535/overview"/>
</contrib>
</contrib-group>
<aff>
<institution>Los Alamos National Laboratory</institution>, <institution>Biosecurity and Public Health Group</institution>, <institution>Bioscience Division</institution>, <addr-line>Los Alamos</addr-line>, <addr-line>NM</addr-line>, <country>United&#x20;States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/999539/overview">Luca Bianco</ext-link>, Fondazione Edmund Mach, Italy</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/313887/overview">Ruslan Kalendar</ext-link>, University of Helsinki, Finland</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/284770/overview">Weilong Hao</ext-link>, Wayne State University, United&#x20;States</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Patrick S. G. Chain, <email>pchain@lanl.gov</email>
</corresp>
<fn fn-type="equal" id="fn1">
<label>
<sup>&#x2020;</sup>
</label>
<p>These authors have contributed equally to this&#x20;work</p>
</fn>
<fn fn-type="other">
<p>This article was submitted to Computational Genomics, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>13</day>
<month>01</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>821715</elocation-id>
<history>
<date date-type="received">
<day>24</day>
<month>11</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>27</day>
<month>12</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Robinson, Daligault, Kelliher, LeBrun and Chain.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Robinson, Daligault, Kelliher, LeBrun and Chain</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Public sequencing databases are invaluable resources to biological researchers, but assessing data veracity as well as the curation and maintenance of such large collections of data can be challenging. Genomes of eukaryotic organelles, such as chloroplasts and other plastids, are particularly susceptible to assembly errors and misrepresentations in these databases due to their close evolutionary relationships with bacteria, which may co-occur within the same environment, as can be the case when sequencing plants. Here, based on sequence similarities with bacterial genomes, we identified several suspicious chloroplast assemblies present in the National Institutes of Health (NIH) Reference Sequence (RefSeq) collection. Investigations into these chloroplast assemblies reveal examples of erroneous integration of bacterial sequences into chloroplast ribosomal RNA (rRNA) loci, often within the rRNA genes, presumably due to the high similarity between plastid and bacterial rRNAs. The bacterial lineages identified within the examined chloroplasts as the most likely source of contamination are either known associates of plants, or co-occur in the same environmental niches as the examined plants. Modifications to the methods used to process untargeted &#x2018;raw&#x2019; shotgun sequencing data from whole genome sequencing efforts, such as the identification and removal of bacterial reads prior to plastome assembly, could eliminate similar errors in the future.</p>
</abstract>
<kwd-group>
<kwd>chloroplast</kwd>
<kwd>plastome</kwd>
<kwd>sequence contamination</kwd>
<kwd>public sequence databases</kwd>
<kwd>genome repositories</kwd>
</kwd-group>
<contract-sponsor id="cn001">U.S. Department of Energy<named-content content-type="fundref-id">10.13039/100000015</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>Publicly available sequence databases, such as those in the International Nucleotide Sequence Database Collaboration (INSDC), are fundamental resources for many types of bioinformatic analyses. With the increased availability of sequencing and a wide array of methods designed for routine use by genomics and bioinformatic novices, there is a constant need to monitor sequence entries and try to assess the quality and veracity of data within these public resources. The errors present in curated and otherwise trusted databases (<xref ref-type="bibr" rid="B15">Steinegger and Salzberg, 2020</xref>; <xref ref-type="bibr" rid="B8">Orakov et&#x20;al., 2021</xref>), such as the National Center for Biotechnology Information (NCBI) RefSeq, are of particular concern as both sequences and associated taxonomic designations reported in these databases are often blindly accepted as accurate by the majority of users, and this database is a common source of reference genomes used for comparative genomic analyses. Given that the nature of most bioinformatic analyses involve similarity searches to references in these databases, it is not uncommon to see transference of annotation and genome errors to other projects or analyses, making it imperative to quickly identify and correct any erroneous submissions within these trusted databases, prior to the propagation of errors.</p>
<p>The NCBI RefSeq database contains a large number of plant and algal nuclear and chloroplast genome reference assemblies, many of which are derived from taxa important to either environmental functioning, agriculture, or medicine. Unfortunately, organelle genome references are not curated as stringently as their nuclear counterparts, increasing the likelihood that erroneous sequences may be published. Additionally, nuclear genomes are almost always published along with the raw sequencing data used to generate the assembly, but based on our examinations, it appears to be quite uncommon to find links to raw data for organelle genomes found in RefSeq or GenBank databases. Particularly in the case of plastids, which are often examined independently from the nuclear genome, the absence of this supporting sequencing data makes it challenging to identify, assess, and correct errors in published assemblies. Additionally, previous screens of the RefSeq and GenBank databases to assess genome assembly quality in terms of potential contamination, to our knowledge, have not included plastome sequences. Evaluation of plastomes would also require special considerations, given their unique evolutionary relationships with bacteria, which complicate assessment of contamination.</p>
<p>While screening plastid sequences in the NCBI RefSeq plastid genome collection (<ext-link ext-link-type="uri" xlink:href="https://www.ncbi.nlm.nih.gov/genome/organelle/">https://www.ncbi.nlm.nih.gov/genome/organelle/</ext-link>), we identified several chloroplast genomes that contained sequence signatures more similar to bacteria than chloroplast. Herein, we present several problematic chloroplast references present in RefSeq which each contain variable and unique bacterial sequence contamination in the regions containing the 16S and 23S rRNA genes. Analysis of the raw sequencing data used to generate these assemblies indicates the inclusion of bacterial DNA. One possible source of the bacterial DNA detected in these chloroplast sequencing projects is potential bacterial associates of the host plant or its environment. This work highlights several concerns unique to the NCBI organelle RefSeq database, and suggests potential changes to help reduce the observed issues.</p>
</sec>
<sec sec-type="methods" id="s2">
<title>Methods</title>
<p>All examined chloroplast genome assemblies were obtained from the 11/2020 NCBI RefSeq release of plastid sequences (<ext-link ext-link-type="uri" xlink:href="https://ftp.ncbi.nlm.nih.gov/refseq/release/plastid/">https://ftp.ncbi.nlm.nih.gov/refseq/release/plastid/</ext-link>). The chloroplast assemblies examined in this work were identified as a result of mapping bacterial metagenomic reads to the entire NCBI RefSeq plastid genome collection (5,569 Plastomes total). Bacterial reads mapped to several chloroplast genome references at a relatively high frequency, which seemed abnormal and were thus investigated more closely. The read mapping results for each of these selected chloroplast assemblies all had similar mapping results, with bacterial reads piling up in the regions containing the 16S and 23S rRNA genes. Annotated chloroplast, plant nuclear, and bacterial sequences and genome assemblies were all obtained from either NCBI RefSeq or GenBank (accessions and sequence ranges provided in results). Alignments and comparisons to the NCBI nucleotide collection (nt) were performed using the BLASTN algorithm (<ext-link ext-link-type="uri" xlink:href="https://blast.ncbi.nlm.nih.gov/Blast.cgi">https://blast.ncbi.nlm.nih.gov/Blast.cgi</ext-link>). Visual alignments presented in figures were generated using the mVISTA and AVID alignment program (<xref ref-type="bibr" rid="B7">Mayor et&#x20;al., 2000</xref>; <xref ref-type="bibr" rid="B3">Frazer et&#x20;al., 2004</xref>) with default settings. All read-based analyses were performed using the EDGE v2.4.0 bioinformatics platform (<xref ref-type="bibr" rid="B6">Li et&#x20;al., 2017</xref>). The following quality-control parameters were used for all examined Illumina and IonTorrent datasets provided by the original submitters: bases from the ends of reads with a Phred score below 20 were trimmed, minimum read length after trimming had to be at least 50&#xa0;bp, trimmed reads containing 10 or more continuous &#x201c;N&#x201d; bases were removed, and sequences comprised of 85% or more low complexity sequence (e.g., mono-/di-nucleotide sequence) were removed. Detailed information on these quality-control parameters can be found in the documentation for FaOCs (<ext-link ext-link-type="uri" xlink:href="https://github.com/LANL-Bioinformatics/FaQCs">https://github.com/LANL-Bioinformatics/FaQCs</ext-link>). Taxonomic classification of quality-controlled reads was performed using GOTTCHA2 (<xref ref-type="bibr" rid="B4">Freitas et&#x20;al., 2015</xref>) with a database generated from the NCBI bacterial RefSeq90 complete genome collection, which contained 14,529 bacterial genomes. Additionally, all reads were aligned to the bacterial RefSeq collection using the BWA-MEM algorithm (<xref ref-type="bibr" rid="B5">Li, 2013</xref>). Reference based analysis and read mapping were also performed using the BWA-MEM algorithm. Bacterial reads were filtered from Illumina and IonTorrent sequencing datasets by aligning reads to both the NCBI bacterial RefSeq collection and/or selected bacterial references, and any sequences with 90% or greater sequence identity were removed from the dataset. Alignments for phylogenetic analysis were performed using Clustal Omega v1.2.4 (<xref ref-type="bibr" rid="B13">Sievers et&#x20;al., 2011</xref>) and phylogenetic trees were produced using RAxML v8.2.12 (<xref ref-type="bibr" rid="B14">Stamatakis, 2014</xref>) using the rapid bootstrap algorithm with 100 iterations and the GTRCAT model of substitution. Phylogenetic trees were edited using FigTree v1.4.4 (<ext-link ext-link-type="uri" xlink:href="https://github.com/rambaut/ftree">https://github.com/rambaut/figtree</ext-link>), and only bootstrap values &#x2265;75 are shown. The sequence data used to generate the erroneous <italic>P. japonica</italic> (NC_037440.1), <italic>R. parvula</italic> (NC_031180.2) and <italic>D. unilobum</italic> (NC_035853.1) chloroplast references have been uploaded to the NCBI SRA database (<xref ref-type="sec" rid="s10">Supplementary Table&#x20;S1</xref>).</p>
</sec>
<sec sec-type="results" id="s3">
<title>Results</title>
<p>The rRNA regions from several RefSeq chloroplast genome assemblies were examined for potential bacterial sequence contamination. These chloroplast genomes, which represent a species of orchid and two red algal species, were all obtained and submitted by separate research groups, using distinct methods for assembly and annotation (<xref ref-type="sec" rid="s10">Supplementary Table S1</xref>). When possible, the raw sequencing data used to generate these chloroplast genome assemblies was screened for bacterial contamination, and bacterial filtered data was compared to the assembly to assess impacts. The ribosomal RNA (rRNA) regions in the original assemblies appear to contain sequences derived from bacterial reads, ranging from relatively short stretches of sequence to an extreme case where &#x223c;1.5&#xa0;kb of bacterial-like 16S sequence was inserted immediately adjacent and upstream of the annotated chloroplast 16S sequence. Descriptions of these anomalous findings detected within the published and annotated rRNA sequences from these chloroplast genome assemblies are detailed&#x20;below.</p>
<sec id="s3-1">
<title>Two <italic>Platanthera japonica</italic> Chloroplast Assemblies Contain Bacterial 23S rRNA Sequence</title>
<p>Alignments between the annotated rRNA operon sequences from the <italic>Platanthera japonica</italic> chloroplast (NCBI accession: NC_037440.1) and <italic>Platanthera chlorantha</italic> (NC_044626.1) show that while the majority of the region is highly similar between both <italic>Platanthera</italic> species, parts of the 23S gene are highly dissimilar (<xref ref-type="fig" rid="F1">Figure&#x20;1</xref>). The annotated <italic>P. japonica</italic> 23S sequence was then aligned to the NCBI nucleotide collection (nt) using BLASTN, which revealed the sequence was highly similar to sequences from the bacterial genus <italic>Klebsiella</italic> (&#x3e;99%). Taxonomic classification of the reads used to generate the <italic>P. japonica</italic> chloroplast assembly (NC_037440.1) indicated the presence of contaminating <italic>Klebsiella</italic> reads in the dataset (<xref ref-type="sec" rid="s10">Supplementary Table S1</xref>). Alignments presented in <xref ref-type="fig" rid="F1">Figure&#x20;1</xref> demonstrate that the 23S sequences from <italic>Klebsiella variicola</italic> strain FH-1 (NZ_CP054254.1) and the <italic>P. japonica</italic> chloroplast (NC_037440.1) are more similar to each other than either is to the <italic>P. chlorantha</italic> (NC_044626.1) sequence. Once bacterial reads were removed from the dataset, the filtered reads were mapped to <italic>P. chlorantha</italic> (NC_044626.1) chloroplast assembly to obtain the correct <italic>P. japonica</italic> (NC_037440.1) 23S sequence. This corrected sequence was provided to the original submitters of the chloroplast genome assembly and was compared to the updated assembly (generated by original submitters) to ensure proper representation (<xref ref-type="sec" rid="s10">Supplementary Table&#x20;S1</xref>).</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Alignments showing similarity between <italic>Platanthera</italic> chloroplast and <italic>Klebsiella variicola</italic> rRNA sequences when <bold>(A)</bold> <italic>P. chlorantha</italic> (NC_044626.1) and <bold>(B)</bold> <italic>K. variicola</italic> (NZ_CP054254.1) are used as references.</p>
</caption>
<graphic xlink:href="fgene-12-821715-g001.tif"/>
</fig>
<p>A separate <italic>P</italic>. <italic>japonica</italic> chloroplast assembly (MN631092) was subsequently sequenced and published after the release of the erroneous <italic>P. japonica</italic> (NC_037440.1) plastome sequence. Both plastomes were derived from leaves collected from the same geographic location, but different references were utilized by the two groups to guide the assembly (<xref ref-type="sec" rid="s10">Supplementary Table S1</xref>) (<xref ref-type="bibr" rid="B2">Dong et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B17">Zhang et&#x20;al., 2020</xref>). BLASTN alignments between these two <italic>P</italic>. <italic>japonica</italic> chloroplast assemblies revealed a high overall similarity (95% coverage, 99.45% identity), including identical 23S sequences in both chloroplast assemblies. Therefore, two separate erroneous chloroplast sequences exist within GenBank for this plant species and after sharing our findings with the original submitters and NCBI staff, both the GenBank (MG925368.2) and RefSeq (NC_037440.2) chloroplast entries have now been updated with the corrected 23S sequences (<xref ref-type="sec" rid="s10">Supplementary Table&#x20;S1</xref>).</p>
</sec>
<sec id="s3-2">
<title>Bacterial Sequences Present in <italic>Rhodochaete parvula</italic> rRNA Regions</title>
<p>Alignments of the annotated 16S sequences from the <italic>Rhodochaete parvula</italic> (NC_031180.2) chloroplast assembly to the NCBI nucleotide collection (nt) revealed that one of the sequences (&#x201c;copy A&#x201d;) was more similar (90&#x2013;95% identity) to other chloroplast sequences, while the other sequence (&#x201c;copy B&#x201d;) shared only 86.64% identity with the top chloroplast match in the database. Phylogenetic trees produced from an alignment of 16S sequences from this <italic>R. parvula</italic> assembly, closely related chloroplasts, cyanobacteria, and bacteria indicate that the <italic>R. parvula</italic> 16S copy A (NC_031180.2:c187891-186416) sequence is most closely related to sequences from another <italic>R. parvula</italic> isolate (KY709212.1), while copy B (NC_031180.2:208102-209574) is more distantly related to all examined plastid and cyanobacterial sequences (<xref ref-type="fig" rid="F2">Figure&#x20;2A</xref>). Similar phylogenetic analyses conducted with the 23S sequences indicate more distant relationships between both copies of <italic>R. parvula</italic> sequences and sequences from the same plant species than what is observed in other plant species (<xref ref-type="fig" rid="F2">Figure&#x20;2B</xref>). Closer examination of these sequences revealed that short stretches in copy B of the 16S gene and in both 23S copies were highly similar to bacterial rRNA sequences from the genus <italic>Marinobacter.</italic> This led to the hypothesis that bacterial sequences were erroneously incorporated into this chloroplast assembly. Read-based taxonomic classification of the raw sequencing data used to generate this chloroplast assembly (SRA accession SRR16979013) indicated that 33% of the total reads were classified as bacterial using the GOTTCHA2 software (see Methods; <xref ref-type="sec" rid="s10">Supplementary Table S1</xref>). Corrected versions of the 16S and 23S sequences from this <italic>R. parvula</italic> (NC_031180.2) assembly were generated by filtering bacterial reads prior to reassembly (see Methods), and these corrected sequences were confirmed to be highly similar to sequences from the other examined <italic>R. parvula</italic> isolate (KY709212.1) (<xref ref-type="fig" rid="F2">Figure&#x20;2</xref>). These corrected sequences were provided to the original submitters of the chloroplast genome assembly and were compared to the updated assembly (generated by original submitters) to ensure proper representation. After sharing our findings with the original submitters and NCBI staff, the GenBank (KX284728.3) entry has now been updated with the corrected 16S and 23S sequences (<xref ref-type="sec" rid="s10">Supplementary Table&#x20;S1</xref>).</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Phylogenetic trees of the 16S <bold>(A)</bold> and 23S <bold>(B)</bold> sequences, showing relationships between bacterial (black), cyanobacterial (blue), chloroplast sequences (green) and the <italic>R. parvula</italic> sequences (red). The original <italic>R. parvula</italic> chloroplast 16S and 23S sequences (2 copies each) from RefSeq assembly (NC_031180.2) are included as well as the corrected sequences obtained after removal of bacterial reads from the original sequencing dataset. NCBI accession identifiers and sequence ranges are shown in parentheses. Branches with bootstrap support greater than 75% (100 bootstrap replicates) are&#x20;shown.</p>
</caption>
<graphic xlink:href="fgene-12-821715-g002.tif"/>
</fig>
</sec>
<sec id="s3-3">
<title>Inclusion of Bacterial 16S Sequence in <italic>Kappaphycus alvarezii</italic> Chloroplast Assembly</title>
<p>The <italic>Kappaphycus alvarezii</italic> chloroplast RefSeq genome (NC_036637.1) also appears to harbor contaminating sequences that do not belong as part of the reference. Alignments between this <italic>K. alvarezii</italic> chloroplast reference and assembled contigs generated from the original whole genome sequencing (WGS) project (NADL02000598.1) indicated that the <italic>K. alvarezii</italic> chloroplast reference harbors an additional 16S sequence (a &#x223c;1.5&#xa0;Kb insertion) immediately adjacent and upstream (NC_036637.1:27964-29353) of the annotated 16S sequence in the <italic>K. alvarezii</italic> chloroplast reference (<xref ref-type="fig" rid="F3">Figure&#x20;3</xref>). This &#x223c;1.5&#xa0;kb insertion shares 88.40&#x2013;88.52% identity with bacterial 16S sequences, with greatest similarity shared with <italic>Serratia spp</italic>. These findings strongly suggest that exogenous bacterial sequence was assembled as part of the chloroplast reference assembly. While there were also additional issues observed within the annotated 16S and 23S genes in the <italic>K. alvarezii</italic> chloroplast reference genome, the raw sequencing data used to generate these assemblies were unfortunately not available at the time of writing to further investigate this&#x20;issue.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Alignment between rRNA sequences from the <italic>K. alvarezii</italic> chloroplast (reference), <italic>K. alvarezii</italic> whole genome assembly and <italic>Serratia plymuthica</italic>.</p>
</caption>
<graphic xlink:href="fgene-12-821715-g003.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>Discussion</title>
<p>Our investigations into assembly errors detected in the above examined chloroplast genomes highlight the susceptibility of plastomes to erroneous bacterial sequence inclusion, specifically in regions of high sequence similarity, including the ribosomal RNA subunits. Due to the bacterial origin of chloroplasts, coupled with the evolutionary and functional constraints of the ribosomal RNA subunits, significant sequence conservation exists among the rRNA subunits of chloroplasts and bacteria. This sequence similarity can, as elucidated from the examples presented in this work, occasionally complicate the generation of an accurate chloroplast reference genome if bacterial reads are present and not removed prior to plastome assembly. While processes are in place to detect foreign sequence contamination during whole genome submissions, these processes are not run on chloroplast or other organelle sequences that may be submitted separately. Sequencing of pure plant samples can be very challenging, due to the presence of often diverse plant-associated microorganisms throughout the various tissues of the plant host. Indeed, the prevalence of associations between bacteria and plants leads to an almost inevitable bacterial presence in most plant DNA extracts (<xref ref-type="bibr" rid="B11">Rosenblueth et&#x20;al., 2004</xref>).</p>
<p>In the examined chloroplast genomes, the bacterial sequences detected belonged to either previously described plant associated taxa, including close relatives of the plant hosts examined in this work, or to taxa that have been isolated from similar environmental niches. The published <italic>Platanthera japonica</italic> assemblies (NC_037440.1 and MN631092.1) examined in this work contain bacterial sequences most closely resembling <italic>Klebsiella</italic>. Members of the bacterial genus <italic>Klebsiella</italic> are capable of forming beneficial associations with orchids closely related to <italic>P. japonica</italic> (<xref ref-type="bibr" rid="B9">Pavlova et&#x20;al., 2017</xref>) and <italic>K. variicola</italic> isolates can form endophytic relationships with plant hosts (<xref ref-type="bibr" rid="B16">Wei et&#x20;al., 2014</xref>). Sequencing data and the resulting chloroplast assembly of the examined <italic>Rhodochaete parvula</italic> strain revealed the presence of bacterial sequences most closely resembling <italic>Marinobacter</italic>. Both <italic>Marinobacter</italic> and <italic>R. parvula</italic> inhabit similar niches in marine environments (<xref ref-type="bibr" rid="B18">Zuccarello et&#x20;al., 2000</xref>; <xref ref-type="bibr" rid="B10">Rani et&#x20;al., 2017</xref>). Sequences closely resembling <italic>Serratia</italic> were detected in the sequencing project and resulting chloroplast assembly from the red algae <italic>Kappaphycus alvarezii</italic>, and while this particular association has not been described, <italic>Serratia</italic> are known to associate with other plants (<xref ref-type="bibr" rid="B1">Asaf et&#x20;al., 2017</xref>). Many of these potential bacterial-plant associations are one possible explanation for the presence of these reads in the examined sequencing data, however, bacterial sequences can also be introduced into samples through reagent and consumable contamination of extraction and library preparation kits (<xref ref-type="bibr" rid="B12">Salter et&#x20;al., 2014</xref>). However, the amount of bacterial reads observed in the provided sequencing datasets was often quite substantial, often covering the main bacterial reference chromosome at &#x3e;20X average fold coverage, a result which is inconsistent with passive or reagent contamination (<xref ref-type="sec" rid="s10">Supplementary Table&#x20;S1</xref>).</p>
<p>Regardless of the origin of these bacterial reads and sequences, our investigation indicates that the presence of these bacterial reads can lead to errors during the chloroplast assembly process. Due in part to the continuing decrease in sequencing costs, and similar to other sequence databases, plastid sequence databases are growing at an increasing pace. Our results suggest that additional quality screening is needed prior to plastid genome inclusion into these reference databases, with particular attention to regions coding for the ribosomal RNA genes. The assembly errors described in this work were all identified through comparisons with databases containing both bacterial and chloroplast sequences, and a similar approach is used by NCBI to identify suspicious contigs and sequences when uploading nuclear and other non-organelle specific genome assemblies. It is relevant to point out that the chloroplast assemblies investigated in this work were generated using a variety of methods and software (<xref ref-type="sec" rid="s10">Supplementary Table S1</xref>). Furthermore, all but a single assembly were generated using a reference genome, suggesting that additional steps are needed to prevent the inclusion of contaminating bacterial sequences.</p>
<p>While this study is not exhaustive and there are likely to be many other types of errors in plastid reference genomes, given the biological samples and modern sample processing methods used to obtain plastid genomes, careful screening for bacterial data could provide some measure of confidence in the resulting genome. The use of tools designed to identify and exclude bacterial sequences from NGS datasets prior to plastid assembly, such as the ones employed in this work (see Methods), could help reduce or possibly eliminate similar errors in future plastid genome work. Failure to detect and correct these errors can lead to inaccurate predictions of presence/absence in metagenomic samples, improper representations of phylogenetic relationships (and ensuing inferences) and increases the probability of propagation of these errors through the use of incorrect references for assembly.</p>
<p>The chloroplast assemblies investigated in this work are representative of larger shortcomings in the field of genomics and highlight some of the vulnerabilities in publicly available sequence databases. None of these assemblies were previously flagged as potentially problematic, and even after we identified the suspicious nature and various inconsistencies in these assemblies, it was often difficult to track down the original data and submitters. Even when research groups reply to queries and share their raw sequencing data, which is not always the case, the investigation of such discrepancies coupled with further correspondence with busy researchers and the activation energy required to amend the database entries can substantially delay corrections to the public repositories. In some cases, additional sequencing data is required prior to submitting a correction, as was the case with the chloroplast entry for <italic>Diplazium unilobum</italic> (NC_035853.1), where we had contacted the original investigators and reported observing short stretches of bacterial-like sequences embedded within the annotated <italic>D. unilobum</italic> chloroplast 16S gene. After being provided the raw sequencing data (SRR16961501), our investigation clearly showed bacterial contamination caused the observed discrepancies. After rigorous treatment of the raw data for reads of bacterial origin, there remained insufficient chloroplast data to obtain a complete genome. Additional sequencing by the original authors was required prior to obtaining the correct sequence (SRR16974227), and this update has now been submitted during the completion of this manuscript (<xref ref-type="sec" rid="s10">Supplementary Table S1</xref>). For bioinformaticists and comparative genomic researchers automating their data analyses using such reference genomes, an additional concern includes the fact that even once correct genomes are submitted to public repositories like GenBank to update those entries, such as with the <italic>P. japonica</italic> chloroplast genome (now MG925368.2), the RefSeq entry may not be updated, and in the case of <italic>P. japonica</italic> (NC_037440.1), this remained identical to the original erroneous submission until additional communication with NCBI staff resulted in a recent updated RefSeq record, over a year after the GenBank update (<xref ref-type="sec" rid="s10">Supplementary Table S1</xref>). This work highlights the importance of being cautious when using public databases, as unintentional mistakes and errors are not always apparent nor easy to identify. Furthermore, the results presented here raise questions about the potential for similar issues to arise in plastomes of other organisms, particularly in bacterial-derived organelle genomes such as mitochondria. With these investigations, we also wish to emphasize the need to have access to original raw data to verify and validate original submissions and inconsistencies in assembled data, and to bring awareness of issues specific to the assembly of plastomes, with the goal of minimizing future mistakes and increasing overall genome data quality.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/<xref ref-type="sec" rid="s10">Supplementary material</xref>.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>AR, HD, JK, and PC contributed to the planning and design of the study and analyses. AR, HD, and JK identified problematic chloroplast reference genomes examined in this work. AR, HD, and EL contributed to the data collection and analyses. AR, HD, JK, and PC contributed to generating figures and tables. AR, JK, HD, and PC wrote sections of the manuscript. All authors contributed to manuscript revision, and all authors read and approved the submitted version.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>This study was supported by the U.S. Department of Energy, Office of Science, Biological and Environmental Research Division, under award number LANLF59T.</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>We wish to thank the authors and submitters of chloroplast genomes for their willingness to share data and to help with corrections to database entries; namely Zhong-Hu Li and Ruo-Nan Wang (Key Laboratory of Resource Biology and Biotechnology in Western China, Ministry of Education, College of Life Sciences, Northwest University, Xi&#x2019;an 710069, China), Ran Wei (State Key Laboratory of Systematic and Evolutionary Botany, Institute of Botany, Chinese Academy of Sciences, Beijing, China), JunMo Lee and Hwan Su Yoon (Department of Biological Sciences, Sungkyunkwan University, Suwon, 16419 Republic of Korea).</p>
</ack>
<sec id="s10">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2021.821715/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fgene.2021.821715/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="Table1.XLSX" id="SM1" mimetype="application/XLSX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Asaf</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Khan</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Khan</surname>
<given-names>A. L.</given-names>
</name>
<name>
<surname>Waqas</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Shahzad</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>A.-Y.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Bacterial Endophytes from Arid Land Plants Regulate Endogenous Hormone Content and Promote Growth in Crop Plants: an Example of <italic>Sphingomonas</italic> Sp. And <italic>Serratia marcescens</italic>
</article-title>. <source>J.&#x20;Plant Interactions</source> <volume>12</volume>, <fpage>31</fpage>&#x2013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1080/17429145.2016.1274060</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dong</surname>
<given-names>W.-L.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>R.-N.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>N.-Y.</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>W.-B.</given-names>
</name>
<name>
<surname>Fang</surname>
<given-names>M.-F.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Z.-H.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Molecular Evolution of Chloroplast Genomes of Orchid Species: Insights into Phylogenetic Relationship and Adaptive Evolution</article-title>. <source>IJMS</source> <volume>19</volume>, <fpage>716</fpage>. <pub-id pub-id-type="doi">10.3390/ijms19030716</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Frazer</surname>
<given-names>K. A.</given-names>
</name>
<name>
<surname>Pachter</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Poliakov</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Rubin</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Dubchak</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>VISTA: Computational Tools for Comparative Genomics</article-title>. <source>Nucleic Acids Res.</source> <volume>32</volume>, <fpage>W273</fpage>&#x2013;<lpage>W279</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkh458</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Freitas</surname>
<given-names>T. A. K.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>P.-E.</given-names>
</name>
<name>
<surname>Scholz</surname>
<given-names>M. B.</given-names>
</name>
<name>
<surname>Chain</surname>
<given-names>P. S. G.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Accurate Read-Based Metagenome Characterization Using a Hierarchical Suite of Unique Signatures</article-title>. <source>Nucleic Acids Res.</source> <volume>43</volume>, <fpage>e69</fpage>. <pub-id pub-id-type="doi">10.1093/nar/gkv180</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2013</year>). <source>Aligning Sequence Reads, Clone Sequences and Assembly Contigs with BWA-MEM</source>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1303.3997">http://arxiv.org/abs/1303.3997</ext-link> (Accessed November 1, 2021)</comment>. </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>P.-E.</given-names>
</name>
<name>
<surname>Lo</surname>
<given-names>C.-C.</given-names>
</name>
<name>
<surname>Anderson</surname>
<given-names>J.&#x20;J.</given-names>
</name>
<name>
<surname>Davenport</surname>
<given-names>K. W.</given-names>
</name>
<name>
<surname>Bishop-Lilly</surname>
<given-names>K. A.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Enabling the Democratization of the Genomics Revolution with a Fully Integrated Web-Based Bioinformatics Platform</article-title>. <source>Nucleic Acids Res.</source> <volume>45</volume>, <fpage>67</fpage>&#x2013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkw1027</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mayor</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Brudno</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Schwartz</surname>
<given-names>J.&#x20;R.</given-names>
</name>
<name>
<surname>Poliakov</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Rubin</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Frazer</surname>
<given-names>K. A.</given-names>
</name>
<etal/>
</person-group> (<year>2000</year>). <article-title>VISTA : Visualizing Global DNA Sequence Alignments of Arbitrary Length</article-title>. <source>Bioinformatics</source> <volume>16</volume>, <fpage>1046</fpage>&#x2013;<lpage>1047</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/16.11.1046</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Orakov</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Fullam</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Coelho</surname>
<given-names>L. P.</given-names>
</name>
<name>
<surname>Khedkar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Szklarczyk</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Mende</surname>
<given-names>D. R.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>GUNC: Detection of Chimerism and Contamination in Prokaryotic Genomes</article-title>. <source>Genome Biol.</source> <volume>22</volume>, <fpage>178</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-021-02393-0</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pavlova</surname>
<given-names>A. S.</given-names>
</name>
<name>
<surname>Leontieva</surname>
<given-names>M. R.</given-names>
</name>
<name>
<surname>Smirnova</surname>
<given-names>T. A.</given-names>
</name>
<name>
<surname>Kolomeitseva</surname>
<given-names>G. L.</given-names>
</name>
<name>
<surname>Netrusov</surname>
<given-names>A. I.</given-names>
</name>
<name>
<surname>Tsavkelova</surname>
<given-names>E. A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Colonization Strategy of the Endophytic Plant Growth-Promoting Strains of Pseudomonas fluorescens and Klebsiella Oxytocaon the Seeds, Seedlings and Roots of the Epiphytic orchid, Dendrobium nobile Lindl</article-title>. <source>J.&#x20;Appl. Microbiol.</source> <volume>123</volume>, <fpage>217</fpage>&#x2013;<lpage>232</lpage>. <pub-id pub-id-type="doi">10.1111/jam.13481</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rani</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Koh</surname>
<given-names>H.-W.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Rhee</surname>
<given-names>S.-K.</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>S.-J.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Marinobacter Salinus Sp. nov., a Moderately Halophilic Bacterium Isolated from a Tidal Flat Environment</article-title>. <source>Int. J.&#x20;Syst. Evol. Microbiol.</source> <volume>67</volume>, <fpage>205</fpage>&#x2013;<lpage>211</lpage>. <pub-id pub-id-type="doi">10.1099/ijsem.0.001587</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rosenblueth</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mart&#xed;nez</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Silva</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Mart&#xed;nez-Romero</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Klebsiella Variicola, A Novel Species with Clinical and Plant-Associated Isolates</article-title>. <source>Syst. Appl. Microbiol.</source> <volume>27</volume>, <fpage>27</fpage>&#x2013;<lpage>35</lpage>. <pub-id pub-id-type="doi">10.1078/0723-2020-00261</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Salter</surname>
<given-names>S. J.</given-names>
</name>
<name>
<surname>Cox</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Turek</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Calus</surname>
<given-names>S. T.</given-names>
</name>
<name>
<surname>Cookson</surname>
<given-names>W. O.</given-names>
</name>
<name>
<surname>Moffatt</surname>
<given-names>M. F.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Reagent and Laboratory Contamination Can Critically Impact Sequence-Based Microbiome Analyses</article-title>. <source>BMC Biol.</source> <volume>12</volume>, <fpage>87</fpage>. <pub-id pub-id-type="doi">10.1186/s12915-014-0087-z</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sievers</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Wilm</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Dineen</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Gibson</surname>
<given-names>T. J.</given-names>
</name>
<name>
<surname>Karplus</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>W.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <article-title>Fast, Scalable Generation of High&#x2010;quality Protein Multiple Sequence Alignments Using Clustal Omega</article-title>. <source>Mol. Syst. Biol.</source> <volume>7</volume>, <fpage>539</fpage>. <pub-id pub-id-type="doi">10.1038/msb.2011.75</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stamatakis</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>RAxML Version 8: a Tool for Phylogenetic Analysis and post-analysis of Large Phylogenies</article-title>. <source>Bioinformatics</source> <volume>30</volume>, <fpage>1312</fpage>&#x2013;<lpage>1313</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu033</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Steinegger</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Salzberg</surname>
<given-names>S. L.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Terminating Contamination: Large-Scale Search Identifies More Than 2,000,000 Contaminated Entries in GenBank</article-title>. <source>Genome Biol.</source> <volume>21</volume>, <fpage>115</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-020-02023-1</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>C.-Y.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>L.-J.</given-names>
</name>
<name>
<surname>Xing</surname>
<given-names>Y.-X.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>C.-J.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>L.-T.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Endophytic Nitrogen-Fixing Klebsiella Variicola Strain DX120E Promotes Sugarcane Growth</article-title>. <source>Biol. Fertil. Soils</source> <volume>50</volume>, <fpage>657</fpage>&#x2013;<lpage>666</lpage>. <pub-id pub-id-type="doi">10.1007/s00374-013-0878-3</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>The Complete Plastome Sequence of Platanthera Japonica (Orchidaceae): an Endangered Medicinal and Ornamental Plant</article-title>. <source>Mitochondrial DNA B</source> <volume>5</volume>, <fpage>468</fpage>&#x2013;<lpage>469</lpage>. <pub-id pub-id-type="doi">10.1080/23802359.2019.1704643</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zuccarello</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>West</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bitans</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kraft</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>Molecular Phylogeny of Rhodochaete Parvula (Bangiophycidae, Rhodophyta)</article-title>. <source>Phycologia</source> <volume>39</volume>, <fpage>75</fpage>&#x2013;<lpage>81</lpage>. <pub-id pub-id-type="doi">10.2216/i0031-8884-39-1-75.1</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>