Abstract
Polyketide synthases (PKSs) utilize the products of primary metabolism to synthesize a wide array of secondary metabolites in both prokaryotic and eukaryotic organisms. PKSs can be grouped into three distinct classes, types I, II, and III, based on enzyme structure, substrate specificity, and catalytic mechanisms. The type III PKS enzymes function as homodimers, and are the only class of PKS that do not require acyl carrier protein. Plant type III PKS enzymes, also known as chalcone synthase (CHS)-like enzymes, are of particular interest due to their functional diversity. In this study, we mined type III PKS gene sequences from the genomes of six aquatic algae and 25 land plants (1 bryophyte, 1 lycophyte, 2 basal angiosperms, 16 core eudicots, and 5 monocots). PKS III sequences were found relatively conserved in all embryophytes, but not exist in algae. We also examined gene expression patterns by analyzing available transcriptome data, and identified potential cis-regulatory elements in upstream sequences. Phylogenetic trees of dicots angiosperms showed that plant type III PKS proteins fall into three clades. Clade A contains CHS/STS-type enzymes coding genes with diverse transcriptional expression patterns and enzymatic functions, while clade B is further divided into subclades b1 and b2, which consist of anther-specific CHS-like enzymes. Differentiation regions, such as amino acids 196-207 between clades A and B, and predicted positive selected sites within α-helixes in late appeared branches of clade A, account for the major diversification in substrate choice and catalytic reaction. The integrity and location of conserved cis-elements containing MYB and bHLH binding sites can affect transcription levels. Potential binding sites for transcription factors such as WRKY, SPL, or AP2/EREBP may contribute to tissue- or taxon-specific differences in gene expression. Our data shows that gene duplications and functional diversification of plant type III PKS enzymes played a critical role in the ancient conquest of the land by early plants and angiosperm diversification.
Introduction
Polyketide synthase (PKS) enzymes play critical roles in bridging primary and secondary metabolism in bacteria, fungi, and plants by catalyzing the sequential condensation of two-carbon acetate units into a growing polyketide chain. PKS enzymes are classified into types I, II, and III based on their structural configurations and catalytic mechanisms. Only the type III PKSs are the essential condensing enzymes that act directly on acyl-CoA substracts in the absence of acyl carrier protein (). Type III PKS enzymes in plants, also known as the chalcone synthase (CHS)-like family, synthesize various biologically important components responsible for photoprotection, flower pigmentation, antimicrobial defense, pollen fertility, and the induction of root nodulation (; ; ). Anti-oxidant and anti-cancer properties of polyketides have attracted considerable attention in pharmaceutical engineering (; ).
Chalcone synthase-like family enzymes have acquired multiple activities, and include enzymes such as CHS, stilbene synthase (STS), 2-pyrone synthase (2-PS), bibenzyle synthase (BBS), homoeriodictyol/eriodictyol synthase (HEDS), acridone synthase (ACS), benzophenone synthase (BPS), phlorisovalerophenone synthase (VPS), coumaroyl triacetic acid synthase (CTAS), benzalacetone synthase (BAS), C-methylchalcone synthase (PstrCHS2), biphenyl synthase (BIS), stilbenecarboxylate synthase (STCS), pentaketide chromone synthase (PCS), hexaketide synthase (HKS), aloesone synthase (ALS), octaketide synthase (OKS), and anther-specific CHS-like (ASCHSLE) (). These enzymes use CoA ester substrates that vary from aliphatic to aromatic, from small molecules to large molecules, and also from polar to non-polar ().
The crystalline structures of several plant type III PKS enzymes have been characterized. All of these PKSs function through a homodimer and each monomer possesses an active site (). Slight but specific modifications in the active site lead to remarkable functional diversity by influencing substrate selection, the number of polyketide chain extensions, and the mechanism of cyclization reactions. This “steric modulation” model is supported by crystallographic elucidation along with catalytic validation (). For example, the Thr197Leu, Gly256Leu and Ser338Ile (numbering as in Medicago sativa CHS) substitutions in the CHS amino acid sequence resulting in a complete transformation to a 2-PS enzyme (). The Ser132Thr, Ala133Ser and Val265Phe substitutions are enough to change ACS to CHS (). The diketide forming activity of Rheum palmatum BAS is attributed to Phe215Leu substitution (). In addition, a single change of His to Glu at position 166 alters the substrate preference of AhSTS from p-coumaroyl-CoA to cinnamoyl-CoA (). But current efforts in CHS and STS conversion only partially alter the catalytic reactions through active sites or their geometry (; ). More evidence is needed for these CHS and STS enzymes, even with the help of the crystalline structures of STS ().
Soon after the CHS gene was cloned from Petroselinum crispum (), tissue-specific and environmentally sensitive expression of CHS was found to be widespread in plants. The expression of PhCHS-A and J in Petunia hybrida are restricted to the flower with high levels, while PhCHS-B and G transcripts are induced by UV light in vegetative tissues (). Except for family members that are expressed in the leaves and flowers, several M. sativa CHS genes are preferentially expressed in roots and nodules, and can be induced by pathogen inoculation (; ). In Ipomoea purpurea, IpCHS-D and E are predominantly expressed in the flower limb and tube, while IpCHS-A/B/C are expressed at very low levels (; ). Among two characterized members of the Antirrhinum majus CHS-like gene family, AmCHS1 mRNA specifically accumulates in the petal, whereas AmCHS2 expression is negligible in the petal and other organs (; ). In Gerbera hybrida, GhCHS1 and GhCHS3 are specifically expressed in the pappus, GhCHS4 expression is dominant in petals and red vegetative tissues, and Gh2PS1 (GhCHS2) is universally expressed in all tissues (; ). The MYB-bHLH-WDR ternary transcriptional activation complex was proposed to regulate CHS transcription, and this has been confirmed in a number of species (; ). In addition, a group of anther-specific CHS-like genes highly express themselves in uninuclear microspores and the tapetum. They exhibited significant differences in amino acid sequences compared to other CHS-like family members (). However, the regulatory mechanisms for these CHS-like genes are not well studied.
We have observed that there is great diversity in PKS III enzyme gene coding sequences or in the regulatory elements. Amino acids in the active site provide substrate- and product-specificity to PKSs in vivo or in vitro. Also, in vivo expression levels of genes in different cell types constitute another dimension of specificity. Similar to key effective sites in the protein, potential cis-elements in regulatory regions can undergo equally dramatic changes and stabilization, corresponding to developmental or environmental regulation (). However, gain-of-function and loss-of-function changes in protein coding sequences or regulatory regions always occur alternatively during evolution, a combination that offers the most fitness for the populations that will be selected by nature (; ). When considering the PKS III family, product categories determined by enzyme structure/function and relative expression levels determined by regulatory elements all have specific biological significance. It will be meaningful to dissect out the evolutionary features out of these interlaced aspects by answering questions such as: (1) What diversification patterns did the PKS III family follow? (2) How is the diversification of enzyme sequences and regulatory elements connected with each other?
Fortunately, the public omics database, along with experimental validations, provide convenient source of molecular genetic data. Here we performed genome-wide searches for type III PKS gene sequences from 25 land plant species, and collected tissue-specific transcriptional abundance information from RNA-seq or MicroArray transcriptomes of core eudicots, in order to investigate the comprehensive evolutionary pattern of this family from enzyme sequences, regulatory elements, and the combination of the two.
Materials and Methods
Identification of PKS III Homologous Genes in Plants
The genome sequences of 31 plant species were downloaded from several genomics data portals, i.e., Phytozome1 (), Ensembl Plants2 (), TAIR3 (), BRAD4 (), and CuGenDB5 (full list, Supplementary Table S1). Nucleotide searches were performed using BLAST () to identify PKS III homologs against genome reference sequences or CDS sequences, using PKS III sequences from the literatures (; ) as queries. The threshold was set to an E-value ≤ 1E-5. A PKS gene was preliminary determined if it was found in both CDS and the corresponding genome sequence locus. Hits found only in genome sequences but not in CDS files were defined as “fragment” loci (shown in Supplementary Figure S1). Any PKS gene with frame shifts, premature termination codons, or low coverage (the alignment length less than 300 bp) was also defined as a “fragment.” Chromosomal locus of PKS genes and fragments were shown in Supplementary Figure S1.
In addition, the redundant sequences in four species, Physcomitrella patens, Vitis vinifera, Glycine max, and M. truncatula, were filtered by using cd-hit program () (threshold: 0.95 for the former two and 0.98 for the latter two). The remaining gene sequences with correct and complete open reading frames (ORFs) were used to construct the phylogenetic trees and estimate expression levels. A list of all sequences is in Supplementary Table S2.
Phylogenetic Analysis of PKS III Proteins
Predicted amino-acid sequences translated from protein coding nucleotide sequences were aligned with MAFFT (), then transformed into corresponding codon sequences using PAL2NAL (). A test of substitution saturation was performed in DAMBE (). The best-fit amino acid substitution model (JTT+G) was selected by MEGA (). Maximum likelihood (ML) analyses were performed in RAxML () with 1000 bootstrap replications. A Bayesian inference (BI) tree was conducted in MrBayes (). Two independent MCMC runs, each with four chains (three hot, one cold) were run simultaneously starting from a random tree, with sampling stopping when the convergence diagnostic falls below 0.01. The first 25% samples were discarded as burnin, and the remaining trees were used to construct the 50% majority-rule consensus tree.
Selection Analysis
To detect the changes in evolutionary rates and signatures of positive selection, we analyzed the alignments of codon sequences and the ML tree under a ML framework using CODEML program in PAML 4.8 (). The one-ratio model assumes the same ω (ω = dN/dS; where dN is the non-synonymous substitution rate and dS is the synonymous substitution rate) for all branches. The two-ratio model assumes a foreground ω parameter for each appointed branch and a background ω for all other branches (). Models were compared using likelihood ratio tests (LRTs) of the log likelihood (InL). 2| ΔlnL| values between models prepared and degrees of freedom were used in a chi-square test with a significance threshold of P < 0.01. Because two-ratio models showed that the ω-values for several branches were significantly different from the one-ratio models, we further used branch-site model A to test for sites that were potentially under positive selection on the branch. The model assumes four classes of sites. The first two sites have ω0 (0 < ω0 < 1) and ω1 (ω1 = 1) along all lineages in the phylogeny, whereas the third and fourth have ω2 along the appointed branch, but ω0 and ω1 along other background branches. The branch-site model A was compared with the null model and with the nearly neutral model (M1) (). Results from PAML are given in Supplementary Table S3. Ancestral state reconstruction analysis was performed in MEGA ().
Tissue Specific Expression Levels
Gene expression data for different tissues (root, stem, leaf, flower, and fruit) were obtained from public databases (Supplementary Table S1). Microarray data was normalized using the RMA method (). RNA-seq read data was first filtered using the NGSQCtoolkit, then mapped to reference genome sequences with TopHat (). FPKM values were calculated and normalized with the Cuffquant and Cuffnorm pipeline in Cufflinks (). All values were Log2-transformed.
In order to compare transcript abundance between species, expression levels were transformed to a range of 0–1 within each species by formula: (target value-minimum value)/(maximum value-minimum value). Figures were generated by ggplot2 package in R (). Expression values were listed in Supplementary Table S4.
Conserved Motif Analysis
Upstream sequences of PKS III genes from -2000 to the initiation codon were obtained by using BioMarts () or PERL scripts. These sequences were first submitted to PLACE () or PlantCARE () for searching the annotated motifs. Motifs gathered to regions from -1000 upstream to the initiation codon. Then, the MEME suite () was used to analyze conserved motifs among all upstream sequences, or among sequences from each lineage de novo. We ran the MEME program under the “anr” (any number of repetitions) mode to find motifs exhaustively, and then used TOMTOM (program in MEME suite) to compare motifs found in different lineages. By doing this, a motif distributed around -300 to -100 bp away from the ATG codon was found to be conserved in the majority of upstream sequences. Again using MEME, sequences from -1,000 to the ATG sequences were executed under the “zoops” (zero or one occurrence per sequence) mode, among all upstream sequences, or among sequences from each lineage. Motifs identified de novo were annotated by GOMO (program in MEME suite), or submitted to PLACE () or PlantCARE () to annotate.
Results
Result 1 Genome-Wide Distribution and Copy Number Variation in PKS III Gene Family: Genes from Moss to Flowering Plants
Thirty-one plant species with whole-genome sequences were chose. Amongst these are six algae (Micromonas pusilla, Coccomyxa subellipsoidea, Ostreococcus lucimarinus, Volvox carteri, Chlamydomonas reinhardtii, and Klebsormidium flaccidum), one moss (P. patens), one lycophyte (Selaginella moellendorffii), two basal angiosperms (Amborella trichopoda and Aquilegia coerulea), 16 core eudicots (V. vinifera, Citrus sinensis, C. clementine, Arabidopsis thaliana, Brassica oleracea, B. rapa, Populus trichocarpa, M. truncatula, G. max, Cucumis sativus, Fragaria vesca, Prunus persica, Malus domestica, Mimulus guttatus, Solanum lycopersicum, and S. tuberosum), and five monocots (Zostera marina, Spirodela polyrhiza, Dendrobium officinale, Oryza sativa, and Zea mays) (Supplementary Table S1). Whole-genome sequences and putative gene sequences were both used in BLAST searches to identify type III PKS genes. The BLASTN with threshold E-value = 1E-5 was not available to identify any PKS III gene from all six algae genome data. In contrast, all of the 25 land plant species gave positive hits. A phylogenetic species tree of these land plants is shown based on the APG III and updates (Angiosperm Phylogeny Group; ) (Figure 1A). The copy number variation (CNV) of PKS genes with complete ORF are given in Figure 1B left and right columns. The total number of putative genes in each genome is also shown (Figure 1C).
FIGURE 1
The numbers of complete ORFs and fragments revealed that PKS III family genes occupy a very small proportion of genomes, and their numbers expand in the genomes of moss, grape, and two bean species. Association with recent whole-genome duplication events occurring at the family or genus levels; for example, in the Brassicaceae, after speciation from the common ancestor of Arabidopsis and Brassica, the ancestor of B. rapa and B. oleracea experienced a triploidization event. Whole-genome polyploidization resulted in an increase in chromosome and gene numbers; although many species subsequently experienced a reduction in chromosome numbers and gene loss (
When surveying the chromosomal distributions of PKS genes or homologous fragments, we found that most PKS loci show scattered chromosomal distributions, while some copies exhibit tandem repeats. Family members embedded within small chromosomal regions always showed higher similarity (95–100%) than other copies (<95%; Supplementary Table S2), e.g., P. patens: chr2 (24507741-24814900), chr19 (2508254-3629901), V. vinifera: chr10 (14216111-14306520), chr16 (16238965-16711898), M. truncatula: chr1 (44127878-44142083), chr7 (5283754-5315993), and G. max: chr8 (8384741-8519303) (full list, Supplementary Figure S1). These highly similar sequences were excluded from our following analyses since they brought redundant calculations.
A preliminary ML tree (Figure 2), which was constructed using 234 PKS III protein sequences with complete function structure from all 25 land plants. The PKS III ML tree mirrored the species phylogenetic relationships between these taxa, as the branches of Leguminosae, Rosaceae, Scrophulariaceae-Solanaceae, and Gramineae were identified. Associating with the highly determined evolutionary relationships of these species, we speculated the expansion and diversification history of PKS III family in the lineage of core eudicots (Figure 1A). Diversifications mainly happened at the divergence points of Superrosides and Superasterids, and Fabidae and Malvidae, and fixed at the family level during the specification and gene loss events. The last two branches of the tree, composed of sequences covering all of the angiosperm taxa, revealed extreme conservation of this kind of PKS III enzyme. It would be interesting to further analyze the protein structures and expression patterns of genes in these branches. The monocots are estimated with an early divergence time from basal angiosperms than the core eudicot lineage (
FIGURE 2

Maximum likelihood (ML) tree of 234 protein sequences (Supplementary Table S2) from 25 land plants (one bryophyte, one lycophyte, two basal angiosperms, 16 core eudicots, and 5 monocots).
Result 2 Phylogenetic Relationship and Protein Sequence Diversification within PKS III Family
In order to understand diversification patterns of PKS III proteins, we focused on core eudicot lineage sequences. A number of 226 sequences were collected, including 190 complete PKS III ORFs from projected sampled species (Figure 1, bold letter species) with monocots removed, and 36 functional validated reference sequences in previous work (full list, footnote of Figure 4).
We first calculated the pairwise distances (transitions and transversions) for these PKS III gene sequences, and calculate their frequency distribution (Figure 3). As a result, three peaks ranging from 0–0.25, 0.25–1.3, and 1.3–2.5 are displayed separately. Values forming the 0–0.25 peak originated mainly from clustered genes and homologous genes of very closely related species. The sequence relationships which forming the values in the latter two peaks could be distinguished by phylogenetic analysis using protein sequences.
FIGURE 3

Frequency distribution of pairwise genetic distances (transitions and transversions) for PKS III genes. Three peaks, with genetic distances ranging from 0–0.25, 0.25–1.3, and 1.3–2.5, are shown in green, blue, and red, respectively.
Polyketide synthases III protein sequences were used to construct the ML (Figure 4A) and BI trees (Figure 5A). When constructing the ML tree, PKS III sequences with crystal structures and enzymatic activity verification summarized in reviews (
FIGURE 4

Maximum likelihood tree showing evolutionary relationships for basal angiosperm and core eudicots PKS III proteins, with corresponding enzyme active sites and estimates of selective pressure. (A) ML tree of plant PKS III proteins. The protein names in color are 190 sequences from 20 species (bold letter species in Figure 1). Thirty-six names in gray are proteins translated from nucleotide sequences mentioned in previous works. List as follows: Petroselinum crispum PcCHS (V01538.1) (
FIGURE 5

Bayesian inference tree of 190 plant PKS III proteins from 20 species, with corresponding gene expression levels and conserved cis-elements predicted. (A) BI tree of plant PKS III proteins. (B) Gene expression levels were normalized from tissue-specific transcriptomes. The length of the colored bars indicated the relative expression level within a genome. Spaces without colored bar represent missing value. (C) Locations of conserved sequences. (D) Conserved sequences and cis-elements estimated.
We show the amino acids of the catalytic triad (red squares), coenzyme A binding site (green squares), and functional diversity (blue squares) sites for 226 sequences in the ML tree, based on the annotated sequence and secondary structure of alfalfa CHS (MsCHS2 in the tree; Figure 4B). Among the three types of residues, the catalytic triad was the most conserved, while the other two showed slight diversification in different phylogenetic groups. If we only consider the conserved amino acid residues of the active site, clade B shows relatively large differences when compared to moss and fern sequences and clade A. We combined the active site residues with whole sequence alignments, and although large differences were present in clade B, the catalytic triad residues Cys164, His303, and Asn336, and “gatekeeper” residues Phe215 and Phe265 of the core chemical machinery were conserved. Significant differences were found in two sections in β4 and β6, 98–138 and 196–207, both functional diversification hot-spot regions of PKS III family enzymes. The alfalfa CHS Met137 and Pro138 residues are regarded as the contact point of two monomers (
In order to detect changes in selection pressure, we performed one-ratio, two-ratio and branch-site models in each branch using PAML (Figure 4C; Supplementary Table S3). Selective pressures in clades A and B in the two-ratio models were all significantly different from the one-ratio model. Positive selection was detected in clade A, while clade B members were under more constrained purifying selection (0.006). In clade B, selective pressure values for the two subclades b1 (0.10) and b2 (0.15) were relatively similar. Divergence in regions of clade B enzymes also show further divergence in subclades b1 and b2, such as active site residues 98 and 196. In clade A, branch-site models against each branch revealed that positive selection was mainly in branch a6’. The positive sites were estimated to be 121 K→I, 208 S→V, 276 S→G, 300 W→Y, 340 A→P, located in α-helix type secondary structures, and 264 T→E, 265 F→Y, 266 H→Y, located in β-turn number 11 (Figure 6). The evolutionary process in branch a6’ sequences at these amino acid sites through ancestral state reconstruction analysis are also shown (Supplementary Figure S2).
FIGURE 6

Sequence frequency of amino acid residues of branch a6’ (colored) and other branches (black). Generated by WebLogo (http://weblogo.berkeley.edu/logo.cgi) (
Result 3 Transcriptional Expression Levels of PKS III Genes Were Related to Cis-Elements
Polyketide synthases III enzymes participate in secondary metabolic processes that are important to plant development and defense. How and when the various genes are expressed is an important aspect of functional diversification. Currently, large amounts of transcriptome data have been submitted to public databases, with sampling from various tissues and growth conditions. Tissue samples taken under natural growth conditions can allow for direct comparison between all species after data normalization. Therefore, we collected gene expression data from different tissues, including root, stem, leaf, flower, and fruit. We also included some specialized organs, such as tendrils of cucumber and nodules of soybeans. Because the expression levels of PKS III genes were highly correlated with certain tissues, we found that five main tissues could adequately represent their gene expression profiles. We found ten angiosperm species with such tissue-specific data sets. After technical normalization to eliminate experimental error caused by sampling and sequencing, Log2-transformed expression values were then 0–1 transformed within each species (see Materials and Methods).
The relative gene expression levels for the five tissues are shown adjacent to the phylogenetic tree in Figure 5B. Although they are from different species and data types, gene expression levels displayed regular features along phylogenetic lines. Most genes in clade B showed expression at low abundance, but some showed increased expression in floral organs, which were most obvious in A. thaliana and V. vinifera. The increased expression in flowers might come from the anthers, as confirmed experimentally in other work (
DNA sequences -1000 bp upstream (5′) of the ATG initiation codon were used to search for conserved motifs and annotated cis-elements. When we queried the annotated motif databases PLACE or PLANTCARE, large numbers of cis-elements were found, which are annotated as responsive to light (G-box, ATCT-motif, AE-box, ACE, ATCC-motif, GAG-motif, and GARE-motif), hormones (ABRE, ERE, AuxRR-core, GARE-motif, and P-box), and elicitors (EIRE and Box-w1). This matches the general defensive function of the PKS gene family. In general, the numbers of cis-elements present in upstream sequences were positively correlated with gene expression levels. However, these cis-elements always overlapped with each other, or appeared to be distributed randomly throughout the sequences, making it difficult to find meaningful patterns.
We then attempted to discover conserved motifs by MEME suites, using all of the upstream sequences or within each lineage. We first ran the MEME program using all of the upstream sequences, but no conserved motifs were found. This might be because of the high divergence between clades A and B, so we analyzed each clade separately. As a result, the most conserved motifs in clades A and B were identified. As shown in Figure 5C, the most conserved motif of all clades distributed from -300 to -100 bp upstream (5′) of the ATG start codon. Except for the missing values, genes expressed in at least one tissue all had a complete motif that was suitably located (Figures 5B,C). As expected, the most conserved motif of clade A genes embraced adjacent MYB and bHLH binding sites, two transcription factors of the MBW transcriptional regulatory complex, which have been experimentally verified in several species (
Discussion
Polyketide synthase enzymes use products of primary metabolism to synthetize chemical molecules participating in defense and reproduction. The expansion and stabilization of the PKS III gene family was important for the history of organismal adaptation and evolution. Early studies deduced the origins of the flavonoid pathway leading by the CHS enzymes through chemical substance analysis in extant species, and the time scale roughly dated back to the formation of the mosses (
Successive expansions of this multigene family in extant seed plants are thought to have depended on both genome-wide duplications and small scale duplication events. In the ML (Figure 4A) and BI (Figure 5A) trees of land plant PKS III, basal angiosperms and core eudicots formed two clades, clades A and B. Previous work has distinguished the ASCHSLE group of sequences as a monophyletic clade (
During the evolutionary process, extant chromosomal distribution and retention bias of PKSs often adapt to their modes of action whatever the duplication patterns are. In clade B, natural selection conserved PKS III genes with scattered chromosomal distributions. This may tend to support the “gene balance hypothesis,” as genes with a tendency to interact with each other were less likely to retain in tandem (
Following gene duplications, the diversification of seed plant PKS III can be observed in both enzyme structures and gene regulatory elements (
In addition to the protein sequences, gene expression patterns and potential cis-elements in upstream regulation regions also showed significant diversification. Anther-specific expression CHS-like proteins, which are embedded in clade B, have the highest expression levels in flowers (Figure 5B). The presence of a putative AP2 binding element implies that genes in this clade can respond to endogenous floral developmental signals. In different branches within clade A, expression patterns showed various tissue-specific patterns among closely related species with different biological effects. Leguminosae genes in branch a1 had root-specific expression patterns, and may be related to nitrogen fixation. Distribution of cis-elements in the different lineages indicated auxiliary trans-regulatory factors, except for MYB and bHLH transcription factors (Figure 5D). WRKY binding elements are present in genes from branches a1, a2, and a3, and ERF or DREB binding elements in genes in branches a4 and a5. Genes in a2, a3, a4, and a5 are expressed universally in vegetative or other above-ground tissues, suggesting their induction by broad acting hormones or external environmental signals. AP2, ERF, and EREBP are three subfamilies in AP2/EREBP superfamily, which occupies large proportion of plant transcriptional factors, and proteins in this superfamily widely participate in growth and response regulations (
Since changes in protein sequence and expression patterns have been found under regular diversification, they appear to be connected (Figures 4 and 5). PKS III in the phylogenetic tree can be roughly classified into two types: the proteins Metr007723-Metr058470-Metr007740 and Soly098090-Sotu043464-Sotu043447-Sotu022254-Sotu022255-Sotu022258 in clade A all have properties like longer branch lengths, loss of cis-binding sites, and inability to transcript. The other type, which comprises the majority of members with full functionality, sequence accuracy, motif integrity, and a high level of transcript in at least one tissue. It seemed that protein activity and gene expression levels are two synergetic aspects in the evolution of a multigene family. In the case of PKS III enzymes, the three sites (164C, 303H, and 336N) which make up the catalytic triad (Figure 4B), and the predicted MYB binding elements -300 to -100 bases upstream of the initiation codon (Figure 5C), are extremely conserved in all PKS III genes in all species, which implies the critical importance of these components. The stationary combination of the catalytic triad and cis-elements are outcomes of a long-term historical evolutionary process. Additionally, duplication events that occurred in the recent ancestors of certain taxa caused gene redundancy within a relatively short period, leading to relaxed selection pressures on some copies. Considering this, slight alterations in sites adjacent to the key region or cis-elements can be more flexible. This might be the reason for the formation of lineage specific patterns of the catalytic reaction and tissue-related expression.
Because of the different metabolic flows in various cell types, in vitro experiments cannot reflect the true catalytic reactions in vivo completely. At the same time, tissue-specific expression patterns can be influenced by trans-acting regulatory factors or even differential epigenetic modifications. These factors make it difficult to elucidate the functional diversification of individual family members. However, as a family of structural genes, in which their end effects on phenotypes result from natural selection, sequence features of structural genes are reflections of plant adaptability. We expect that the results from this study will provide a reference for further evolutionary studies or engineering programs in the field of polyketide synthases.
Statements
Author contributions
RS and LX designed the study. LX collected and analyzed the data and drafted the manuscript. ZZ, ShiZ, ShuZ, FL, and HZ helped to collect data. PL, GL, and YW helped to analyze data and draft the manuscript.
Funding
This research was supported by China Agricultural Research System (CARS-25-A-01), a Chinese 973 Program Grant (2012CB113900) and a Chinese 863 Program Grant (2012AA100100), both to RS.
Acknowledgments
The authors thank Oxford Science Editing for the professional language editing. We thank Farahnoz Khojayori for her effective advices on language and details to this manuscript.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Supplementary material
The Supplementary Material for this article can be found online at: http://journal.frontiersin.org/article/10.3389/fpls.2016.01312
FIGURE S1Chromosomal distribution of predicted polyketide synthases (PKS) III genes and fragments.
FIGURE S2Ancestral state reconstruction of eight positive sites (121, 208, 264, 265, 266, 276, 300, and 340) in clade a6’.
TABLE S1Information on the genomes and transcriptomes used in this study that were downloaded from public databases.
TABLE S2Sequences list of polyketide synthases (PKS) III genes and fragments in 25 land plant genomes. Gray filled means sequences with incomplete open reading frames (ORFs). Blue filled means sequences deleted as redundancy by cd-hit. Red filled means sequences located within the same gene cluster.
TABLE S3Results of one-ratio, two-ratio and branch-site models from PAML.
TABLE S4Expression levels of PKS III genes calculated from transcriptomes data.
Footnotes
2.^http://plants.ensembl.org/index.html
3.^http://www.arabidopsis.org/
References
1
AbeT.MoritaH.NomaH.KohnoT.NoguchiH.AbeI. (2007). Structure function analysis of benzalacetone synthase from Rheum palmatum.Bioorg. Med. Chem. Lett.173161–3166. 10.1016/j.bmcl.2007.03.029
2
AdamsK. L.WendelJ. F. (2005). Polyploidy and genome evolution in plants.Curr. Opin. Plant Biol.8135–141. 10.1016/j.pbi.2005.01.001
3
AltschulS. F.GishW.MillerW.MyersE. W.LipmanD. J. (1990). Basic local alignment search tool.J. Mol. Biol.215403–410. 10.1016/S0022-2836(05)80360-2
4
AtanassovI.RussinovaE.AntonovL.AtanassovA. (1998). Expression of an anther-specific chalcone synthase-like gene is correlated with uninucleate microspore development in Nicotiana sylvestris.Plant Mol. Biol.381169–1178. 10.1023/A:1006074508779
5
AustinM. B.NoelJ. P. (2003). The chalcone synthase superfamily of type III polyketide synthases.Nat. Prod. Rep.2079–110. 10.1039/b100917f
6
BaileyT. L.BodenM.BuskeF. A.FrithM.GrantC. E.ClementiL.et al (2009). MEME SUITE: tools for motif discovery and searching.Nucleic Acids Res.37202–208. 10.1093/nar/gkp335
7
BlancG.WolfeK. H. (2004). Widespread paleopolyploidy in model plant species inferred from age distributions of duplicate genes.Plant Cell161667–1678. 10.1105/tpc.021345
8
Castillo-DavisC. I.HartlD. L.AchazG. (2004). cis-Regulatory and protein evolution in orthologous and duplicate genes.Genome Res.141530–1536. 10.1101/gr.2662504
9
ChawS. M.ChangC. C.ChenH. L.LiW. H. (2004). Dating the monocot-dicot divergence and the origin of core eudicots using whole chloroplast genomes.J. Mol. Evol.58424–441. 10.1007/s00239-003-2564-9
10
ChengF.LiuS.WuJ.FangL.SunS.LiuB.et al (2011). BRAD, the genetics and genomics database for Brassica plants.BMC Plant Biol.11:136. 10.1186/1471-2229-11-136
11
ChengF.WuJ.WangX. (2014). Genome triplication drove the diversification of Brassica plants.Hortic. Res.11–8. 10.1038/hortres.2014.24
12
ColpittsC. C.KimS. S.PosehnS. E.JepsonC.KimS. Y.WiedemannG.et al (2011). PpASCL, a moss ortholog of anther-specific chalcone synthase-like enzymes, is a hydroxyalkylpyrone synthase involved in an evolutionarily conserved sporopollenin biosynthesis pathway.New Phytol.192855–868. 10.1111/j.1469-8137.2011.03858.x
13
CrooksG. E.HonG.ChandoniaJ. M.BrennerS. E. (2004). WebLogo: a sequence logo generator.Genome Res.141188–1190. 10.1101/gr/849004
14
De BodtS.MaereS.Van de PeerY. (2005). Genome duplication and the origin of angiosperms.Trends Ecol. Evol.20591–597. 10.1016/j.tree.2005.07.008
15
DengX.BashandyH.AinasojaM.KontturiJ.PietiainenM.LaitinenR. A.et al (2013). Functional diversification of duplicated chalcone synthase genes in anthocyanin biosynthesis of Gerbera hybrida.New Phytol.20131–15. 10.1111/nph.12610
16
DeyN.SarkarS.AcharyaS.MaitiI. B. (2015). Synthetic promoters in planta.Planta2421077–1094. 10.1007/s00425-015-2377-2
17
DhawaleS.SoucietG.KuhnD. N. (1989). Increase of chalcone synthase mRNA in pathogen inoculated soybeans with race-specific resistance is different in leaves and roots.Plant Physiol.91911–916. 10.1104/pp.91.3.911
18
DurbinM. L.McCaigB.CleggM. T. (2000). Molecular evolution of the chalcone synthase multigene family in the morning glory genome.Plant Mol. Biol.4279–92. 10.1023/A:1006375904820
19
FerrerJ.JezJ. M.BowmanM. E.DixonR. A.NoelJ. P. (1999). Structure of chalcone synthase and the molecular basis of plant polyketide biosynthesis.Nat. Struct. Biol.6775–784. 10.1038/11553
20
Flores-SanchezI. J.VerpoorteR. (2009). Plant polyketide synthases: a fascinating group of enzymes.Plant Physiol. Biochem.47167–174. 10.1016/j.plaphy.2008.11.005
21
Franco-ZorrillaJ. M.López-VidrieroI.CarrascoJ. L.GodoyM.VeraP.SolanoR. (2014). DNA-binding specificities of plant transcription factors and their potential to define target genes.Proc. Natl. Acad. Sci. U.S.A.1112367–2372. 10.1073/pnas.1316278111
22
FreelingM. (2009). Bias in plant gene content following different sorts of duplication: tandem, whole-genome, segmental, or by transposition.Annu. Rev. Plant Biol.60433–453. 10.1146/annurev.arplant.043008.092122
23
GoodsteinD. M.ShuS.HowsonR.NeupaneR.HayesR. D.FazoJ.et al (2011). Phytozome: a comparative platform for green plant genomics.Nucleic Acids Res.401178–1186. 10.1093/nar/gkr944
24
HatayamaM.OnoE.Yonekura-SakakibaraK.TanakaY.NishinoT.NakayamaT. (2006). Biochemical characterization and mutational studies of a chalcone synthase from yellow snapdragon (Antirrhinum majus) flowers.Plant Biotechnol.23373–378. 10.5511/plantbiotechnology.23.373
25
HelariuttaY.KotilanenM.ElomaaP.KalkkinenN.BremerK.TeeriT. H.et al (1996). Duplication and functional divergence in the chalcone synthase gene family of Asteraceae: evolution with substrate change and catalytic simplification.Proc. Natl. Acad. Sci. U.S.A.939033–9038. 10.1073/pnas.93.17.9033
26
HigoK.UgawaY.IwamotoM.KorenagaT. (1999). Plant cis-acting regulatory DNA elements (PLACE) database: 1999.Nucleis Acids Res.27297–300. 10.1093/nar/27.1.297
27
HopwoodD. A.ShermanD. H. (1990). Molecular genetics of polyketids and its comparison to fatty acid biosynthesis.Annu. Rev. Genet.2437–66. 10.1146/annurev.ge.24.120190.000345
28
HuelsenbeckJ. P.RonquistF. (2001). MRBAYES: Bayesian inference of phylogenetic trees.Bioinformatics17754–755. 10.1093/bioinformatics/17.8.754
29
IrizarryR. A.BolstadB. M.CollinF.CopeL. M.HobbsB.SpeedT. P. (2003). Summaries of Affymetrix GeneChip probe level data.Nucleic Acids Res.311–8. 10.1093/nar/gng015
30
JezJ. M.AustinM. B.FerrerJ.BowmanM. E.SchroderJ.NoelJ. P. (2000). Structural control of polyketide formation in plant-specific polyketide synthases.Chem. Biol.7919–930. 10.1016/S1074-5521(00)00041-7
31
JiangC.KimS. Y.SuhD. (2008). Divergent evolution of the thiolase superfamily and chalcone synthase family.Mol. Phylogenet. Evol.49691–701. 10.1016/j.ympev.2008.09.002
32
JiaoY.WickettN. J.AyyampalayamS.ChanderbaliA. S.LandherrL.RalphP. E.et al (2011). Ancestral polyploidy in seed plants and angiosperms.Nature47397–100. 10.1038/nature09916
33
Johzuka-HisatomiY.HoshinoA.MoriY.HabuY.IidaS. (1999). Characterization of the chalcone synthase genes expressed in flowers of the common and Japanese morning glories.Genes Genet. Syst.74141–147. 10.1266/ggs.74.141
34
JunghansH.DalkinK.DixonR. A. (1993). Stress responses in alfalfa (Medicago sativa L.). 15. Characterization and expression patterns of members of a subset of the chalcone synthase multigene family.Plant Mol. Biol.22239–253. 10.1007/BF00014932
35
KagaleS.RobinsonS. J.NixonJ.XiaoR.HuebertT.CondieJ.et al (2014). Polyploid evolution of the Brassicaceae during the Cenozoic era.Plant Cell262777–2791. 10.1105/tpc.114.126391
36
KatohK.MisawaK.KumaK.-I.MiyataT. (2002). MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform.Nucleic Acids Res.303059–3066. 10.1093/nar/gkf436
37
KerseyP. J.AllenJ. E.ChristensenM.DavisP.FalinL. J.GrabmuellerC.et al (2014). Ensembl genomes 2013: scaling up access to genome-wide data.Nucleic Acids Res.42546–552. 10.1093/nar/gkt979
38
KinsellaR. J.KähäriA.HaiderS.ZamoraJ.ProctorG.SpudichG.et al (2011). Ensembl BioMarts: a hub for data retrieval across taxonomic space.Database20111–9. 10.1093/database/bar030
39
KoduriP. K. H.GordonG. S.BarkerE. I.ColpittsC. C.AshtonN. W.SuhD. (2009). Genome-wide analysis of the chalcone synthase superfamily genes of Physcomitrella patens.Plant Mol. Biol.72247–263. 10.1007/s11103-009-9565-z
40
KoesR. E.SpeltC. E.MolJ. N. M. (1989). The chalcone synthase multigene family of Petunia hybrida (V30): differential, light-regulated expression during flower development and UV light induction.Plant Mol. Biol.12213–225. 10.1007/BF00020506
41
KreuzalerF.RaggH.HellerW.TeschR.WittI.HammerD.et al (1979). Flavanone synthase from Petroselinum hortense.Eur. J. Biochem.9989–96. 10.1111/j.1432-1033.1979.tb13235.x
42
LameschP.BerardiniT. Z.LiD.SwarbreckD.WilksC.SasidharanR.et al (2012). The Arabidopsis information resource (TAIR): improved gene annotation and new tools.Nucleic Acids Res.401202–1210. 10.1093/nar/gkr1090
43
LescotM.DéhaisP.ThijsG.MarchalK.MoreauY.Van de PeerY.et al (2002). PlantCARE, a database of plant cis-acting regulatory elements and a portal to tools for in silico analysis of promoter sequences.Nucleic Acids Res.30325–327. 10.1093/nar/30.1.325
44
LiW.GodzikA. (2006). Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences.Bioinformatics221658–1659. 10.1093/bioinformatics/btl158
45
LilaM. A. (2004). Anthocyanins and human health: an in vitro investigative approach.J. Biomed. Biotechnol.5306–313. 10.1155/S111072430440401X
46
LukačinR.SchreinerS.MaternU. (2001). Transformation of acridone synthase to chalcone synthase.FEBS Lett.508413–417. 10.1016/S0014-5793(01)03061-7
47
MaereS.De BodtS.RaesJ.CasneufT.Van MontaguM.KuiperM.et al (2005). Modeling gene and genome duplications in eukaryotes.Proc. Natl. Acad. Sci. U.S.A.1025454–5459. 10.1073/pnas.0501102102
48
MooreM. J.HassanN.GitzendannerM. A.BruennR. A.CroleyM.VandeventerA.et al (2011). Phylogenetic analysis of the plastid inverted repeat for 244 species: insights into deeper-level angiosperm relationships from a long, slowly evolving sequence region.Int. J. Plant Sci.172541–558. 10.1086/658923
49
MooreR. C.PuruggananM. D. (2005). The evolutionary dynamics of plant duplicate genes.Curr. Opin. Plant Biol.8122–128. 10.1016/j.pbi.2004.12.001
50
OlsenJ. L.RouzéP.VerhelstB.LinY. C.BayerT.CollenJ.et al (2016). The genome of the seagrass Zostera marina reveals angiosperm adaptation to the sea.Nature530331–335. 10.1038/nature16548
51
ReimoldU.KrögerM.KreuzalerF.HahlbrockK. (1983). Coding and 3’ non-coding nucleotide sequence of chalcone synthase mRNA and assignment of amino acid sequence of the enzyme.EMBO J.21801–1805.
52
RiechmannJ. L. (2000). Arabidopsis transcription factors: genome-wide comparative analysis among eukaryotes.Science2902105–2110. 10.1126/science.290.5499.2105
53
SandersonM. J.ThorneJ. L.WikströmN.BremerK. (2004). Molecular evidence on plant divergence times.Am. J. Bot.911656–1665. 10.3732/ajb.91.10.1656
54
SchröderG.SchröderJ. (1992). A single change of histidine to glutamine alters the substrate preference of a stilbene synthase.J. Biol. Chem.26720558–20560.
55
ShangY.VenailJ.MackayS.BaileyP. C.SchwinnK. E.JamesonP. E.et al (2011). The molecular basis for venation patterning of pigmentation and its effect on pollinator attraction in flowers of Antirrhinum.New Phytol.189602–615. 10.1111/j.1469-8137.2010.03498.x
56
ShomuraY.TorayamaI.SuhD. Y.XiangT.KitaA.SankawaU.et al (2005). Crystal structure of stilbene synthase from Arachis hypogaea.Proteins60803–806. 10.1002/prot.20584
57
SoltisD. E.BellC. D.KimS.SoltisP. S. (2008). Origin and early evolution of angiosperms.Ann. N. Y. Acad. Sci.11333–25. 10.1196/annals.1438.005
58
SommerH.SaedlerH. (1986). Structure of the chalcone synthase gene of Antirrhinum majus.Mol. Gen. Genet.202429–434. 10.1007/BF00333273
59
StaffordH. A. (1991). Flavonoid evolution: an enzymic approach.Plant Physiol.96680–685. 10.1104/pp.96.3.680
60
StamatakisA.HooverP.RougemontJ. (2008). A rapid bootstrap algorithm for the RAxML web servers.Syst. Biol.57758–771. 10.1080/10635150802429642
61
SteynW. J.WandS. J. E.HolcroftD. M.JacobsG. (2002). Anthocyanins in vegetative tissues: a proposed unified function in photoprotection.New Phytol.155349–361. 10.1046/j.1469-8137.2002.00482.x
62
SuhD.FukumaK.KagamiJ.YamazakiY.ShibuyaM.EbizukaY.et al (2000). Identification of amino acid residues important in the cyclization reactions of chalcone and stilbene synthases.Biochem. J.350229–235. 10.1042/0264-6021:3500229
63
SuyamaM.TorrentsD.BorkP. (2006). PAL2NAL: robust conversion of protein sequence alignments into the corresponding codon alignments.Nucleic Acids Res.34609–612. 10.1093/nar/gkl315
64
TamuraK.PetersonD.PetersonN.StecherG.NeiM.KumarS. (2011). MEGA5: molecular evolutionary genetics analysis using maximum likelihood, evolutionary distance, and maximum parsimony methods.Mol. Biol. Evol.282731–2739. 10.1093/molbev/msr121
65
ThomassetS.TellerN.CaiH.MarkoD.BerryD. P.StewardW. P.et al (2009). Do anthocyanins and anthocyanidins, cancer chemopreventive pigments in the diet, merit development as potential drugs?Cancer Chemother. Pharmacol.64201–211. 10.1007/s00280-009-0976-y
66
TrapnellC.HendricksonD. G.SauvageauM.GoffL.RinnJ. L.PachterL. (2013). Differential analysis of gene regulation at transcript resolution with RNA-seq.Nat. Biotechnol.3146–53. 10.1038/nbt.2450
67
TrapnellC.PachterL.SalzbergS. L. (2009). TopHat: discovering splice junctions with RNA-Seq.Bioinformatics251105–1111. 10.1093/bioin-formatics/btp120
68
TropfS.KärcherB.SchröderG.SchröderJ. (1995). Reaction mechanisms of homodimeric plant polyketide synthases (stilbene and chalcone synthase).J. Biol. Chem.2707922–7928. 10.1074/jbc.270.14.7922
69
WangH.GuanS.ZhuZ.WangY.LuY. (2013). A valid strategy for precise identifications of transcription factor binding sites in combinatorial regulation using bioinformatic and experimental approaches.Plant Methods91–11. 10.1186/1746-4811-9-34
70
WickhamH. (2009). Ggplot2: Elegant Graphics for Data Analysis.New York, NY: Springer.
71
XiaX.XieZ. (2001). DAMBE: software package for data analysis in molecular biology and evolution.J. Hered.92371–373. 10.1093/jhered/92.4.371
72
XuW.DubosC.LepiniecL. (2015). Transcriptional control of flavonoid biosynthesis by MYB-bHLH-WDR complexes.Trends Plant Sci.20176–185. 10.1016/j.tplants.2014.12.001
73
YangZ. (1998). Likelihood ratio tests for detecting positive selection and application to primate lysozyme evolution.Mol. Biol. Evol.15568–573. 10.1093/oxfordjournals.molbev.a025957
74
YangZ. (2007). PAML 4: phylogenetic analysis by maximum likelihood.Mol. Biol. Evol.241586–1591. 10.1093/molbev/msm088
75
YangZ.NielsenR. (2002). Codon-substitution models for detecting molecular adaptation at individual sites along specific lineages.Mol. Biol. Evol.19908–917. 10.1093/oxfordjournals.molbev.a004148
76
ZengL.ZhangQ.SunR.KongH.ZhangN.MaH.et al (2014). Resolution of deep angiosperm phylogeny using conserved nuclear genes and estimates of early divergence times.Nat. Commun.5:4956. 10.1038/ncomms5956
77
ZhangJ. (2003). Evolution by gene duplication: an update.Trends Ecol. Evol.18292–298. 10.1016/s0169-5347(03)00033-8
78
ZhuZ.WangH.WangY.GuanS.WangF.TangJ.et al (2015). Characterization of the cis elements in the proximal promoter regions of the anthocyanin pathway genes reveals a common regulatory logic that governs pathway regulation.J. Exp. Bot.663775–3789. 10.1093/jxb/erv173
Summary
Keywords
PKS III multigene family, CHS, STS, phylogenetic reconstruction, functional diversification, gene expression, cis-elements
Citation
Xie L, Liu P, Zhu Z, Zhang S, Zhang S, Li F, Zhang H, Li G, Wei Y and Sun R (2016) Phylogeny and Expression Analyses Reveal Important Roles for Plant PKS III Family during the Conquest of Land by Plants and Angiosperm Diversification. Front. Plant Sci. 7:1312. doi: 10.3389/fpls.2016.01312
Received
15 June 2016
Accepted
16 August 2016
Published
30 August 2016
Volume
7 - 2016
Edited by
Alessio Mengoni, University of Florence, Italy
Reviewed by
Yanbin Yin, University of Georgia, USA; Juan Caballero, Autonomous University of Queretaro, Mexico
Updates

Check for updates
Copyright
© 2016 Xie, Liu, Zhu, Zhang, Zhang, Li, Zhang, Li, Wei and Sun.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Rifei Sun, sunrifei@caas.cn Lulu Xie, xielulu_1003@163.com
This article was submitted to Bioinformatics and Computational Biology, a section of the journal Frontiers in Plant Science
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.