Abstract
In synthetic biology, precise control over protein expression is required in order to construct functional biological systems. A core principle of the synthetic biology approach is a model-guided design and based on the biological understanding of the process, models of prokaryotic protein production have been described. Translation initiation rate is a rate-limiting step in protein production from mRNA and is dependent on the sequence of the 5′-untranslated region and the start of the coding sequence. Translation rate calculators are programs that estimate protein translation rates based on the sequence of these regions of an mRNA, and as protein expression is proportional to the rate of translation initiation, such calculators have been shown to give good approximations of protein expression levels. In this review, three currently available translation rate calculators developed for synthetic biology are considered, with limitations and possible future progress discussed.
Introduction
Synthetic biology is a recently emerged field concerned with engineering complex living systems by assembling individually characterized biological parts in novel combinations. The discipline arose from the discovery of the mathematical logic of gene pairings, as well as from advances made in genetic engineering and recombinant DNA technology (Andrianantoandro et al., ). The development of de novo DNA synthesis, protein engineering, and the designs of artificial gene networks have greatly contributed to the field’s advancement (Heinemann and Panke, ). Synthetic biology seeks to determine the behavior of organisms and their parts, and then to modify and combine them into complete specific tasks. The field is based on the engineering principles of design and fabrication and focuses on the concept of standardized parts (Serrano, ). Precise control over the levels of protein expression is an important requirement for the robust operation of complex synthetic circuits built from many parts.
Despite improving characterization and assembly methods, cycles of design, fabrication, and testing in synthetic biology can be slow. Production of circuits with desired properties can require several rounds of testing and modifying, each time editing imperfect parts by mutation or identifying alternatives. Directed evolution has been shown to provide a short cut through this phase (Yokobayashi et al., ), but is complicated by the additional work needed to couple networks to selective pressures. Instead, use of predictive mathematical modeling to rationally guide the design of gene networks can greatly improve design cycles to accelerate advances in synthetic biology (Ellis et al., ).
Levels of protein expression are affected by both the transcription and translation rates but early genetic engineering approaches usually focused solely on transcription (Lipniacki et al., ). The transcription rate’s heavy dependence on the promoter strength and the relative ease of estimating binding affinity of RNA polymerase helped its early popularity (Alper et al., ). However, to gain more accurate and efficient control over protein expression translation rates must also be considered.
Translation initiation is one of the major steps in translation and plays a large role in determining the overall translation rate (Laursen et al., ; Kudla et al., ). While other factors such as the elongation rate and the termination rate also significantly affect translation (Lithwick and Margalit, ; Mehra and Hatzimanikatis, ), the initiation rate is of particular interest for synthetic biology as it provides a means to tune protein production over many orders of magnitude by only varying the relatively short RNA sequences at the start of mRNAs that determine the initiation rate. Modeling this step is therefore hugely valuable for designing biological systems.
Ribosome-mRNA Interactions at Initiation
Modeling translation initiation requires an accurate understanding of ribosome interactions with the mRNA 5′-untranslated region (5′-UTR) ahead of protein synthesis. When a ribosome docks with an mRNA to begin translation, only the 30S subunit of the ribosome binds the 5′-UTR. The 16S ribosomal RNA (rRNA) within this subunit binds to a sequence in the 5′-UTR known as the ribosome binding site (RBS), while the initiator transfer RNA (fMET-tRNA) binds to the start codon (AUG) of the protein-coding sequence. The spacing between these sites on the 5′-UTR is important, with a distance of 6–8 nucleotides between the RBS and AUG being optimal (Vellanoweth and Rabinowitz, ). Within the RBS, the 3′ end of the 16S rRNA subunit is complementary to a short sequence named the Shine–Dalgarno (SD) sequence.
The factors that influence the rate of translation initiation can be grouped into three categories (Figure 1). Firstly, the global folding and unfolding of transcribed mRNAs, whose secondary structures can hinder the binding of the ribosome: during translation initiation the transcribed mRNA folds in and out of the secondary structures, which may interfere with ribosome binding (de Smit and van Duin, ). Secondly, the regional folding and unfolding of nucleotides in the RBS region: the ribosome docking site (RDS), a sequence roughly 30 nucleotides around the start codon, must be unfolded and exposed for the ribosome recognition sequence to bind. Lastly, there is the efficiency of ribosome binding itself, which is determined by the binding affinities between the SD sequence and the complementary 16S rRNA anti-SD sequence (Na et al., ).
Figure 1
Ribosome Binding Models and Calculators
Three different translation rate calculators have been developed. The first, released in 2009 and updated in 2011 is the RBS Calculator (Salis et al.,
All the translation rate calculators use a proportional scale for their estimated translation initiation rate rather than any definitive units. For example, a predicted output of 500 should produce 10 times more protein than an output of 50, if all other effects are equal. The relative scales are not the same between the different calculators. The three calculators have been initially designed to predict translation initiation rates and estimate protein expression from a given mRNA sequence. This feature is known as “reverse-engineering” as the sequence has been pre-defined and a property of this sequence is calculated. Each calculator also incorporates a “forward-engineering” feature, where a 5′-UTR sequence (if required) and coding sequence are inputted with a desired translation initiation rate. An algorithm is then used to generate a suitable RBS sequence to go between the 5′-UTR sequence and coding sequence to give the desired rate. To accomplish this, a random RBS seed sequence is created and varied until the translation rate matches the desired rate. Each calculator has its own algorithm for efficiently generating and selecting suitable sequences from the combinatorially huge number of possibilities.
The RBS Calculator
Predicting the rate of translation initiation for different 5′-UTR sequences requires a biophysical model of the process. To do this, Salis et al., developed an equilibrium statistical thermodynamic model using previously characterized free energies of key molecular interactions involved in translation initiation (Salis et al.,
The translation initiation rate relates to ΔGtotal according to the exponential relationship where r is the translation initiation rate and β is the Boltzmann factor for the system. Similarly, the total protein expression E is proportional to the translation initiation rate r by a constant k, which accounts for ribosomal and mRNA interactions independent of the 5′-UTR sequence and parameters unaffected by translation (Salis,
The currently available RBS Calculator (Version 1.1) released in 2011, uses the ViennaRNA suite (Gruber et al.,
The Salis Lab RBS Calculator is run from a web-based server and can be found at https://salis.psu.edu/software/. The results page for reverse-engineering shows the entire inputted mRNA sequence, highlighting any possible start codons. For each possible start codon the calculated translation initiation rate is given, followed by the ΔGtotal and all component ΔG values. Also, as an advantage over other software, an estimation of confidence is given. A green result indicates relatively high confidence, while various error codes indicate potential inaccuracies. For example, there may be multiple closely spaced or overlapping start codons that could cause unpredictable ribosome–ribosome interactions.
For forward-engineering, the thermodynamic model is combined with a stochastic optimization method to design synthetic sequence. A particular translation rate may be chosen or the “maximize” function selected to give the highest possible translation rate for the given coding sequence. By accurately considering the context effects the software can design synthetic RBSs far stronger than previously possible by manual design or by copying strong natural sequences. A benefit of the forward-engineering mode is the ability to only design synthetic sequences that always satisfy the model’s assumptions, which leads to higher predictive accuracy (Salis,
The UTR Designer
Seo et al. (
The output ΔGUTR term, equal to ΔGfinal − ΔGintial, is used to estimate relative translation rate (r) using the Salis et al. (
The UTR Designer can be found at http://sbi.postech.ac.kr/rbs. The reverse-engineering results page gives the imputed sequences with position of each possible start codon and the predicted core RBS sequence highlighted. The calculator indicates the standby location where the ribosome may bind to contribute to the ΔGindirect term and the nucleotide spacing between the start codon and the RBS. The forward-engineering mode designs an optimal sequence to achieve a given expression level. Unlike other calculators, however, the UTR Designer can also alter the codons of the coding sequence in order to reduce secondary structures and improve translation rate when the variations in 5’-UTR cannot satisfy the desired expression levels. It also features a UTR Library Designer that designs degenerate sequences to give translation rates across a specified range.
The RBS Designer
A third translation rate calculator, the RBS Designer, was developed by Na and Lee (
Possible mRNA secondary structures are next considered and the ΔG of each are determined using UNAFold. For each structure, an RDS exposure probability (the probability that the RDS will be accessible to the ribosome) is determined by calculating individual nucleotide unpairing probabilities for each nucleotide within the RDS. Individual probabilities of each structure forming are calculated then multiplied by each structure’s individual RDS exposure probability. All terms are then summed to give the total exposure probability for the RDS.
Ribosome binding is modeled with ordinary differential equations and the steady state is assumed. The probability of a given mRNA being bound to a ribosome (translation efficiency) is then calculated from the total RDS exposure probability and ribosome binding affinity, with other parameters taken from the literature. This translation efficiency is approximately proportional to protein production level (de Smit and van Duin,
The RBS Designer must be downloaded and run locally. Installation instructions and relevant links can be found at http://rbs.kaist.ac.kr/. A notable difference compared to other software is the requirement for at least 300 nucleotides of mRNA sequence. This allows better prediction of secondary structures by considering long-range interaction but is more computationally intensive. The RBS Designer can estimate translation rate for a given mRNA sequence in reverse-engineering mode and in forward-engineering mode it uses a genetic algorithm to vary and select optimal nucleotide sequence, designing a 5′-UTR sequence to give a specified translation rate. Unlike the other calculators, however, the program lacks any library design features.
Discussion
Each of the currently available calculators show similarly accurate predictions compared with experimental data in their respective publications. The RBS Calculator was tested with 29 synthetic RBSs and predictions correlated well with experimental results with R2 = 0.84 (Salis,
There are several areas of improvement for RBS Calculators. Salis acknowledges several limitations of his model and these simplifications are also present in the other models (Salis,
Accounting for these limitations and refining the parameters of the models will lead to improvements in accuracy. There is also room for widening applicability. All calculators were designed for use with E. coli and acknowledge that they would not be as accurate for other organisms (though models should hold for similar Gram-negative bacteria). With further testing, models could be adapted to include Gram-positive bacteria. These cells exhibit differences in translational machinery with a major difference in optimum spacing requirements between the SD sequence and the start codon, which would significantly affect the ΔGspacing terms (Vellanoweth and Rabinowitz,
Conclusion
Ribosome binding site calculators are increasingly valuable tools for synthetic biologists. They allow translation strengths to be estimated from the mRNA sequence so genetic designs can be better informed. Three calculators have been created with two (RBS Calculator and UTR Designer) using a thermodynamic model and run from online servers, and a third (RBS Designer) using a steady-state kinetic model with a downloadable application (Table 1). All of the models seek to simplify the complex natural phenomenon of translation and will continue to be improved and refined to increase predictive accuracy.
Table 1
| Feature | RBS Calculator (Salis et al., | UTR Designer (Seo et al., | RBS Designer (Na and Lee, |
|---|---|---|---|
| Location | Online | Online | Locally run |
| Forward and reverse engineering | Yes | Yes | Yes |
| RBS library design | Yes | Yes | No |
| External software used for RNA free energy calculations | ViennaRNA (v1.1) | NUPACK | UNAfold |
| NUPACK (v1.0) | |||
| Unique selling points | The most frequently updated model, gives indications of confidence | Can edit codon usage to limit unwanted secondary structures | Considers very long-range interactions within the mRNA |
Key differences between the calculators.
Statements
Acknowledgments
Work at CSYNBI is supported by the UK Engineering and Physical Sciences Research Council (EPSRC), and Benjamin Reeve is additionally co-funded by TMO Renewables Ltd.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
References
1
AlperH.FischerC.NevoigtE.StephanopoulosG. (2005). Tuning genetic control through promoter engineering. Proc. Natl. Acad. Sci. U.S.A.102, 12678–12683.10.1073/pnas.0504604102
2
AndrianantoandroE.BasuS.KarigD. K.WeissR. (2006). Synthetic biology: new engineering rules for an emerging discipline. Mol. Syst. Biol.2, 2006.0028.10.1038/msb4100073
3
de SmitM. H.van DuinJ. (1990). Secondary structure of the ribosome binding site determines translational efficiency: a quantitative analysis. Proc. Natl. Acad. Sci. U.S.A.87, 7668–7672.10.1073/pnas.87.19.7668
4
DirksR. M.BoisJ. S.SchaefferJ. M.WinfreeE.PierceN. A. (2007). Thermodynamic analysis of interacting nucleic acid strands. SIAM Rev.49, 65–88.10.1137/060651100
5
EllisT.WangX.CollinsJ. J. (2009). Diversity-based, model-guided construction of synthetic gene networks with predicted functions. Nat. Biotechnol.27, 465–471.10.1038/nbt.1536
6
Espah BorujeniA.ChannarasappaA. S.SalisH. M. (2013). Translation rate is controlled by coupled trade-offs between site accessibility, selective RNA unfolding and sliding at upstream standby sites. Nucleic Acids Res.10.1093/nar/gkt1139
7
GruberA. R.LorenzR.BernhartS. H.NeubockR.HofackerI. L. (2008). The Vienna RNA websuite. Nucleic Acids Res.36, W70–W74.10.1093/nar/gkn188
8
HeinemannM.PankeS. (2006). Synthetic biology – putting engineering into biology. Bioinformatics22, 2790–2799.10.1093/bioinformatics/btl469
9
KudlaG.MurrayA. W.TollerveyD.PlotkinJ. B. (2009). Coding-sequence determinants of gene expression in Escherichia coli. Science324, 255–258.10.1126/science.1170160
10
LaursenB. S.SørensenH. P.MortensenK. K.Sperling-PetersenH. U. (2005). Initiation of protein synthesis in bacteria. Microbiol. Mol. Biol. Rev.69, 101–123.10.1128/MMBR.69.1.101-123.2005
11
LipniackiT.PaszekP.Marciniak-CzochraA.BrasierA. R.KimmelM. (2006). Transcriptional stochasticity in gene expression. J. Theor. Biol.238, 348–367.10.1016/j.jtbi.2005.05.032
12
LithwickG.MargalitH. (2003). Hierarchy of sequence-dependent features associated with prokaryotic translation. Genome Res.13, 2665–2673.10.1101/gr.1485203
13
MarkhamN. R.ZukerM. (2008). UNAFold: software for nucleic acid folding and hybridization. Methods Mol. Biol.453, 3–31.10.1007/978-1-60327-429-6_1
14
MehraA.HatzimanikatisV. (2006). An algorithmic framework for genome-wide modeling and analysis of translation networks. Biophys. J.90, 1136–1146.10.1529/biophysj.105.062521
15
NaD.LeeS.LeeD. (2010). Mathematical modeling of translation initiation for the estimation of its efficiency to computationally design mRNA sequences with desired expression levels in prokaryotes. BMC Syst. Biol.4:71.10.1186/1752-0509-4-71
16
NaD.LeeD. (2010). RBSDesigner: software for designing synthetic ribosome binding sites that yields a desired level of protein expression. Bioinformatics26, 2633–2634.10.1093/bioinformatics/btq458
17
ParkY. S.SeoS. W.HwangS.ChuH. S.AhnJ. H.KimT. W.et al (2007). Design of 5’-untranslated region variants for tunable expression in Escherichia coli. Biochem. Biophys. Res. Commun.356, 136–141.10.1016/j.bbrc.2007.02.127
18
QuX.LancasterL.NollerH. F.BustamanteC.TinocoI.Jr. (2012). Ribosomal protein S1 unwinds double-stranded RNA in multiple steps. Proc. Natl. Acad. Sci. U.S.A.109, 14458–14463.10.1073/pnas.1208950109
19
SalisH. M. (2011). The ribosome binding site calculator. Meth. Enzymol.498, 19–42.10.1016/B978-0-12-385120-8.00002-4
20
SalisH. M.MirskyE. A.VoigtC. A. (2009). Automated design of synthetic ribosome binding sites to control protein expression. Nat. Biotechnol.27, 946–950.10.1038/nbt.1568
21
SeoS. W.YangJ.JungG. Y. (2009). Quantitative correlation between mRNA secondary structure around the region downstream of the initiation codon and translational efficiency in Escherichia coli. Biotechnol. Bioeng.104, 611–616.10.1002/bit.22431
22
SeoS. W.YangJ.-S.KimI.YangJ.MinB. E.KimS.et al (2013). Predictive design of mRNA translation initiation region to control prokaryotic translation efficiency. Metab. Eng.15, 67–74.10.1016/j.ymben.2012.10.006
23
SerranoL. (2007). Synthetic biology: promises and challenges. Mol. Syst. Biol.3, 158.10.1038/msb4100202
24
VellanowethR. L.RabinowitzJ. C. (1992). The influence of ribosome-binding-site elements on translational efficiency in Bacillus subtilis and Escherichia coli in vivo. Mol. Microbiol.6, 1105–1114.10.1111/j.1365-2958.1992.tb01548.x
25
YokobayashiY.WeissR.ArnoldF. H. (2002). Directed evolution of a genetic circuit. Proc. Natl. Acad. Sci. U.S.A.99, 16587–16591.10.1073/pnas.252535999
26
ZadehJ. N.SteenbergC. D.BoisJ. S.WolfeB. R.PierceM. B.KhanA. R.et al (2011). NUPACK: analysis and design of nucleic acid systems. J. Comput. Chem.32, 170–173.10.1002/jcc.21596
Summary
Keywords
synthetic biology, translation rate, translation efficiency, ribosome binding site, 5′-untranslated region, RBS Calculator, RBS Designer, UTR Designer
Citation
Reeve B, Hargest T, Gilbert C and Ellis T (2014) Predicting Translation Initiation Rates for Designing Synthetic Biology. Front. Bioeng. Biotechnol. 2:1. doi: 10.3389/fbioe.2014.00001
Received
14 November 2013
Accepted
06 January 2014
Published
20 January 2014
Volume
2 - 2014
Edited by
Lixin Zhang, IMCAS, China
Reviewed by
Weiwen Zhang, Tianjin University, China; Dong-Yup Lee, National University of Singapore, Singapore
Copyright
© 2014 Reeve, Hargest, Gilbert and Ellis.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Tom Ellis, Department of Bioengineering, Imperial College London, 704 Bessemer Building, South Kensington Campus, London SW7 2AZ, UK e-mail: t.ellis@imperial.ac.uk
This article was submitted to Synthetic Biology, a section of the journal Frontiers in Bioengineering and Biotechnology.
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.