MINI REVIEW article

Front. Bioeng. Biotechnol., 28 October 2020

Sec. Synthetic Biology

Volume 8 - 2020 | https://doi.org/10.3389/fbioe.2020.584178

Modeling Cell-Free Protein Synthesis Systems—Approaches and Applications

  • Institute of Biochemical Engineering, University of Stuttgart, Stuttgart, Germany

Abstract

In vitro systems are ideal setups to investigate the basic principles of biochemical reactions and subsequently the bricks of life. Cell-free protein synthesis (CFPS) systems mimic the transcription and translation processes of whole cells in a controlled environment and allow the detailed study of single components and reaction networks. In silico studies of CFPS systems help us to understand interactions and to identify limitations and bottlenecks in those systems. Black-box models laid the foundation for understanding the production and degradation dynamics of macromolecule components such as mRNA, ribosomes, and proteins. Subsequently, more sophisticated models revealed shortages in steps such as translation initiation and tRNA supply and helped to partially overcome these limitations. Currently, the scope of CFPS modeling has broadened to various applications, ranging from the screening of kinetic parameters to the stochastic analysis of liposome-encapsulated CFPS systems and the assessment of energy supply properties in combination with flux balance analysis (FBA).

Introduction

Cell-free protein synthesis (CFPS) technology has a long history in life sciences, which started with fundamental research on deducing the genetic code (). Over several decades, the system was developed stepwise into a polypeptide production machinery (). Since the early adaptations of the system for commercial use (), an increasing number of applications have emerged in the market (e.g., PURExpress, PUREfrex, PUREfrex2.0, myTXTLkit). These are based either on synthetic transcription–translation systems with a well-defined composition or on crude cell extracts that contain a more complex component assembly. While the product titer and production volume of such systems have increased from a few microliters to hundreds of liters, several limitations of the system remain. To date, and even within the best commercial systems, the protein titer with CFPS systems is orders of magnitude lower than that of in vivo whole-cell production due to resource expense and reduced longevity (; ; ). Here, purpose-driven modeling can be a crucial tool to push the boundaries forward and identify bottlenecks.

CFPS is an “open” experimental system that allows defined reaction setups, which is ideal for simulation approaches. It has emerged not only as a research tool for the processes of transcription and translation but as a biomanufacturing platform for rapidly prototyping production systems in silico and in vitro (; ). For a variety of transcriptional and translational components, kinetic parameters are known, allowing the study of their behavior. Yet, one of the most fundamental principles of modeling always restricts the approach; the accuracy of model predictions cannot exceed the granularity of the model itself. In other words, distinctions between experimental observations and simulations are likely to occur if model predictions, extrapolated data sets, or fundamental model structures do not reflect the real problem. Consequently, such discrepancies may motivate a more thorough study of the experimental problem. Hence, proper model design aims to reflect reality with sufficient granularity (e.g., should the maturation of a reporter be considered?), thereby building on a solid mixture of experimentally validated data supported by assumptions. In this regard, we provide a brief overview of the existing models for CFPS and related systems and how they are applied to specific cases. It must be stated that many models have been developed with different objectives regarding the system environment, model approach (deterministic, stochastic), and granularity. Therefore, the models are not categorized as “good” or “bad” but clustered and assessed with respect to their particular purpose. In contrast to the mini-review of , which focuses on deterministic models for CFPS, we expand the scope to adjacent fields and highlight qualitative and quantitative model characteristics.

When developing a model, it is necessary to know the components that should be considered. A CFPS system typically consists of, at least, the core components of transcription and translation: a mRNA polymerase, ribosomes, translational factors, amino acyl-tRNA synthetases, amino acids, tRNAs, an energy regeneration system, and nucleotides (). Additionally, the DNA substrate, the produced mRNA, and the product (in most cases presented here, GFP derivatives) must be considered. If a crude cell extract is used, the system becomes much more complex, as the concentration of many of the components is unknown. The modeling studies on CFPS presented in detail in section “Development and Application of CFPS Models” share the common goal of identifying key model parameters by parameter regression on experimental data. However, the complexity of the models differs. By trend, the models may be divided into four groups of different granularities (Figure 1):

FIGURE 1

“Minimum model”: Minimal models, presented in section “Identifying Bottlenecks in CFPS Systems,” take into account up to ten parameters or equations, mainly focusing on macromolecular components such as mRNA and DNA. They are the backbone for more detailed descriptions. Additionally, most of the genetic circuit models presented in section “Extending the Scope of CFPS Modeling” can be described as minimum models.

“Structured model”: Medium-scale models that introduce structured descriptions of certain aspects of the transcription-translation network. In structured models, kinetic models such as Michaelis–Menten or Hill are implemented in combination with larger ODE systems of up to 100 equations.

“Unstructured model”: Large-scale models are fine-grained. They are meant to describe the CFPS in a holistic way and comprise networks of several hundred reactions. These models typically use simple individual reactions without further structural elements in order to save computational costs (e.g., ).

“Hybrid model”: A special case are hybrid models, connecting structured models to other networks such as metabolic networks, or unstructured parts of lumped elements. Intrinsically, the approach increases the model complexity and computational costs. However, it offers an in-depth analysis of CFPS.

Development and Application of CFPS Models

Continuous models are typically used to simulate CFPS systems. They make use of ordinary differential equations (ODEs) and algebraic equations to dynamically describe model states. In a structured model, equations incorporate affinity constants and other parameters. To reduce complexity, models can be formulated following an unstructured black-box approach, considering only apparent kinetics (). The differences between the model types are dynamic. Often, models are partly structured to focus on selected segments of the reaction network with particular interest. We call these approaches “hybrid models.” The quality of CFPS mechanistic models relies heavily on the proper model structure and the correct identification of model parameters. Given the complex nature of CFPS, the precondition of independent datasets for parameter identification is challenging and may require repeated careful consideration for each regression analysis (; ). Stochastic effects play only a minor role in most of the classical CFPS modeling approaches. In liposome or droplet-based CFPS, due to small reaction volumes, low numbers of molecules may cause rendering reactions between different molecules in stochastic events. Under such conditions, a description with a discrete and stochastic model is preferable (; ).

Identifying Bottlenecks in CFPS Systems

The most straightforward way of describing in vitro expression of GFP is to consider the macromolecular components, DNA, mRNA, and proteins in a black-box model (; ; ; ). This allows fitting kinetic equations to the experimental results of GFP production, mRNA production, and mRNA degradation. RNA polymerase and ribosomes are considered as catalytic components. For simplification, we call such approaches “minimum models.” proposed a coarse-grained dynamic model consisting of four enzymatic reactions. Kinetic studies were performed for crude cytoplasmic extract from Escherichia coli to identify biosynthesis and degradation parameters. A similar granularity was chosen by to simulate and analyze the results gathered with the PURExpress system. Here, the model neglects protein degradation but covers a broader experimental range, identifying the plateau phase when the translational system expires. A comparable ODE-based model was applied to a variety of regulatory elements (promotor strength) to identify limitations in the resources of the transcription–translation system (). Here, the commercially available “myTXTLkit” was used. Despite the application of different CFPS systems, all models revealed a saturation effect in GFP production under increased DNA template concentrations. extended the minimal model to describe the expression of different fluorescence proteins under various regulatory elements. Here, limitations of current CFPS models were addressed, namely the specificity for only narrow experimental data sets, limited prediction capacity, and neglecting biophysical factors (e.g., RNA secondary structure).

With a model system of similar complexity, it was shown that limitations of CFPS may occur, which could not be mirrored by minimum models (). Using a comprehensive experimental data set of commercial E. coli CFPS, depletion of tRNAs and translation initiation were identified as limiting factors. By extending the minimum model to a structured description with additional terms for inactive mRNA states, it was possible to improve the prediction quality for the experimental data.

More fine-grained models were necessary to identify the challenging substrate limitations in silico. Such a model was introduced in our laboratory by and recently renovated (). In this hybrid model, a simplified transcriptional model was connected to a detailed description of the translation process. The unique approach uses a ribosome flow model to simulate the movement of the ribosome along a one-dimensional discrete template (; ). This approach enables a careful study of the influence of different components on the translation rate. The elongation factor Tu and tRNA concentration were identified as the most sensitive parameters hampering the translation rate. As a key difference between in vitro and in vivo conditions, a control shift from the ternary complex to translation initiation was identified.

An equally complex system was developed to describe the synthesis of a short Met-Gly-Gly peptide in an E. coli-based in vitro system by incorporating 968 reactions and 241 components (). The approach evaluated the stability of pseudo-steady states, revealing the temporal stability of metabolite clusters, their collapse, and re-merge, until a final steady state is reached. Interestingly, increasing tRNA supply also led to a slight increase in translation rates (observed as increased poly-peptide production), but the effect was much less dominant, as shown by .

Analysis and Prediction of Liposome-Encapsulated Protein Synthesis

A special case of in vitro protein synthesis is the encapsulation of CFPS components in liposomes. In this model, only a few stochastically distributed components may be balanced, creating different reaction conditions in the vesicle and outside the vesicle. The initial studies showed that GFP production kinetics strongly depend on liposome size and lipid composition (). In a first attempt to simulate protein synthesis inside liposomes, a medium-sized CFPS model considering 30 species and 106 reactions was connected to a stochastic model for encapsulation (). Later, the model was extended to 280 species and 270 reactions, comprising a coupled transcription–translation model (). The approach allowed screening of GFP production with different start conditions either by looking at different liposome diameters or by considering different quantities of CFPS components inside the liposome. In agreement with continuum CFPS simulations, optimal DNA levels were identified for maximizing GFP formation. Oversaturation of the system with DNA decreased GFP yields. This finding reflected the enormous energy needs for the transcription process. Follow-up studies showed that some of the results could be achieved with a much less complicated model. Here, around 10 reactions were incorporated by lumping reactions for tRNA charging, transcription, translation, and energy regeneration. The simplified model described the behavior of the PURE system under 27 different compositions, rendering resource availability from standard conditions to limitation, remarkably well (; ).

Extending the Scope of CFPS Modeling

CFPS systems have emerged as ideal test beds for genetic circuits, allowing easier and faster prototyping than traditional in-cell engineering. Consequently, mathematical models to describe those systems have been developed (; ). They cover a wide range of regulatory circuits: two-gene cascades (), sigma factor guided regulation (), complex genetic ring oscillators (), and experimentally verified RNA circuit controllers (, ; ). The developed “minimum models” typically consist of three to ten ODEs and mass balances, considering mRNA, regulatory RNAs, or protein products as model species, and mass action, Hill, or Michaelis–Menten kinetics for regulatory descriptions. Even with coarse-grained models, the highly dynamic systems could be mirrored and predicted successfully. The model-guided circuit design significantly reduced development times.

The use of in silico models is not limited to well-defined model systems such as E. coli crude extract or commercial products. broadened their application to study the CFPS capacities of Bacillus megaterium, linking robotic liquid handling with a coarse-grained ODE model (26 parameters, 14 species, and 18 reactions) for the TX-TL system. Key kinetic parameters of the xylose-repressor system were approximated from DNA titration experiments. Simulations were performed using parameters identified by Bayesian parameter inference. Extending the model to describe the concurrent expression of two targets, plasmids carrying GFP and mCherry derivatives revealed competition for translational resources. In general, the reported translation elongation rates (between 0.10 and 0.02 aa s–1) were slower than those reported for CFPS systems. The inefficient use of available energy accounted for the low performance. Another model approach for investigating resource competition in CFPS was formulated by with a minimal model for genetic circuits. Here, the authors could successfully quantify the burden of two targets expressed simultaneously on the resources of a CFPS system.

A constraint-based model to approximate energy and substrate supply from E. coli CFPS extract was presented by Varner and colleagues (; ; ). They connected a simplified description of protein production () and allosteric enzyme regulation () with the metabolic network. Flux balance analysis (FBA) was applied to estimate the flux patterns of central carbon metabolism, amino acid biosynthesis, and energy metabolism using the objective function of maximizing the production rate of chloramphenicol acetyltransferase (CAT). Analysis of different amino acid supply scenarios in silico revealed inefficient energy yields of the experimental in vivo setup, most likely due to unfavorable side reactions. Similar scenarios may have also occurred in the experimental setup for B. megaterium described above.

Discussion

A historical trend can be observed regarding the objectives of CFPS models. Early models (; ; ) are focused on the basic CFPS system using GFP as an experimental readout, beyond classical targets such as ß-galactosidase, chloramphenicol transferase, luciferase, or other likewise “easy to quantify” targets. In the past decade, an increasing number of diverging scientific branches have developed broadening the scope of model building and simulations. Yet, the GFP-based system is described in most detail and is the focus of current investigations. Currently, derivatives of the initial GFP are commonly used, such as its “enhanced” and “super folder” variants (; ). The mRNA product is typically quantified with RNA aptamer reporters such as the malachite green RNA aptamer. The broadening of scientific approaches and the extension of cell-free genetic circuits will increase the need for easy and reliable reporter systems based on short nucleotide or peptide sequences ().

When analyzing the different model granularities (Figure 1), major differences are observed. For coarse-grained models, compromises are made by assuming certain states of the model system by neglecting components or by lumping different metabolites (e.g., all amino acids) to one species. Fine-grained models consider these species in detail. System complexity has been increased from small systems with around 10 equations to large models with hundreds of reaction components (Table 1). Increasing complexity can offer the possibility to resolve bottlenecks by getting insights into reactions or reaction networks. Experimental access to all process elements is hardly possible, and only subsets of information are normally available, even for best investigated bacterial strains such as E. coli. As a result, complex models usually rely on multiple data resources covering different experiments (; ), whereas small models may be well identified by single experiments. Interestingly, it was shown that results gathered with complex systems can also be mimicked with reduced systems (e.g., ). Consequently, deciding a proper CFPS model structure should be driven by the questions to be answered, and should critically reflect the database for model identification.

TABLE 1

ReferencesFeaturesNo. of parametersNo. of speciesNo. of equationskTX [nt s–1]kTL [aa s–1]
Minimal model of the CFPS system10740.504.00
Refined minimal model of the CFPS system8542.200.03
and Simplified CFPS model for screening of different CFPS compositions16961.670.09
Simulation of different transcription (promoter) and translation initiation (ribosome binding site) configurations145510.002.50
Study on different fluorescence protein targets, regulatory elements, and critical evaluation of model prediction121010––
Model description and kinetic parameter estimation for CFPS of non-model bacteria2614188.13–11.470.09–0.11
Identification of translation initiation as bottleneck of CFPS, analysis of different commercial kits131010––
and In lipo protein synthesis, stochastic distribution of components in liposomes24–280106–270–19.004.00
Quasi-stationary state analysis of complex model networks483241968––
(based on )First detailed description of coupled transcription and translation model (Arnold). Comparison of in vitro and in vivo conditions, metabolic control analysis>70174 + no. of codons>500–1.12
; , and Coupling of CFPS to flux balance network of the central carbon metabolism, implementation of allosteric regulation–146264––

Overview of the different granularities of CFPS models.

Each reference is listed with its unique features alongside the number of parameters, species, and equations incorporated in the model. If available, the transcriptional and translational rate constants kTX and kTL as readouts of the model are given. Models that share the same background are lumped together (note that genetic circuit models are not considered here).

The quality of CFPS models is checked by challenging model predictions with experimental observations. Typically, rates for transcription (kTX), translation, and elongation (kTL) are experimental readouts. However, the range of these parameters is broad (Table 1). kTX has been reported from 0.5 () to 19 nt s–1 (). kTL ranged from 0.03 () to 4.00 aa s–1 (; ). The apparent differences may reflect the intrinsic problem of using relatively few experimental readouts to identify models of different complexities (). It has been shown that even simultaneously planned and performed CFPS experiments can lead to significant outcomes between different laboratory sites (). As the modeling studies presented here are based on a wide range of commercial and homemade CFPS systems and extracts, this might explain the deviance of calculated parameters.

Conclusion

Currently, CFPS models can identify bottlenecks in the transcriptional and translational processes as well as infer kinetic parameters from model data. The consensus of most model predictions is the identification of the translational rather than the transcriptional process as one of the key targets for further developments in CFPS systems. Potential starting points are translation initiation, tRNA supply, and recycling. In most approaches, the modeled mechanisms of the translational process seem to be oversimplified. Inspiring approaches for in vivo translation have been published by and that could be adapted to in vitro descriptions. For many modeling purposes, hybrid models can be the ideal tradeoff between complexity and acceptable computational costs. As CFPS systems and genetic circuits get more complex and consider multiple targets (RNAs/proteins), models that consider the joint burden on resources will come into focus (; ). A feedback loop between the model investigation and experimental setup has to be established. The works on genetic circuit models have proved that fast and easy prototyping is possible with CFPS. To unravel the key mechanisms for designing models, data from metabolomics and proteomics have to be integrated. Recent research addresses this need and offers a variety of datasets that could be harnessed by the CFPS modeling community (; ; ). This significantly increases the possibility to describe CFPS with an improved mechanistic resolution, up to a complete dynamic description of the CFPS system components. The development will open the door for a thorough application of tools of statistical systems analysis and metabolic control analysis to translate simulation results into system engineering advice.

Statements

Author contributions

JM reviewed the literature, designed the concept, wrote the manuscript, and prepared the figures. MS-H and RT co-edited and supervised the manuscript. All authors approved the manuscript for publication.

Funding

JM was supported by the UfIB Ph.D. program of “Bundesministerium für Bildung und Forschung – BMBF” (Grant No. 031B0725).

Conflict of interest

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

References

Summary

Keywords

in vitro protein synthesis, cell-free synthetic biology, mathematical model, in silico, ribosomes, transcription and translation, modeling

Citation

Müller J, Siemann-Herzberg M and Takors R (2020) Modeling Cell-Free Protein Synthesis Systems—Approaches and Applications. Front. Bioeng. Biotechnol. 8:584178. doi: 10.3389/fbioe.2020.584178

Received

16 July 2020

Accepted

29 September 2020

Published

28 October 2020

Volume

8 - 2020

Edited by

Simon J. Moore, University of Kent, United Kingdom

Reviewed by

Matthew Lux, U.S. Army Edgewood Chemical Biological Center (ECBC), United States; Ashty Karim, Northwestern University, United States; James T. MacDonald, Imperial College London, United Kingdom

Updates

Copyright

*Correspondence: Ralf Takors, ;

This article was submitted to Synthetic Biology, a section of the journal Frontiers in Bioengineering and Biotechnology

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics