Abstract
Codon optimization has evolved to enhance protein expression efficiency by exploiting the genetic code’s redundancy, allowing for multiple codon options for a single amino acid. Initially observed in E. coli, optimal codon usage correlates with high gene expression, which has propelled applications expanding from basic research to biopharmaceuticals and vaccine development. The method is especially valuable for adjusting immune responses in gene therapies and has the potenial to create tissue-specific therapies. However, challenges persist, such as the risk of unintended effects on protein function and the complexity of evaluating optimization effectiveness. Despite these issues, codon optimization is crucial in advancing gene therapeutics. This study provides a comprehensive review of the current metrics for codon-optimization, and its practical usage in research and clinical applications, in the context of gene therapy.
1 Introduction
Codon optimization first appeared due to the search for an approach to increase the efficiency of expression of target proteins in bacterial cultures. The known property of degeneracy of the genetic code allows mRNA to encode the same proteins in different ways since 20 proteinogenic amino acids can be encoded by 61 codons (Welch et al., 2009). This property formed the basis of the codon optimization method, when, with the advent of genetic sequencing, it became evident that the usage of codons is non-random. Bias in codon usage occurs between different organisms, tissues, and sometimes even between parts of the same gene (; Pouyet et al., 2017). Thus, it became clear that the selection of the most common codons deemed suitable for an organism or cell line during genetic engineering research allows significantly changing approaches to conducting experiments.
Escherichia coli was the first organism with an analyzed codon usage system. Knowing the sequences of anticodons and the abundance of various tRNAs in the cell, the authors identified criteria for codon optimality (Ikemura, 1981). The first criterion was high codon recognition, the second was the highest abundance of tRNA. Highly expressed genes had a bias in frequency of use towards optimal codons, while genes with low expression were characterized by high randomness in the choice of codons ().
Currently, codon optimization has found application in a wide range of topics. In addition to fundamental research, control of the efficiency of protein expression through the selection of synonymous codons is also used for the development and production of biotherapies (), most of which are based on the expression of recombinant proteins. The method has become indispensable for molecular pharming on plants, where the problem of low expression efficiency is most pressing (Perlak et al., 1991; ; Thomas and Walmsley, 2014).
Differentiated cells determine the formation of tissues of various types. This complicated process can be modulated at the cellular and molecular level (Simon et al., 2018). At the molecular level, this diversity is reflected in particular in differences in protein expression - proteins that are abundant in one tissue may be absent in another (Thul and Lindskog, 2018). Differences in protein abundance are, in turn, caused by differences in RNA expression. One of the possible factors affecting such patterns is the different frequency of use of synonymous codons encoding the same amino acid during translation (Kames et al., 2020) (Figure 1). Indeed, either the rarity of codon usage (Plotkin et al., 2004) or the frequency of tRNA variants (; ) both vary between tissues. This can potentially be exploited for the construction of tissue-specific gene therapy. At the same time, to our knowledge, there is currently only one paper in peer-reviewed journals that has experimentally tested this hypothesis (Hernandez-Alias et al., 2023). This study is evidence that tissue-specific codon usage can potentially be used to design tissue-specific transgenes. At the same time, this metric is only one additional tool in the gene design toolbox whose implementation needs to be further explored and cannot be considered in isolation from several other indicators discussed below (Hernandez-Alias et al., 2023).
FIGURE 1
One of the most relevant and important areas of codon optimization application is the development of vaccines. The current way to create non-live vaccines is the use of attenuated viruses. Several research groups have experimented with attenuating poliovirus by changing codon bias in the gene encoding the poliovirus capsid protein, which involved replacing more frequent codons with less frequent ones (; Mueller et al., 2006). Moreover, increasing transgene expression in vaccines may improve the effectiveness of immunization and can be achieved through codon optimization (; ). In addition, a new class of vaccines—mRNA vaccines—has recently been introduced into clinical practice in the context of the COVID-19 pandemic (Oliver et al., 2020). Currently, the possibility of a similar approach for the prevention of infectious diseases such as rabies (Wan et al., 2023), influenza virus (Lee et al., 2023), Zika virus (), Lassa virus (Ronk et al., 2023) is the subject of active research and development. Remarkably, codon optimization of mRNA vaccines can significantly improve their stability and immunogenicity (Zhang et al., 2023). Despite the benefits of codon optimization, it is important to maintain a balance in the use of these techniques. Excessive interest in codon optimization can possibly lead to the accumulation of substances that are poorly excreted from the body, such as, for example, modified mRNA and the corresponding antigen (; Röltgen et al., 2022).
Currently, various approaches could be used for the development of gene therapeutics. Control of the immunogenicity of the administered drug is one of the most vital tasks not only in the preparation of vaccines but also for gene therapies. For the drug to work effectively, it is necessary to reduce the viral vector’s immunogenicity. It has been shown that by varying synonymous codons in the transgene and vector, it is possible to increase the effectiveness of therapy by lowering immunogenicity (; ), which provides optimism for simplifying vector selection and expanding the application of this type of therapy.
Regrettably, codon optimization techniques, while widely employed in the development of gene therapies, are far from perfect and are fraught with several challenges. One prominent issue lies in the incomplete synonymy of substitutions. This drawback carries the potential to disrupt natural post-transcriptional modification sites or, alternatively, give rise to novel sites, leading to critical alterations in the final protein’s structure, properties, and functions (; Irimia et al., 2012). Furthermore, overlooking the existence of alternative translation initiation sites (Malarkannan et al., 1999; Matsuda and Mauro, 2010)can lead to the unintended production of new proteins, adding another layer of complexity to the process. Beyond these intrinsic challenges, the selection of an appropriate numerical method for evaluating the effectiveness of codon optimization poses an additional obstacle. The abundance of metrics available complicates the task, requiring careful consideration to ensure a meaningful assessment. Despite the above difficulties, codon optimization approaches are actively used in clinical trials around the world and, furthermore, COVID-19 mRNA vaccines Pfizer/BioNTech and Moderna employ codon optimization.
Codon optimization can be carried out in many different ways today. It is often not clear which of these approaches is best suited to fulfill a particular task. The purpose of this review is to cover the current state of this problem and future directions for codon optimization approaches for gene therapies.
2 The quantitative assessment of codon usage and optimization
2.1 Measures of codon usage
The codon usage bias (CUB), also known as codon usage preferences (CUP), is influenced by a combination of factors that vary among species. Such factors include mutation frequency (Pizzo et al., 2015), selection for translation efficiency (Navon and Pilpel, 2011), and the presence of transfer RNA (tRNA) molecules that recognize specific codons (; Wei et al., 2019), ribosome binding efficiency (Shi et al., 2020), and translation speed and co-translational protein folding (Mitarai et al., 2008; Liu, 2020).
Based on the non-random usage of codons in the genomes of different species and the previously demonstrated positive correlation between codon bias and gene expression efficiency, Sharp and Li developed the relative synonymous codon usage (RSCU) scale (Sharp and Li, 1986). The RSCU value was calculated for a set of genes as the ratio of the observed codon frequency to the expected frequency, assuming equal usage of synonymous codons. This research has made a substantial contribution to the creation of various metrics, including but not limited to codon adaptation index (CAI) (Sharp and Li, 1987), average ratio of RSCU (ARSCU) (), and genetic tRNA adaptation index (gtAI) (). CAI continues to be a widely employed metric in both commercial and academic applications. CAI reflects the level of species-specific codon adaptation and is calculated as the geometric mean of RCSU values for each codon in the gene relative to the value of the most frequently used triplet encoding a single amino acid.
To date, numerous metrics for quantitative assessment of the level sequence optimization have been developed. Table 1 offers concise descriptions of commonly used metrics. To give the readers an idea of the frequency of metric usage, we added the citation rate of the original sources. However, it is important to emphasize that this approach does not reflect the level of usage of optimization tools based on the mentioned metrics.
TABLE 1
| Index | Mathematical formula | Principle | Interpretation | Original source | Citation count |
|---|---|---|---|---|---|
| Frequency of optimal codons (Fop) | Fop = Nopt/Ntotal, where Nopt denotes number of optimal codons in a gene, Ntotal indicates number of total codons in sequence | The rationale behind Fop is that genes with a higher expression level tend to preferentially use certain codons, and this bias can be quantified using this ratio | Representing the occurrence of the optimal codon within a gene sequence. Values close to 1 indicate a strong bias toward optimal codons, while values closer to 0 suggest a more even or random usage of synonymous codons | Ikemura, 1981 (1982) | 1724 and 768 |
| Relative synonymous codon usage (RSCU) | where xij denotes the number of occurrences of the j-th codon for the i-th amino acid, ni represents the degeneracy for the i-th amino acid | Calculated as the ratio of the observed codon frequency to the expected frequency assuming equal usage of synonymous codons | The codon with RSCU of 1 value indicates average (random or equally) synonymous codon usage; RSCU value greater than 1 indicates positive codon usage bias (overrepresented); RSCU value less than 1 indicates negative codon usage bias (underrepresented) | Sharp and Li (1986) | 653 |
| Codon adaptation index (CAI) | , where L is the length of a gene measured in codons. Wk is the relative adaptiveness value for the k-th codon in the gene | Quantifies the geometric mean of the RSCU for each codon with respect to the codon usage of a reference set of highly expressed genes. Genes that possess higher scores are anticipated to demonstrate enhanced efficiency in translation and elevated levels of protein expression | A higher CAI score implies that the gene’s codon usage is better aligned with the preferred codons identified in a highly expressed reference set. A CAI value of 1 suggests that the codon usage in the gene is perfectly adapted to the preferred codons observed in a highly expressed reference set; value of 0 indicates that the codon usage in the gene is not adapted | Sharp and Li (1987) | 4091 |
| Codon pair adaptation index (CPAI) | , where N − 1 is the number of codon pairs in gene g and w-i, j is the relative adaptiveness of the codon pair (i, j). , where fi, j is the frequency of codon pair (i, j) and f(max(i,j)) is the frequency of the codon pair most often used to code for the amino acid pair (aa(i),aa(j)) in a set of highly expressed genes G | The advantage of the index is the automatic weight selection algorithm . Combined use of CPAI and CAI have been shown to better predict gene expression | Similar to CAI. | 43 and 402 | |
| Codon bias index (CBI) | , where Nopt represents the number of preferred optimal codons, Ntot is the total number of codons in the gene, and Nrand denotes the expected number of preferred codons if random codon assignments were made for each amino acid | Codon bias refers to the unequal usage of synonymous codons encoding the same amino acid in a DNA sequence. The codon bias index is a measure that quantifies the extent of this bias | Values range from 0 to 1: 0 indicates random selection of codons, and 1 indicates a significant bias towards preferred codons | 1955 | |
| Effective number of codons (ENc or Nc) | , where , where is the degeneracy of the amino acid (number of synonymous codons), is the fraction of each synonymous codon () out of the total codons () for that amino acid, | The ENc is a measure designed to assess codon usage bias by evaluating how far the observed usage of synonymous codons deviates from an anticipated equal distribution | The ENc values range from 20 (maximum codon bias, only one codon used for each amino acid) to 61 (no bias, all synonymous codons used equally). If the ENc value is less than 35, it indicates a pronounced bias | Wright (1990) | 2306 |
| Effective number of codon-pair (ENcp) | , where Fm is the average degeneracy of all amino acid pairs with m synonymous codon pairs, and km is the number of amino acid pairs with m synonymous representations | ENcp is a metric designed to assess the bias in the usage of codon pairs, defined analogously to ENc, with the addition of a square root | Similar to ENc | 100 | |
| Codon usage similarity index (COUSIN) | COUSIN evaluates the codon CUB of a given query in comparison to a reference, and it standardizes the results by applying a Null Hypothesis that assumes random codon usage. In COUSIN18, all 18 families of synonymous codons contribute equally to the overall index, whereas in COUSIN59, each family contributes proportionally based on the frequency of the corresponding amino acid in the query | A COUSIN score of 1 signifies that the CUB in the query closely resemble those in the reference dataset. A COUSIN score of 0 indicates that the CUB in the query align with those in the Null Hypothesis. For scores exceeding 1, the CUB in the query share similarities with those in the reference but on a larger scale. Scores falling between 0 and 1 suggest that the CUPrefs in the query are akin to those in the reference but with a smaller magnitude. A score below 0 implies that the CUPrefs in the query are opposite to those in the reference | 51 | ||
| , where N is the number of amino acids present in both the query and the reference and is the set of synonymous codons coding amino acid a. weight for each codon (Wc,a) in reference and test set is the frequency of the amino acid a in the query | |||||
| Average ratio of RSCU (ARSCU) | , where aa is amino acid, a is RSCU of GC end codons and b is RSCU of AT end codons (any a and b with a value of zero is arbitrarily assigned a value of 0.1) | Measures the ratio of RSCUs with GC-ending codons to the AT-ending codons for all amino acids in a gene | Genes exhibiting ARSCU values surpassing 13 (a subjective threshold) are anticipated to showcase heightened expression levels. Genes with ARSCU values within the range of 9–13 are predicted to have either high or intermediate expression, while genes possessing ARSCU values below 9 are expected to demonstrate low or intermediate expression | 21 | |
| Relative codon bias strength (RCBS) | , where L is the length, in codons, of the gene, f (x,y,z) the observed frequency of codon xyz and f1(x), f2(y) and f3(z) the observed frequencies of bases x, y and z at, respectively, codon positions 1, 2 and 3 | It calculates the observed frequency of specific codons relative to the expected frequency, considering biases in base composition at three codon sites, providing a reliable measure of codon preferences while accounting for sequence-specific features like GC content | An RCBS value near 0 implies an absence of codon usage bias, while a value exceeding 0.5 indicates a notable preference for specific codon usage | Roymondal et al. (2009) | 109 |
| Directional codon bias score (DCBS) | DCBS is based on RCBS and allows the measurement of both positive and negative codon usage bias | Similar to RCBS. | Sabi and Tuller (2014) | 116 | |
| Modified relative codon bias strength (MRCBS) | , where fxyz is the normalized codon frequency of a codon xyz and fn(m) is the normalized frequency of base m at codon position n in a gene. RCBSaa, max is the maximum value of RCBS of codon encoding the same amino acid aa in the same reference set, and N is the codon length of the query sequence | Ribosomal protein is used as a reference set of genes and the dependence on gene length is overcome | The score of the modified relative codon bias ranges from 0 to 1 | 1 | |
| Relative codon adaptation (RCA) | , where fxyz is the observed relative frequency of codon xyz in any reference gene set, fi(m) is the observed relative frequency of base m at codon position i in the same reference set and L is the length of the query sequence | The RCA index calculates the anticipated frequency of a codon within a provided reference set by considering the positional base frequencies. Then, it assesses codon adaptation by comparing the observed codon frequency with the anticipated frequency | The score of the RCA ranges from 0 to 2 | 107 | |
| Codon deviation coefficient (CDC) | The calculation of the metric can be found in the original source | The index takes into account the nucleotide composition of the sequence, GC content, and purine content. CUB is estimated using the cosine distance between the expected and observed codon usage vectors. Assessing the statistical significance of results using bootstrap resampling | Values range between 0 and 1: value 0 - no bias, 1 - maximum bias | Zhang et al. (2012) | 62 |
| Index of Translation Elongation (ITE) | , where is the frequency of codon , is the number of sense codons (excluding those in single-codon families) | CAI-like index, but it involves determining a weight for each codon by considering its frequency within the NNR and NNY codon subfamilies in the reference set. Subfamilies are distinguished due to different translation by different tRNAs and susceptibility to different mutational errors | ITE values range between 0 and 1. A greater score is assigned to genes containing codons that are more commonly found in highly expressed genes | [25] | 136 |
| Synonymous codon usage order (SCUO) | where j is the codon i-th amino acid. SCUOi is the SCUO for i-th amino acid in each sequence and Hi and Hmaxi are the entropy and maximum for an i-th amino acid in a sequence | The SCUO index assesses how much a sequence deviates from a uniform distribution, using Shannon entropy as a basis. It involves the normalized difference between the maximum entropy and the observed entropy | SCUO the value varies between 0 and 1, and higher values indicate a stronger codon usage bias | Wan et al. (2006) | 21 |
Metrics for codon optimization with formal definition and description. The number of citations was retrieved from the Scopus database.
TABLE 2
| Synonymous codon variants | 1st | 2nd | 3rd | 4th |
|---|---|---|---|---|
| Amino acid | ||||
| A | GCT | GCC | GCA | GCG |
| D | GAT | GAC | ||
| G | GGT | GGC | GGA | GGG |
| Y | TAT | TAC | ||
Example representation of the 4-letter amino acid sequence ADGY (alanine-aspartic acid-glycine-tyrosine) via synonymous codons. Nucleotide sequence of wild-type GCC-GAT-GGT-TAT. There are 4 codon variants for the first and third amino acids, and 2 variants for the second and fourth amino acids. Total 64 possible variants of nucleotide presentation of this sequence.
Numerous metrics can be easily calculated with a reference set of genes to obtain the codon usage frequency. For example, Fop is calculated as the ratio of optimal codons to the total number of codons, excluding stop codons and codons without alternatives for amino acids (methionine, tryptophan) (Ikemura, 1981; 1982). The index aids in gauging the prevalence of synonymous codon usage. Other metrics are grounded in the assumption that the usage of codons is non-random. The metrics quantify the difference in codon usage frequency from a uniform distribution within the coding sequence. When all codon variants for a specific amino acid are utilized with equal frequency, such difference is minimal. Conversely, the maximum is achieved when only one codon out of the possible ones is utilized. Examples of such indices include ENC, CDC, SCUO, and others.
2.2 Codon adaptation metrics for assessing mRNA properties
Codon optimization is a strategy aimed at increasing the efficiency of mRNA translation and overcoming protein expression limitations. The use of synonymous codons affects the stability of mRNA in human cells (Narula et al., 2019; Wu et al., 2019). The thermodynamic stability of mRNA within a cell significantly influences translation efficiency (; ). mRNA is inherently unstable and can undergo transient states and adopt multiple stable structures. One approach to selecting synonymous amino acids for the purpose of thermodynamic stabilization is aimed at minimizing the free energy ΔG (MFE) released during RNA folding (Zuker and Stiegler, 1981; Zuker, 1994). Ringner and Krogh demonstrated in Saccharomyces cerevisiae that the folding free energy in the vicinity of the 5′-UTR correlates positively with transcription efficiency and mRNA half-life (Ringnér and Krogh, 2005).
An alternative approach suggests that the optimal structure will possess the maximum number of chemical bonds (Wayment-Steele et al., 2021). The AUP (Average Unpaired Probability) and SUP (Sum of Unpaired Probabilities) metrics, employed to assess RNA stability against hydrolytic degradation, operate under the premise that structures formed by paired bases exhibit lower susceptibility to hydrolysis.
Cluster analysis discovered that different mRNAs preferentially use different types of codons. Some mRNAs predominantly use optimal codons, while others prefer non-optimal codons. Furthermore, they observed that mRNAs with a higher proportion of optimal codons tend to be more stable, while those with a lower proportion of optimal codons are more unstable. Based on conducted experimental research, a metric called the codon stability coefficient (CSC) has been proposed. It is calculated as the Pearson correlation coefficient between the frequency of each codon and mRNA half-lives (Presnyak et al., 2015).
In the standard genetic code, the first two positions of a codon play a decisive role in coding an amino acid, while the third position is variable for one amino acid. Collection of metrics developed GC1, GC2, and GC3 represents the frequency of G + C usage at the first, second, and third positions, respectively (Stenico et al., 1994). Another evaluation derived from RSCU is the Average RSCU Ratio (ARSCU) (). Its noteworthy feature involves considering the base at the third position of the codon. The optimization of protein expression often involves the frequent usage of GC content. The model of post-transcriptional mRNA regulation involving P-bodies, 5′-3′ exonuclease XRN1, RNA helicase DDX6, and enhancer of decapping PAT1B shows that GC-rich coding sequences (CDS) result in higher protein production compared to AU-rich ones, and are controlled by a mechanism involving degradation factors DDX6 and XRN1 (). On the contrary, reducing the GC content in the 5′-UTR leads to an increase in free energy and also enhances protein yield, presumably due to mRNA destabilization in the translation initiation region and greater accessibility of the ribosome binding site (). The GC3 content varies depending on the type of tissue but is not an exhaustive characteristic for tissue-specific gene separation (Plotkin et al., 2004). GC3 codons are also associated with a longer half-life of mRNA (Kudla et al., 2006; Hia et al., 2019).
2.3 Metrics for adaptation to tRNA pool
Codon usage bias is closely linked to translational selection, which is the process of selecting codons that match abundant tRNAs, the molecules responsible for carrying amino acids during protein synthesis. Highly expressed genes tend to use such preferred codons, resulting in enhanced translation rates and accuracy. showed that the expression levels of nuclear and mitochondrial tRNAs vary between human tissues, indicating tissue-specific translational selection. However, minor differences in mouse mitochondrial RNA have only been detected for cardiac tissue, while significant differences between the central nervous system and other tissues have been demonstrated at the level of tRNA isodecoders, i.e., transcripts with the same anticodon but encoded by numerous different genes (Pinkard et al., 2020). It is important to note that the strength of translational selection varies across different organisms based on their genome sizes and genomic tRNA content (Reis, 2004).
To account for the role of intracellular tRNA content in translation efficiency, the following indices have been developed: P2index () and tRNA adaptation index (tAI) ().
Initially, tAI was only applicable to S. cerevisiae, but its subsequent modifications, stAI (Sabi et al., 2017) and gtAI ()—overcome this limitation by incorporating species-specific weights through algorithmic approaches to find extrema. gtAI demonstrated greater efficiency by employing a genetic algorithm to identify the optimal set of weights. In its calculation, indices ENc and RSCU are also incorporated. gtAI ranges from 0 to 1, where a higher value implies better adaptation of the codon to the tRNA pool.
The P2 Index is a metric used for the quantitative assessment of the efficiency of interactions between codons and their corresponding anticodons during the translation process. Based on the frequency of specific types of codons, values exceeding 0.5 indicate the presence of translational selection influencing the coding sequence.
2.4 Algorithmic approaches and tools for codon optimization
Currently, various optimization algorithms are utilized, such as the genetic algorithm (), multi-objective artificial bee colony (), Ribotree Monte Carlo (Leppek et al., 2022), and dynamic programming (Pham et al., 2004; Taneda and Asai, 2020), to identify codon combinations with desired characteristics. In several studies, the use of recurrent neural networks for codon optimization in heterologous protein expression has been presented in Chinese hamster (Gricetulus griseus) ovary cells () and E. coli (Jain et al., 2023). The Bidirectional Long Short-Term Memory (LSTM) deep learning model has also been trained for E. coli ().
Other studies applied machine learning methods for mRNA stabilization, such as integrated deep learning-based mRNA optimization (iDRO) (Jain et al., 2023), which provides a two-step optimization for the open reading frame and the untranslated regions. S. Castillo-Hair and G. Seelig trained a model on the 5′UTR polysome profile dataset to predict ribosome loading and protein expression (). The predictive power of such models strongly depends on the quantity and quality of the training datasets. At the same time, the accumulation of experimentally verified data sets is often not as fast as the development of machine learning methods. For example, to date (February 2024) only 6,142, of which 1,416 are human, experimentally validated RNA structures have been deposited in the Protein Data Bank (). This indicates that the high-precision prediction of RNA 3D structures using machine learning methods may be accurate for training data, but not for new data (Sato and Hamada, 2023).
Several software tools that utilize statistical and algorithmic solutions are available for commercial and free use. Here, we present some current tools that can be used for various tasks, including those related to gene therapy: ATGme (), OPTIMIZER (Puigbo et al., 2007), CHARMING (Wright et al., 2022), %MinMax (Rodriguez et al., 2018), JCat (), Optipyzer (LeRoy and Roleck, 2023), IDT (Owczarzy et al., 2008), gtAI ().
3 Codon optimization for gene therapy vectors
Above, the elucidation of metrics and principles related to codon optimization has been expounded. At the same time, it should be noted that the resources required to test the functionality of in silico predicted RNA variants significantly exceed the cost of the prediction itself. For this reason, studies often mainly present unconfirmed hypotheses in in vitro or in vivo experiments. Nevertheless, we present below some examples where codon optimization has been successfully applied in vitro. Proceeding to in vitro studies, it should be noted that gene therapeutics consist of a delivery vector and a therapeutic gene. Currently many types of vectors are used as a transgene vehicle (e.g., lipoplexes (), polyplexes (), virus-like particles (Pitoiset et al., 2017)).
Some of these vectors are a cassette with the selected viral genes, others do not contain nucleic acids. In some cases, wild-type viral genes in the gene therapy vector are not optimized for efficient application (). At the same time, codon-optimized variants of these sequences increase the efficacy of gene therapy, although they may lead to unfavorable results such as undesirable conformational changes and subsequently alterations in protein activity and function. Examples of codon optimization of adenoviral (), retroviral and lentiviral vectors () are discussed below.
Since adeno-associated vectors have recently become the most widely used platform for gene transfer (Mendell et al., 2021) and adenoviruses have long been successfully used to deliver genes (), we will consider the application of optimizations on their example.
It has been shown that in adenoviruses, the genes responsible for highly abundant late structural proteins tend to use codons frequently used in humans (optimal codons), while early regulatory use less optimal codons (Villanueva et al., 2016). However, the adenoviral fiber protein specifically uses suboptimal codons for efficient viral replication. Surprisingly, analysis of transgenes expressed in oncolytic adenoviruses, that are used for the oncoselective expression of a wide range of therapeutic molecules (; Huang et al., 2019) shows that most transgenes also use suboptimal codons. This contradicts the recommendation to use optimal host codons in transgenes to maximize gene expression. The study investigates the impact of transgene codon usage on viral fitness and finds that transgenes with higher GC3 content (optimal codon usage) have higher gene expression and viral replication, while those with lower GC3 content have lower expression and replication (Núñez-Manchón et al., 2021). By tuning the codon usage of transgenes, it is possible to achieve better transgene expression without compromising viral replication, thus optimizing the therapeutic outcome.
In the development of gene therapies, the problem arises of achieving high titers and a high ratio of empty to full capsids in viral vectors. One of the solutions to this obstacle is codon optimization of viral genomes encoding capsid proteins and assembly proteins. Thus, not only transgenes but also the coding sequences of the viral vector itself are subjected to codon optimization. For AAV-based (adeno-associated virus) vectors a novel codon optimization method was presented (Localized Codon-Optimization or LCO) ().
This method aims to preserve functional elements of the capsid genes and improve capsid shuffling efficiency for AAV engineering. The LCO algorithm performs localized optimization of codons at each position independently, based on the usage frequency of codons observed in the input variants of AAV sequences. A codon usage frequency table is generated for each amino acid position, and this table is used to optimize individual sequences (Table 3). The LCO-modified capsid genes showed increased sequence identity between parental AAV capsids and novel AAV capsid variants.
TABLE 3
| Probability of finding a codon at a given position | ||||
|---|---|---|---|---|
| codon | Position 1 | Position 2 | Position 3 | Position 4 |
| GCT | 0.4 | |||
| GCC | 0.17 | |||
| GCA | 0.24 | |||
| GCG | 0.19 | |||
| GAT | 0.68 | |||
| GAC | 0.32 | |||
| GGT | 0.26 | |||
| GGC | 0.22 | |||
| GGA | 0.36 | |||
| GGG | 0.16 | |||
| TAT | 0.39 | |||
| TAC | 0.61 | |||
An example of how the LCO method works to optimize the four codons of the mRNA encoding ADGY (see Table 2). A probability is calculated for all possible codons for a particular amino acid at a particular position. The most probable codons are marked in bold. Accordingly: GCC-GAT-GGT-TAT (wild-type nucleotide sequence)—would be optimized to GCT-GAT-GGA-TAC (final LCO-optimised sequence).
Functionality tests demonstrated that the optimized capsids retained their function, and transduction efficiency was similar to unoptimized counterparts. The LCO method also improved the efficiency of capsid shuffling, resulting in a highly shuffled library with increased complexity and reduced size of donor sequence segments. The shuffled clones generated using LCO-encoded capsids demonstrated successful transduction, indicating the effectiveness of LCO in generating novel AAV variants.
Ironically, the extensive use of codon optimization occurred simultaneously with abundant research findings that revealed the impact of synonymous mutations on protein function. This has been shown on a variety of proteins (; Kirchner et al., 2017).
The mechanism being discussed involves the comparison between codon-optimized (CO) and wild-type (WT) variants of a protein named FIX (coagulation factor IX). The results highlight that the CO and WT FIX variants exhibit distinct conformations, suggesting that the codon optimization process has influenced the protein’s structure. Ribosome profiling analyses uncover altered ribosomal distribution patterns and local translational kinetics in the CO variant when compared to the WT variant. Notably, these differences are unique to the CO FIX variant, as control genes demonstrate comparable ribosome distribution profiles ().
Despite the observed differences in translational kinetics, the overall efficiency of protein synthesis between the CO and WT variants remained similar. This finding is consistent with previous studies conducted in vitro (outside of a living organism) and suggests that the rate of protein synthesis is comparable between the two variants. The researchers propose that differences in translational kinetics within these domains may contribute to the observed conformational differences between the CO and WT FIX variants.
Codon optimization can be approached not only by a global view of codon usage in general, but also by a local optimization for each individual position in a particular amino acid. Moreover, it is also important to check that the functions of the essential elements and the optimized protein of interest remain unchanged.
4 The effect of codon optimization on immunogenicity
The immune response to an administered foreign substance or molecule can be defined as immunogenicity. It should be noted that higher immunogenicity increases the efficacy of the drug in some cases, but decreases it in others (Figure 2). For example, the purpose of immunization is to generate an immune response against a pathogen. In this case, methods should be used to increase the immunogenicity of the drug. It should be noted that in the development of mRNA vaccines, an excessive overreaction of the immune system is undesirable due to possible damage to the human organism (Igyártó and Qin, 2024) and should be taken into account during codon optimization. On the other hand, if a transgene introduced into the organism is intended to lead to the production of the corresponding protein, any degree of immunogenicity will reduce the effectiveness of the therapy. The innate and adaptive immune response to gene therapy may vary depending on the source of immunogenicity. These may be factors related to the capsid of the virion or to the viral genome. In relation to the capsid, binding of TLR2 or TLR9 can potentially activate the innate immune response and initiate the MyD88 signaling cascade, which in turn stimulates the production of proinflammatory cytokines such as TNF-alpha or induces the synthesis of IFN-gamma (Yang et al., 2022). Depending on the composition of the viral vector, the innate immune response can lead to enhanced adaptive immune responses. For example, AAVs, which are often used as gene therapy vectors, circulate naturally between humans. As a result, most people develop antibodies against natural AAV serotypes due to previous exposure. These antibodies are even known to cross-react with engineered vectors (). As a result, these antibodies can lead to either complement activation or neutralization of the capsid. The adaptive immune response is characterized by the degradation of the capsid protein by the proteasome and peptide presentation on MHC class I molecules. CD8+ cytotoxic T-cell lymphocytes can bind to the MHC, which leads to cell death (Martino et al., 2013). Peptide presentation on MHC class II molecules after phagocytosis and proteolysis can be recognized by CD4+ T lymphocytes, which can then stimulate the proliferation of B cells and the production of capsid-specific antibodies (Li et al., 2013). Studies have shown that plasmacytoid dendritic cells (pDCs) and conventional dendritic cells (cDCs) co-operate to achieve cross-priming of CD8+ T cells (Rogers et al., 2017). pDCs recognize the AAV genome via TLR9, while cDCs present the antigen on MHC I. The binding of cytokine-produced IFN to its receptor on cDCs is necessary for this process, indicating a direct relationship between pDC-produced cytokines and the activation of cDCs. Cross-priming of CD8+ T cells against AAV capsids requires CD40−CD40L co-stimulation, which is performed in addition to T1 IFN from CD4+ Th cells (Shirley et al., 2020b).
FIGURE 2
After viral uncoating, TLR9 receptors can recognize unmethylated CpG motifs in the released single-stranded DNA, which also leads to activation of the innate immune system and stimulates cytokine production. The humoral and cellular innate immune responses described above for AAV capsids also occur for the transgene protein. The adaptive immune response can depend on various factors such as the target tissue, vector design and dose. Depending on the specificity of the promoter, there is a potential risk of immunogenicity (Shirley et al., 2020a). For example, a ubiquitous promoter can increase the risk of an adaptive cellular immune response of target and non-target cells (Sun et al., 2005).
It should be noted that the appearance of a foreign protein in the human organism is associated with the development of autoimmune diseases due to the similarity of individual epitopes of foreign and self proteins (Rojas et al., 2018). For example, it was recently shown that the same antibodies cross-react with the Epstein-Barr virus protein and the human alpha-crystallin B protein (Thomas et al., 2023). This phenomenon of molecular mimicry could be associated with the development of multiple sclerosis. The possibility of molecular mimicry of proteins resulting from the translation of the nucleic acids used must therefore be taken into account in the development of gene therapeutics. As already mentioned, codon optimization of the RNA can influence the structure of the translated protein (). As a result, depending on the different variants of the synonymous substitutions, the presentation of different epitopes of the same protein is possible.
It is of interest to reduce these CpG motifs to circumvent the possible human immune response, which can be achieved by codon optimization. For example, various elements of an AAV vector such as the CMV enhancer and promoter, ITR regions, UTR regions and the therapeutic transgene itself may contain CpG motifs. The CpGs within the promoter sequence can be removed, but with unpredictable effects on the activity and specificity of the promoter. For example, the authors have shown that the removal of CpGs within the CMV promoter gene significantly reduces its activity (Yew and Cheng, 2004). Although CpGs can be removed from the expression cassette, as in the case of human coagulation factor IX (hFIX) (), this does not always increase efficiency—CpG elimination had only reduced antibody formation against the transgene and not against the capsid itself. There are several studies in which this strategy was used, but mostly with a modification of the transgene. They have shown that the elimination of CpG motifs may lead to a significant reduction in the CD8+ T cell response (Yew and Cheng, 2004; ; Herzog et al., 2019; Wright, 2020; ; Konkle et al., 2021).
Several codon optimization strategies, including the chemical modification of nucleosides (Karikó et al., 2005) and the incorporation of pseudouridine (Karikó et al., 2008; ; Thess et al., 2015), have been shown to improve translation and reduce the immune response to mRNAs. pDCs exposed to such modified RNA exhibit a significant reduction in cytokines and activation markers. Nucleoside modification at a single position in a chemically synthesized oligoribonucleotide (ORN) is sufficient to abrogate TLR activation. In addition, the incorporation of pseudouridine in particular has been shown to facilitate evasion of recognition by Toll-like receptors (Karikó et al., 2005), although the molecular differences contributing to this mechanism has not yet been elucidated. Although the implementation of pseudouridine increases the stability of the mRNA and its translational capacity, it is important to note the disadvantages of replacing uridine with pseudouridine (Xia, 2021; Mueller, 2023). A recent study has shown that the presence of pseudouridine in IVT mRNA increases ribosomal + 1 frameshifting during mRNA translation. In addition, new peptides were generated that triggered an immune response (Mulroney et al., 2024). The presence of pseudouridine in the stop codon region suppresses translation termination and allows non-canonical base pairing, which is particularly detrimental for in vitro transcribed mRNAs (Loomis et al., 2016). The negative effects of pseudouridine synthases have been associated with various cancers (Xue et al., 2022) and autoimmune diseases (). This strongly suggests that the influence of codon optimization and pseudouridine incorporation on mRNA expression needs to be further investigated. A limitation of the present review is that it does not focus on a detailed description of the specific effects of codon optimization on the mRNA vaccines against COVID-19 per se that have been introduced into clinical practice (reviewed in Xia, 2021), but aims to discuss the advantages and disadvantages of the different options for the use of codon optimization in gene therapy in general.
To summarize, a common strategy to avoid immunogenicity is to eliminate redundant CpG motifs, implement chemical modifications of ORNs and replace uridine with pseudouridine. However, it should be noted that the implementation of codon optimization to eliminate CpG motifs and pseudouridine modification must be performed strategically to avoid the negative consequences of both approaches. Given the various unresolved factors leading to potential immunogenicity as a consequence of gene therapy, developing metrics for prediction is a complicated task. Nevertheless, a recent report (Wright, 2020) proposed a metric for prediction focusing exclusively on CpG motifs and their potential immunogenicity. Three formulas were developed that take into account the amount of unmethylated CpG motifs in the vector sequence. Known immunostimulatory sequences commonly used in DNA vaccines were also considered in the development of the formulae (). Although these formulae still need to be improved for full validation and accurate prediction, they reflect the beginning of a deeper understanding of how codon optimization can contribute to the reduction of immunogenicity.
5 Experimental testing of codon optimized sequences
There are numerous strategies for optimizing codons in nucleic acids. The methods mentioned above enable the creation of numerous optimized sequence variants. However, experimental verification of properties such as mRNA stability and protein expression levels is necessary before further experimentation can be conducted. Depending on the goals and available resources, it may be possible to select the best candidates based on chosen criteria from the range of design variants. These candidates can then be examined using routine laboratory methods. Alternatively, a pool of hundreds of sequences can be studied, in which case high-throughput protocols must be developed (Figure 3).
FIGURE 3
When studying a small number of variants, it is possible to determine the expression level separately for each construct after transfecting the cells. To quantify transgene expression in this case, the most common method is to use target-specific primers with cDNA obtained from RNA by reverse transcription as a matrix and perform qPCR (Leppek et al., 2022). Expression can be quantified at both the transcriptional and translational levels. The latter involves the analysis of synthesized proteins and can be performed using antibodies specific to the target protein. For instance, Zhang (Zhang et al., 2023) described the properties of the optimized structure of the SARS-CoV-2 virus S protein using flow cytometry. A possible alternative method for determining protein concentrations is to use SDS-PAGE gels for Western blot analysis, along with specific antibodies (Raab et al., 2010; ).
Although codon optimization of the target sequence can provide certain benefits, it may also result in reduced mRNA stability in solution, which impairs its functionality. Therefore, it is necessary to experimentally confirm the stability of the structure of optimized nucleic acids. The stability of mRNA molecules is inversely proportional to their degradation rate in solution. To determine the degradation rate, mRNAs are incubated in PBS buffer containing Mg2+ ions. Samples are collected at various time intervals of 1–2 h, and the number of fragments produced is estimated using capillary electrophoresis (Zhang et al., 2023) or polyacrylamide gel electrophoresis with urea. Therefore, the RNA is less stable if it degrades more quickly after being incubated in solution.
However, the laboratory approaches described above are time-consuming when testing multiple variants of codon-optimized sequences. In light of this, there is a great need to create high-throughput methods for studying many sequences simultaneously.
Most methods that allow mass screening of sequences follow a general principle: a unique barcode, a sequence of several nucleotides, is inserted into each variant. All the sequences to be tested can then be pooled and processed in a multiplex format. The presence of the barcode makes it possible to identify a variant using high-throughput sequencing platforms after all the necessary protocol steps have been completed.
Massively parallel variant analysis requires the synthesis of a library of DNA templates. The next steps in the study can be performed in two ways. The first involves transcription and modification (3′ polyA tail and 5′ m7G capping) in vitro, followed by transfection of the resulting mRNA pool into cells for further experiments. The “PERSIST-seq” method was developed based on this approach. It enables the simultaneous evaluation of stability and translation efficiency of over 200 mRNA molecules, making it a convenient tool for messenger RNA development (Leppek et al., 2022). In this case, the design of the DNA must take into account the presence of a promoter in the initial sequence. The second approach involves creating a vector library with cassettes that contain the sequence under study and regions of homology. The cells are then transfected with the library, and the sequences are integrated into the genome using CRISPR/Cas. This process enables the direct synthesis of mRNA within the cells. A study of the motifs that cause ribosome slowdown in a yeast model system describes a similar approach (). The next steps for experimental validation in both cases involve isolating RNA from cell culture, analyzing it through high-throughput sequencing, and quantifying the results. To identify inserts in the pool of isolated nucleic acids, unique barcodes are introduced into the library construct, which is a common aspect of the described strategies.
The presence of unique barcodes in the original DNA matrices allows quantitative assessment of the expression level for each individual variant using high-throughput RNA sequencing.
Translation of sequence variants has been demonstrated to be a crucial determinant in mammalian gene expression (). However, genomic expression profiling alone cannot reveal the precise regulation provided by post-transcriptional mechanisms, such as 5′ capping, splicing, polyadenylation, nuclear export, translation, and decay. To overcome this limitation, a polysome profiling method can be used to isolate ribosome-free and polysome-associated RNAs for further independent analysis (Pereira et al., 2018) This method involves separating mRNA in a sucrose gradient into two fractions: polysome-bound and polysome-free. The mRNA is then isolated from both fractions and sequenced using one of the available high-throughput platforms.
When studying multiple variants, stability assessment is also important. To identify full-length molecules that have not degraded, it is necessary to amplify the cDNA that was reverse transcribed from the RNA and then sequence it to quantify the amount of intact mRNA at each time point. This method can evaluate mRNA stability in both solution and cells. The solution replicates the conditions in which the molecules may be present during therapy, typically high pH and positively charged media. It is important to note that the outcomes obtained after incubation in solution differ significantly from those obtained after isolation from cells. This is likely due to cellular mechanisms of RNA degradation (Leppek et al., 2022).
Therefore, there are approaches that allow for the evaluation of the efficiency and stability of nucleic acid sequences obtained during codon optimization. The choice of a particular method depends on the number of variants to be analyzed. If there are only a few variants, it is possible to describe the properties of each variant separately, providing a fairly accurate understanding of its characteristics. When dealing with hundreds or thousands of variants, high-throughput methods are necessary. This allows for a pool of samples to be tested instead of individual samples, greatly increasing the productivity of experimental work. It is important to note that massively parallel sequencing methods provide high accuracy analysis, while polysome profiling can offer additional insights into the impact of codon optimization on the final product’s quality.
6 Future directions
Currently, there are some gene therapies that use different codon optimization metrics and are approved by the FDA (). To analyse other therapies that are in clinical trials and where codon optimization has been used, we conducted a thorough examination of the data available on ClinicalTrials.gov () until December 2023. A systematic search strategy was devised using the keyword “gene therapy” in the Condition/disease field. In addition to the specified search criteria, it is important to note that the term “vector” was included in the “Other terms” considered in the search. The algorithm did not include any specified values for the “Intervention/treatment” and “Location” categories in the search process. After searching, the algorithm automatically incorporated synonyms for the given query: gene: “Genes,” gene therapy: “Gene transfer”; “Gene Transfer Procedure,”, therapy: “treatment”; “Therapeutic”; “therapeutics”.
Furthermore, a comprehensive search was conducted using the specific only Condition/disease of “codon optimized” and excluded any specified values for the “Other terms,” “Intervention/treatment” and “Location” categories in the search process. However, it is crucial to mention that studies explicitly referring to monoclonal antibodies and enzymes as drugs in the Study URL and Brief Summary columns were manually excluded from the sample. This careful exclusion strategy ensured that the selected studies focused specifically on codon optimization. The search was conducted over a period of 20 years to capture an extensive range of relevant clinical studies.
Of the 395 clinical studies analyzed, only 12 contained information on codon optimization (Figure 4).
FIGURE 4
Prior to experimental testing of codon-optimized sequences using any of the aforementioned methods, it is essential to synthesize these sequences, often in large quantities. The most widely used method currently is phosphoramidite synthesis, which involves the interaction of nucleotide phosphoramidite monomers protected by acid-labile groups with an activating agent, binding to the growing oligonucleotide (Sinyakov et al., 2021). There are two main types of implementation for this approach, depending on the equipment used: synthesis on columns or on microarrays. The former option allows for the synthesis of oligonucleotides at a relatively low cost and with an error rate of 1 per 600 base pairs or less on average. However, it does not provide sufficient throughput for mass synthesis of oligonucleotides (Ma et al., 2012). Furthermore, if the sequence of interest exceeds 200 base pairs (some estimates suggest 300 (Palluk et al., 2018)), an additional assembly step via molecular cloning is required (). These factors significantly limit the speed of testing and represent the primary bottleneck in experimental design.
This problem can be solved by integrating higher-throughput oligonucleotide microarray synthesisers into laboratory practice (Song et al., 2021). Commercially available technologies are also based on phosphoramidite synthesis, albeit with slight modifications. Although microarray-based nucleotide synthesis is more error-prone due to heterogeneity and edge effects, it enables the synthesis of oligonucleotide pools and also reduces the cost per nucleotide by 2–4 orders of magnitude compared to column synthesis (Kosuri and Church, 2014). This suggests that advances in de novo DNA synthesis and experimental verification of codon-optimized sequences are likely to be associated with the microarray approach.
Since 2020, a trend towards an increase in the proportion of codon-optimized studies has been observed. In 2020, 1 in 34 (2.9%) clinical trials used codon optimization, compared to 4 in 42 (9.5%) in the first 11 months of 2023 (Figure 4). The main aim of codon optimization was to increase the level of transgene expression and the stability of the mRNA. In addition, a study using codon optimization to reduce immunogenicity was reported in 2021.
To effectively achieve the goals of codon optimization in research, it is important to follow established metrics. However, today there is no single generally accepted standard for codon optimization. Therefore, it is possible to use a large number of combinations of the methods described above to create optimal RNA variants. Some of these approaches significantly increase the efficacy of gene therapeutics. Therefore, several drug options have been registered in clinical trials, for example.
Codon optimization has played an important role in the development of RNA-based COVID-19 vaccines. Current research efforts are focused on further advancing the field of codon optimization for COVID-19 vaccines to address new strains of the coronavirus (Wu et al., 2023). Unfortunately, it was not possible to provide here the specific metrics used for codon optimization in the above-mentioned studies for commercial product development. This limitation results from the intellectual property of the original codon-optimized constructs. In this article, we have explored various metrics for assessing codon usage, based on both the composition of the coding sequence and the composition of a reference set of genes. One widely used metric is the Codon Adaptation Index (CAI). Although these measures provide useful information about adaptation to the host organism, they do not necessarily indicate an increase in translational efficiency due to selection pressure (Rahman et al., 2018; ). Furthermore, CAI is also interpreted as an indicator of the speed of translational elongation (Kudla et al., 2009). In turn, an increase in translation speed may not necessarily result in the production of a protein with similar properties in greater quantities.
Apparently, during translation, the most important regions for codon optimization are the areas around the start codon. This is supported by work demonstrating the contribution of the CDS position near the start codon (Höllerer and Jeschek, 2023; Nieuwkoop et al., 2023) and the 5′UTR sequence region (). The efficiency of translation is significantly dependent on the energy of mRNA folding, particularly in the vicinity of the start codon (). This is associated with the fact that unfolding more stable RNA secondary structures require greater energy before the initiation of translation (Figure 5). Additionally, the presence of hairpin, stem-loop, and pseudoknot structures in mRNA can hinder ribosome translocation and tRNA binding, thus impeding translation elongation (Kozak, 2005; ).
FIGURE 5
Thus, advancements in gene therapy could be directed towards a more comprehensive exploration of the impact of codon optimization on the characteristics and secondary structure of mRNA.Also, it is possible to apply optimization metrics locally to the start region, but there are limitations since many of them are based on codon usage frequency without taking into account the features of untranslated regions.
In addition, consideration of local codon optimization is a critical aspect that must be taken into account during codon optimization for a particular protein of interest. Furthermore, essential protein functions may change due to the possible influence of codon optimization on the conformation of the resulting protein, which should also be taken into account.
Statements
Author contributions
AP: Writing–original draft, Writing–review and editing. AK: Writing–original draft, Writing–review and editing. AM: Writing–original draft, Writing–review and editing. DN: Writing–original draft, Writing–review and editing. AS: Writing–original draft, Writing–review and editing. IA: Writing–original draft, Writing–review and editing. SF: Writing–original draft, Writing–review and editing. OM: Writing–original draft, Writing–review and editing. AD: Writing–original draft, Writing–review and editing. PV: Writing–original draft.
Funding
The author(s) declare that financial support was received for the research, authorship, and/or publication of this article. This study was supported by the Russian Science Foundation (Grant No. 23-64-00002).
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
References
1
AlexakiA.HettiarachchiG. K.AtheyJ. C.KatneniU. K.SimhadriV.Hamasaki-KatagiriN.et al (2019a). Effects of codon optimization on coagulation factor IX translation and structure: implications for protein and gene therapies. Sci. Rep.9, 15449. 10.1038/s41598-019-51984-2
2
AlexakiA.KamesJ.HolcombD. D.AtheyJ.Santana-QuinteroL. V.LamP. V. N.et al (2019b). Codon and codon-pair usage tables (CoCoPUTs): facilitating genetic variation analyses and recombinant gene design. J. Mol. Biol.431, 2434–2441. 10.1016/j.jmb.2019.04.021
3
AndersonB. R.MuramatsuH.NallagatlaS. R.BevilacquaP. C.SansingL. H.WeissmanD.et al (2010). Incorporation of pseudouridine into mRNA enhances translation by diminishing PKR activation. Nucleic Acids Res.38, 5884–5892. 10.1093/nar/gkq347
4
AnwarA. M.KhodaryS. M.AhmedE. A.OsamaA.EzzeldinS.TaniosA.et al (2023). gtAI: an improved species-specific tRNA adaptation index using the genetic algorithm. Front. Mol. Biosci.10, 1218518. 10.3389/fmolb.2023.1218518
5
AthanasopoulosT.FosterH.FosterK.DicksonG. (2011). Codon optimization of the microdystrophin gene for Duchene muscular dystrophy gene therapy. Gene Ther.709, 21–37. 10.1007/978-1-61737-982-6_2
6
AtheyJ.AlexakiA.OsipovaE.RostovtsevA.Santana-QuinteroL. V.KatneniU.et al (2017). A new and updated resource for codon usage tables. BMC Bioinforma.18, 391. 10.1186/s12859-017-1793-7
7
AyyarB. V.AroraS.RaviS. S. (2017). Optimizing antibody expression: the nuts and bolts. Methods116, 51–62. 10.1016/j.ymeth.2017.01.009
8
BainbridgeJ. W. B.SmithA. J.BarkerS. S.RobbieS.HendersonR.BalagganK.et al (2008). Effect of gene therapy on visual function in leber’s congenital amaurosis. N. Engl. J. Med.358, 2231–2239. 10.1056/NEJMoa0802268
9
BansalS.PerincheriS.FlemingT.PoulsonC.TiffanyB.BremnerR. M.et al (2021). Cutting edge: circulating exosomes with covid spike protein are induced by BNT162b2 (Pfizer–BioNTech) vaccination prior to development of antibodies: a novel mechanism for immune activation by mRNA vaccines. J. Immunol.207, 2405–2410. 10.4049/jimmunol.2100637
10
BaoC.LoerchS.LingC.KorostelevA. A.GrigorieffN.ErmolenkoD. N. (2020). mRNA stem-loops can pause the ribosome by hindering A-site tRNA binding. Elife9, e55799. 10.7554/eLife.55799
11
BellP.WangL.ChenS.-J.YuH.ZhuY.NayalM.et al (2016). Effects of self-complementarity, codon optimization, transgene, and dose on liver transduction with AAV8. Hum. Gene Ther. Methods27, 228–237. 10.1089/hgtb.2016.039
12
BennetzenJ. L.HallB. D. (1982). Codon selection in yeast. J. Biol. Chem.257, 3026–3031. 10.1016/S0021-9258(19)81068-2
13
BermanH. M. (2000). The protein Data Bank. Nucleic Acids Res.28, 235–242. 10.1093/nar/28.1.235
14
BertoliniT. B.ShirleyJ. L.ZolotukhinI.LiX.KaishoT.XiaoW.et al (2021). Effect of CpG depletion of vector genome on CD8+ T cell responses in AAV gene therapy. Front. Immunol.12, 672449. 10.3389/fimmu.2021.672449
15
BłażejP.WnętrzakM.MackiewiczD.MackiewiczP. (2018). Optimization of the standard genetic code according to three codon positions using an evolutionary algorithm. PLoS One13, e0201715. 10.1371/journal.pone.0201715
16
BodeC.ZhaoG.SteinhagenF.KinjoT.KlinmanD. M. (2011). CpG DNA as a vaccine adjuvant. Expert Rev. Vaccines10, 499–511. 10.1586/erv.10.174
17
BollmanB.NunnaN.BahlK.HsiaoC. J.BennettH.ButlerS.et al (2023). An optimized messenger RNA vaccine candidate protects non-human primates from Zika virus infection. npj Vaccines8, 58. 10.1038/s41541-023-00656-4
18
BourretJ.AlizonS.BravoI. G. (2019). COUSIN (COdon usage similarity INdex): a normalized measure of codon usage preferences. Genome Biol. Evol.11, 3523–3528. 10.1093/gbe/evz262
19
BoutinS.MonteilhetV.VeronP.LeborgneC.BenvenisteO.MontusM. F.et al (2010). Prevalence of serum IgG and neutralizing factors against adeno-associated virus (AAV) types 1, 2, 5, 6, 8, and 9 in the healthy population: implications for gene therapy using AAV vectors. Hum. Gene Ther.21, 704–712. 10.1089/hum.2009.182
20
BreckpotK.EscorsD.ArceF.LopesL.KarwaczK.Van LintS.et al (2010). HIV-1 lentiviral vector immunogenicity is mediated by toll-like receptor 3 (TLR3) and TLR7. J. Virol.84, 5627–5636. 10.1128/JVI.00014-10
21
BuchanJ. R. (2006). tRNA properties help shape codon pair preferences in open reading frames. Nucleic Acids Res.34, 1015–1027. 10.1093/nar/gkj488
22
BuhrF.JhaS.ThommenM.MittelstaetJ.KutzF.SchwalbeH.et al (2016). Synonymous codons direct cotranslational folding toward different protein conformations. Mol. Cell61, 341–351. 10.1016/j.molcel.2016.01.008
23
BulchaJ. T.WangY.MaH.TaiP. W. L.GaoG. (2021). Viral vector platforms within the gene therapy landscape. Signal Transduct. Target. Ther.6, 53. 10.1038/s41392-021-00487-6
24
BurkeP. C.ParkH.SubramaniamA. R. (2022). A nascent peptide code for translational control of mRNA stability in human cells. Nat. Commun.13, 6829. 10.1038/s41467-022-34664-0
25
BurnsC. C.ShawJ.CampagnoliR.JorbaJ.VincentA.QuayJ.et al (2006). Modulation of poliovirus replicative fitness in HeLa cells by deoptimization of synonymous codon usage in the capsid region. J. Virol.80, 3259–3272. 10.1128/JVI.80.7.3259-3272.2006
26
Cabanes-CreusM.GinnS. L.AmayaA. K.LiaoS. H. Y.WesthausA.HallwirthC. V.et al (2019). Codon-optimization of wild-type adeno-associated virus capsid sequences enhances DNA family shuffling while conserving functionality. Mol. Ther. - Methods Clin. Dev.12, 71–84. 10.1016/j.omtm.2018.10.016
27
CapellA.FellererK.HaassC. (2014). Progranulin transcripts with Short and long 5′ untranslated regions (UTRs) are differentially expressed via posttranscriptional and translational repression. J. Biol. Chem.289, 25879–25889. 10.1074/jbc.M114.560128
28
CarboneA.ZinovyevA.KépèsF. (2003). Codon adaptation index as a measure of dominating codon bias. Bioinformatics19, 2005–2015. 10.1093/bioinformatics/btg272
29
CasiniA.StorchM.BaldwinG. S.EllisT. (2015). Bricks and blueprints: methods and standards for DNA assembly. Nat. Rev. Mol. Cell Biol.16, 568–576. 10.1038/nrm4014
30
Castillo-HairS. M.SeeligG. (2022). Machine learning for designing next-generation mRNA therapeutics. Acc. Chem. Res.55, 24–34. 10.1021/acs.accounts.1c00621
31
Chamani MohassesF.SoloukiM.GhareyazieB.FahmidehL.MohsenpourM. (2020). Correlation between gene expression levels under drought stress and synonymous codon usage in rice plant by in-silico study. PLoS One15, e0237334. 10.1371/journal.pone.0237334
32
ChenK. Y.ParkH.SubramaniamA. R. (2023). Massively parallel identification of sequence motifs triggering ribosome-associated mRNA quality control. bioRxiv, 2023.09.27.2023.09.27.559793. 10.1101/2023.09.27.559793
33
ChenM.-W.ChengT.-J. R.HuangY.JanJ.-T.MaS.-H.YuA. L.et al (2008). A consensus–hemagglutinin-based DNA vaccine that protects mice against divergent H5N1 influenza viruses. Proc. Natl. Acad. Sci.105, 13538–13543. 10.1073/pnas.0806901105
34
ChenW.LiH.LiuZ.YuanW. (2016). Lipopolyplex for therapeutic gene delivery and its application for the treatment of Parkinson’s disease. Front. Aging Neurosci.8, 68. 10.3389/fnagi.2016.00068
35
ClinicalTrials.govClinicalTrials.gov (2024).
36
CoughlanL. (2020). Factors which contribute to the immunogenicity of non-replicating adenoviral vectored vaccines. Front. Immunol.11, 909. 10.3389/fimmu.2020.00909
37
CourelM.ClémentY.BossevainC.ForetekD.Vidal CruchezO.YiZ.et al (2019). GC content shapes mRNA storage and decay in human cells. Elife8, e49708. 10.7554/eLife.49708
38
DanielE.OnwukweG. U.WierengaR. K.QuagginS. E.VainioS. J.KrauseM. (2015). ATGme: open-source web application for rare codon identification and custom DNA sequence optimization. BMC Bioinforma.16, 303. 10.1186/s12859-015-0743-5
39
DasS. (2017). Analysis of gene expression using modified relative codon bias strength in nanoarchaeum equitans. Biosci. Biotechnol. Res. Asia14, 793–799. 10.13005/bbra/2510
40
DesaiP. N.ShrivastavaN.PadhH. (2010). Production of heterologous proteins in plants: strategies for optimal expression. Biotechnol. Adv.28, 427–435. 10.1016/j.biotechadv.2010.01.005
41
de SostoaJ.FajardoC. A.MorenoR.RamosM. D.Farrera-SalM.AlemanyR. (2019). Targeting the tumor stroma with an oncolytic adenovirus secreting a fibroblast activation protein-targeted bispecific T-cell engager. J. Immunother. Cancer7, 19. 10.1186/s40425-019-0505-4
42
DewiK. S.FuadA. M. (2020). Improving the expression of human granulocyte colony stimulating factor in Escherichia coli by reducing the GC-content and increasing mRNA folding free energy at 5’-terminal end. Adv. Pharm. Bull.10, 610–616. 10.34172/apb.2020.073
43
DiezM.Medina-MuñozS. G.CastellanoL. A.da Silva PescadorG.WuQ.BazziniA. A. (2022). iCodon customizes gene expression based on the codon composition. Sci. Rep.12, 12126. 10.1038/s41598-022-15526-7
44
DittmarK. A.GoodenbourJ. M.PanT. (2006). Tissue-specific differences in human transfer RNA expression. PLoS Genet.2, e221. 10.1371/journal.pgen.0020221
45
dos ReisM. (2003). Unexpected correlations between gene expression and codon usage bias from microarray data for the whole Escherichia coli K-12 genome. Nucleic Acids Res.31, 6976–6985. 10.1093/nar/gkg897
46
FathS.BauerA. P.LissM.SpriestersbachA.MaertensB.HahnP.et al (2011). Multiparameter RNA and codon optimization: a standardized tool to assess and enhance autologous mammalian gene expression. PLoS One6, e17596. 10.1371/journal.pone.0017596
47
FaustS. M.BellP.CutlerB. J.AshleyS. N.ZhuY.RabinowitzJ. E.et al (2013). CpG-depleted adeno-associated virus vectors evade immune detection. J. Clin. Invest.123, 2994–3001. 10.1172/JCI68205
48
FDA (2024).
49
FengH.SegalésJ.WangF.JinQ.WangA.ZhangG.et al (2022). Comprehensive analysis of codon usage patterns in Chinese porcine circoviruses based on their major protein-coding sequences. Viruses14, 81. 10.3390/v14010081
50
FestenE. A. M.GoyetteP.GreenT.BoucherG.BeauchampC.TrynkaG.et al (2011). A meta-analysis of genome-wide association scans identifies IL18RAP, PTPN2, TAGAP, and PUS10 as shared risk loci for crohn’s disease and celiac disease. PLoS Genet.7, e1001283. 10.1371/journal.pgen.1001283
51
FoxJ. M.ErillI. (2010). Relative codon adaptation: a generic codon bias index for prediction of gene expression. DNA Res.17, 185–196. 10.1093/dnares/dsq012
52
FribergM.von RohrP.GonnetG. (2004). Limitations of codon adaptation index and other coding DNA‐based features for prediction of protein expression in Saccharomyces cerevisiae. Yeast21, 1083–1093. 10.1002/yea.1150
53
FuH.LiangY.ZhongX.PanZ.HuangL.ZhangH.et al (2020). Codon optimization with deep learning to enhance protein expression. Sci. Rep.10, 17617–17619. 10.1038/s41598-020-74091-z
54
GaoW.Gallardo-DoddC. J.KutterC. (2022). Cell type–specific analysis by single-cell profiling identifies a stable mammalian tRNA–mRNA interface and increased translation efficiency in neurons. Genome Res.32, 97–110. 10.1101/gr.275944.121
55
Godfried SieC.HeslerS.MaasS.KuchkaM. (2012). IGFBP7’s susceptibility to proteolysis is altered by A‐to‐I RNA editing of its transcript. FEBS Lett.586, 2313–2317. 10.1016/j.febslet.2012.06.037
56
Gonzalez-SanchezB.Vega-RodríguezM. A.Santander-JiménezS.Granado-CriadoJ. M. (2019). Multi-Objective Artificial Bee Colony for designing multiple genes encoding the same protein. Appl. Soft Comput.74, 90–98. 10.1016/j.asoc.2018.10.023
57
GouletD. R.YanY.AgrawalP.WaightA. B.MakA. N.ZhuY. (2023). Codon optimization using a recurrent neural network. J. Comput. Biol.30, 70–81. 10.1089/cmb.2021.0458
58
GouyM.GautierC. (1982). Codon usage in bacteria: correlation with gene expressivity. Nucleic Acids Res.10, 7055–7074. 10.1093/nar/10.22.7055
59
GroteA.HillerK.ScheerM.MunchR.NortemannB.HempelD. C.et al (2005). JCat: a novel tool to adapt codon usage of a target gene to its potential expression host. Nucleic Acids Res.33, W526–W531. 10.1093/nar/gki376
60
GuW.ZhouT.WilkeC. O. (2010). A universal trend of reduced mRNA stability near the translation-initiation site in prokaryotes and eukaryotes. PLoS Comput. Biol.6, e1000664. 10.1371/journal.pcbi.1000664
61
HansonG.CollerJ. (2018). Codon optimality, bias and usage in translation and mRNA decay. Nat. Rev. Mol. Cell Biol.19, 20–30. 10.1038/nrm.2017.91
62
HayatS. M. G.FarahaniN.SafdarianE.RoointanA.SahebkarA. (2019). Gene delivery using lipoplexes and polyplexes: principles, limitations and solutions. Crit. Rev. Eukaryot. Gene Expr.29, 29–36. 10.1615/CritRevEukaryotGeneExpr.2018025132
63
Hernandez-AliasX.BenistyH.RaduskyL. G.SerranoL.SchaeferM. H. (2023). Using protein-per-mRNA differences among human tissues in codon optimization. Genome Biol.24, 34–20. 10.1186/s13059-023-02868-2
64
HerzogR. W.CooperM.PerrinG. Q.BiswasM.MartinoA. T.MorelL.et al (2019). Regulatory T cells and TLR9 activation shape antibody formation to a secreted transgene product in AAV muscle gene transfer. Cell. Immunol.342, 103682. 10.1016/j.cellimm.2017.07.012
65
HiaF.YangS. F.ShichinoY.YoshinagaM.MurakawaY.VandenbonA.et al (2019). Codon bias confers stability to human mRNA s. EMBO Rep.20, e48220. 10.15252/embr.201948220
66
HöllererS.JeschekM. (2023). Ultradeep characterisation of translational sequence determinants refutes rare-codon hypothesis and unveils quadruplet base pairing of initiator tRNA and transcript. Nucleic Acids Res.51, 2377–2396. 10.1093/nar/gkad040
67
HuangH.LiuY.LiaoW.CaoY.LiuQ.GuoY.et al (2019). Oncolytic adenovirus programmed by synthetic gene circuit for cancer immunotherapy. Nat. Commun.10, 4801. 10.1038/s41467-019-12794-2
68
IgyártóB. Z.QinZ. (2024). The mRNA-LNP vaccines – the good, the bad and the ugly?Front. Immunol.15, 1336906. 10.3389/fimmu.2024.1336906
69
IkemuraT. (1981). Correlation between the abundance of Escherichia coli transfer RNAs and the occurrence of the respective codons in its protein genes: a proposal for a synonymous codon choice that is optimal for the E. coli translational system. J. Mol. Biol.151, 389–409. 10.1016/0022-2836(81)90003-6
70
IkemuraT. (1982). Correlation between the abundance of yeast transfer RNAs and the occurrence of the respective codons in protein genes. J. Mol. Biol.158, 573–597. 10.1016/0022-2836(82)90250-9
71
IrimiaM.DenucA.FerranJ. L.PernauteB.PuellesL.RoyS. W.et al (2012). Evolutionarily conserved A-to-I editing increases protein stability of the alternative splicing factor Nova1. RNA Biol.9, 12–21. 10.4161/rna.9.1.18387
72
JainR.JainA.MauroE.LeShaneK.DensmoreD. (2023). ICOR: improving codon optimization with recurrent neural networks. BMC Bioinforma.24, 132. 10.1186/s12859-023-05246-8
73
KamesJ.AlexakiA.HolcombD. D.Santana-QuinteroL. V.AtheyJ. C.Hamasaki-KatagiriN.et al (2020). TissueCoCoPUTs: novel human tissue-specific codon and codon-pair usage tables based on differential tissue gene expression. J. Mol. Biol.432, 3369–3378. 10.1016/j.jmb.2020.01.011
74
KarikóK.BucksteinM.NiH.WeissmanD. (2005). Suppression of RNA recognition by toll-like receptors: the impact of nucleoside modification and the evolutionary origin of RNA. Immunity23, 165–175. 10.1016/j.immuni.2005.06.008
75
KarikóK.MuramatsuH.WelshF. A.LudwigJ.KatoH.AkiraS.et al (2008). Incorporation of pseudouridine into mRNA yields superior nonimmunogenic vector with increased translational capacity and biological stability. Mol. Ther.16, 1833–1840. 10.1038/mt.2008.200
76
KirchnerS.CaiZ.RauscherR.KastelicN.AndingM.CzechA.et al (2017). Alteration of protein function by a silent polymorphism linked to tRNA abundance. PLOS Biol.15, e2000779. 10.1371/journal.pbio.2000779
77
KonkleB. A.WalshC. E.EscobarM. A.JosephsonN. C.YoungG.von DrygalskiA.et al (2021). BAX 335 hemophilia B gene therapy clinical trial results: potential impact of CpG sequences on gene expression. Blood137, 763–774. 10.1182/blood.2019004625
78
KosuriS.ChurchG. M. (2014). Large-scale de novo DNA synthesis: technologies and applications. Nat. Methods11, 499–507. 10.1038/nmeth.2918
79
KozakM. (2005). Regulation of translation via mRNA structure in prokaryotes and eukaryotes. Gene361, 13–37. 10.1016/j.gene.2005.06.037
80
KudlaG.LipinskiL.CaffinF.HelwakA.ZyliczM. (2006). High guanine and cytosine content increases mRNA levels in mammalian cells. PLoS Biol.4, e180. 10.1371/journal.pbio.0040180
81
KudlaG.MurrayA. W.TollerveyD.PlotkinJ. B. (2009). Coding-sequence determinants of gene expression in Escherichia coli. Science324, 255–258. 10.1126/science.1170160
82
LeeI. T.NachbagauerR.EnszD.SchwartzH.CarmonaL.SchaefersK.et al (2023). Safety and immunogenicity of a phase 1/2 randomized clinical trial of a quadrivalent, mRNA-based seasonal influenza vaccine (mRNA-1010) in healthy adults: interim analysis. Nat. Commun.14, 3631. 10.1038/s41467-023-39376-7
83
LeppekK.ByeonG. W.KladwangW.Wayment-SteeleH. K.KerrC. H.XuA. F.et al (2022). Combinatorial optimization of mRNA structure, stability, and translation for RNA-based therapeutics. Nat. Commun.13, 1536. 10.1038/s41467-022-28776-w
84
LeRoyN.RoleckC. (2023). Optipyzer: a fast and flexible multi-species codon optimization server. bioRxiv, 2023.05.22.541759. 10.1101/2023.05.22.541759
85
LiC.HeY.NicolsonS.HirschM.WeinbergM. S.ZhangP.et al (2013). Adeno-associated virus capsid antigen presentation is dependent on endosomal escape. J. Clin. Invest.123, 1390–1401. 10.1172/JCI66611
86
LiuY. (2020). A code within the genetic code: codon usage regulates co-translational protein folding. Cell Commun. Signal.18, 145. 10.1186/s12964-020-00642-6
87
LoomisK. H.KirschmanJ. L.BhosleS.BellamkondaR. V.SantangeloP. J. (2016). Strategies for modulating innate immune activation and protein production of in vitro transcribed mRNAs. J. Mater. Chem. B4, 1619–1632. 10.1039/C5TB01753J
88
MaS.TangN.TianJ. (2012). DNA synthesis, assembly and applications in synthetic biology. Curr. Opin. Chem. Biol.16, 260–267. 10.1016/j.cbpa.2012.05.001
89
MalarkannanS.HorngT.ShihP. P.SchwabS.ShastriN. (1999). Presentation of out-of-frame peptide/MHC class I complexes by a novel translation initiation mechanism. Immunity10, 681–690. 10.1016/S1074-7613(00)80067-9
90
MartinoA. T.Basner-TschakarjanE.MarkusicD. M.FinnJ. D.HindererC.ZhouS.et al (2013). Engineered AAV vector minimizes in vivo targeting of transduced hepatocytes by capsid-specific CD8+ T cells. Blood121, 2224–2233. 10.1182/blood-2012-10-460733
91
MatsudaD.MauroV. P. (2010). Determinants of initiation codon selection during translation in mammalian cells. PLoS One5, e15057. 10.1371/journal.pone.0015057
92
MendellJ. R.Al-ZaidyS. A.Rodino-KlapacL. R.GoodspeedK.GrayS. J.KayC. N.et al (2021). Current clinical applications of in vivo gene therapy with AAVs. Mol. Ther.29, 464–488. 10.1016/j.ymthe.2020.12.007
93
MitaraiN.SneppenK.PedersenS. (2008). Ribosome collisions and translation efficiency: optimization by codon usage and mRNA destabilization. J. Mol. Biol.382, 236–245. 10.1016/j.jmb.2008.06.068
94
MuellerS. (2023). Challenges and opportunities of mRNA vaccines against SARS-CoV-2. Cham: Springer International Publishing. 10.1007/978-3-031-18903-6
95
MuellerS.PapamichailD.ColemanJ. R.SkienaS.WimmerE. (2006). Reduction of the rate of poliovirus protein synthesis through large-scale codon deoptimization causes attenuation of viral virulence by lowering specific infectivity. J. Virol.80, 9687–9696. 10.1128/JVI.00738-06
96
MulroneyT. E.PöyryT.Yam-PucJ. C.RustM.HarveyR. F.KalmarL.et al (2024). N1-methylpseudouridylation of mRNA causes +1 ribosomal frameshifting. Nature625, 189–194. 10.1038/s41586-023-06800-3
97
NarulaA.EllisJ.TaliaferroJ. M.RisslandO. S. (2019). Coding regions affect mRNA stability in human cells. RNA25, 1751–1764. 10.1261/rna.073239.119
98
NavonS.PilpelY. (2011). The role of codon selection in regulation of translation efficiency deduced from synthetic libraries. Genome Biol.12, R12. 10.1186/gb-2011-12-2-r12
99
NieuwkoopT.TerlouwB. R.StevensK. G.ScheltemaR. A.de RidderD.van der OostJ.et al (2023). Revealing determinants of translation efficiency via whole-gene codon randomization and machine learning. Nucleic Acids Res.51, 2363–2376. 10.1093/nar/gkad035
100
Núñez-ManchónE.Farrera-SalM.Otero-MateoM.CastellanoG.MorenoR.MedelD.et al (2021). Transgene codon usage drives viral fitness and therapeutic efficacy in oncolytic adenoviruses. Nar. Cancer3, zcab015. 10.1093/narcan/zcab015
101
OliverS. E.GarganoJ. W.MarinM.WallaceM.CurranK. G.ChamberlandM.et al (2020). The advisory committee on immunization practices’ interim recommendation for use of pfizer-BioNTech COVID-19 vaccine — United States, december 2020. MMWR. Morb. Mortal. Wkly. Rep.69, 1922–1924. 10.15585/mmwr.mm6950e2
102
OwczarzyR.TataurovA. V.WuY.MantheyJ. A.McQuistenK. A.AlmabraziH. G.et al (2008). IDT SciTools: a suite for analysis and design of nucleic acid oligomers. Nucleic Acids Res.36, W163–W169. 10.1093/nar/gkn198
103
PallukS.ArlowD. H.de RondT.BarthelS.KangJ. S.BectorR.et al (2018). De novo DNA synthesis using polymerase-nucleotide conjugates. Nat. Biotechnol.36, 645–650. 10.1038/nbt.4173
104
PereiraI. T.SpangenbergL.RobertA. W.AmorínR.StimamiglioM. A.NayaH.et al (2018). Polysome profiling followed by RNA-seq of cardiac differentiation stages in hESCs. Sci. Data5, 180287. 10.1038/sdata.2018.287
105
PerlakF. J.FuchsR. L.DeanD. A.McPhersonS. L.FischhoffD. A. (1991). Modification of the coding sequence enhances plant expression of insect control protein genes. Proc. Natl. Acad. Sci.88, 3324–3328. 10.1073/pnas.88.8.3324
106
PhamT. D.O’ConnellJ.CraneD. I. (2004). “Constrained codon optimization by dynamic programming,” in Proceedings of 2004 International Symposium on Intelligent Multimedia, Video and Speech Processing, 2004 (IEEE), 153–156. 10.1109/ISIMP.2004.1434023
107
PinkardO.McFarlandS.SweetT.CollerJ. (2020). Quantitative tRNA-sequencing uncovers metazoan tissue-specific tRNA regulation. Nat. Commun.11, 4104. 10.1038/s41467-020-17879-x
108
PitoisetF.VazquezT.LevacherB.Nehar-BelaidD.DérianN.VigneronJ.et al (2017). Retrovirus-based virus-like particle immunogenicity and its modulation by toll-like receptor activation. J. Virol.91, e01230-17. 10.1128/JVI.01230-17
109
PizzoL.IriarteA.Alvarez-ValinF.MarínM. (2015). Conservation of CFTR codon frequency through primates suggests synonymous mutations could have a functional effect. Mutat. Res. Mol. Mech. Mutagen.775, 19–25. 10.1016/j.mrfmmm.2015.03.005
110
PlotkinJ. B.RobinsH.LevineA. J. (2004). Tissue-specific codon usage and the expression of human genes. Proc. Natl. Acad. Sci.101, 12588–12591. 10.1073/pnas.0404957101
111
PouyetF.MouchiroudD.DuretL.SémonM. (2017). Recombination, meiotic expression and human codon usage. Elife6, e27344. 10.7554/eLife.27344
112
PresnyakV.AlhusainiN.ChenY.-H.MartinS.MorrisN.KlineN.et al (2015). Codon optimality is a major determinant of mRNA stability. Cell160, 1111–1124. 10.1016/j.cell.2015.02.029
113
PuigboP.GuzmanE.RomeuA.Garcia-VallveS. (2007). OPTIMIZER: a web server for optimizing the codon usage of DNA sequences. Nucleic Acids Res.35, W126–W131. 10.1093/nar/gkm219
114
RaabA. M.GebhardtG.BolotinaN.Weuster-BotzD.LangC. (2010). Metabolic engineering of Saccharomyces cerevisiae for the biotechnological production of succinic acid. Metab. Eng.12, 518–525. 10.1016/j.ymben.2010.08.005
115
RahmanS. U.YaoX.LiX.ChenD.TaoS. (2018). Analysis of codon usage bias of Crimean-Congo hemorrhagic fever virus and its adaptation to hosts. Infect. Genet. Evol.58, 1–16. 10.1016/j.meegid.2017.11.027
116
ReisM. d. (2004). Solving the riddle of codon usage preferences: a test for translational selection. Nucleic Acids Res.32, 5036–5044. 10.1093/nar/gkh834
117
RingnérM.KroghM. (2005). Folding free energies of 5′-UTRs impact post-transcriptional regulation on a genomic scale in yeast. PLoS Comput. Biol.1, e72. 10.1371/journal.pcbi.0010072
118
RodriguezA.WrightG.EmrichS.ClarkP. L. (2018). %MinMax: a versatile tool for calculating and comparing synonymous codon usage and its impact on protein folding. Protein Sci.27, 356–362. 10.1002/pro.3336
119
RogersG. L.ShirleyJ. L.ZolotukhinI.KumarS. R. P.ShermanA.PerrinG. Q.et al (2017). Plasmacytoid and conventional dendritic cells cooperate in crosspriming AAV capsid-specific CD8+ T cells. Blood129, 3184–3195. 10.1182/blood-2016-11-751040
120
RojasM.Restrepo-JiménezP.MonsalveD. M.PachecoY.Acosta-AmpudiaY.Ramírez-SantanaC.et al (2018). Molecular mimicry and autoimmunity. J. Autoimmun.95, 100–123. 10.1016/j.jaut.2018.10.012
121
RöltgenK.NielsenS. C. A.SilvaO.YounesS. F.ZaslavskyM.CostalesC.et al (2022). Immune imprinting, breadth of variant recognition, and germinal center response in human SARS-CoV-2 infection and vaccination. Cell185, 1025–1040.e14. 10.1016/j.cell.2022.01.018
122
RonkA. J.LloydN. M.ZhangM.AtyeoC.PerrettH. R.MireC. E.et al (2023). A Lassa virus mRNA vaccine confers protection but does not require neutralizing antibody in a Guinea pig model of infection. Nat. Commun.14, 5603. 10.1038/s41467-023-41376-6
123
RoymondalU.DasS.SahooS. (2009). Predicting gene expression level from relative codon usage bias: an application to Escherichia coli genome. DNA Res.16, 13–30. 10.1093/dnares/dsn029
124
SabiR.TullerT. (2014). Modelling the efficiency of codon–tRNA interactions based on codon usage bias. DNA Res.21, 511–526. 10.1093/dnares/dsu017
125
SabiR.Volvovitch DanielR.TullerT. (2017). stAIcalc: tRNA adaptation index calculator based on species-specific weights. Bioinformatics33, 589–591. 10.1093/bioinformatics/btw647
126
SatoK.HamadaM. (2023). Recent trends in RNA informatics: a review of machine learning and deep learning for RNA secondary structure prediction and RNA drug discovery. Brief. Bioinform.24, bbad186. 10.1093/bib/bbad186
127
SharpP.LiW.-H. (1986). Codon usage in regulatory genes in Escherichia coli does not reflect selection for ‘rare’ codons. Nucleic Acids Res.14, 7737–7749. 10.1093/nar/14.19.7737
128
SharpP. M.LiW.-H. (1987). The codon adaptation index-a measure of directional synonymous codon usage bias, and its potential applications. Nucleic Acids Res.15, 1281–1295. 10.1093/nar/15.3.1281
129
ShiF.FanZ.ZhangS.WangY.TanS.LiY. (2020). Optimization of ribosomal binding site sequences for gene expression and 4-hydroxyisoleucine biosynthesis in recombinant corynebacterium glutamicum. Enzyme Microb. Technol.140, 109622. 10.1016/j.enzmictec.2020.109622
130
ShirleyJ. L.de JongY. P.TerhorstC.HerzogR. W. (2020a). Immune responses to viral gene therapy vectors. Mol. Ther.28, 709–722. 10.1016/j.ymthe.2020.01.001
131
ShirleyJ. L.KeelerG. D.ShermanA.ZolotukhinI.MarkusicD. M.HoffmanB. E.et al (2020b). Type I IFN sensing by cDCs and CD4+ T cell help are both requisite for cross-priming of AAV capsid-specific CD8+ T cells. Mol. Ther.28, 758–770. 10.1016/j.ymthe.2019.11.011
132
SimonC. S.HadjantonakisA.SchröterC. (2018). Making lineage decisions with biological noise: lessons from the early mouse embryo. WIREs Dev. Biol.7, e319. 10.1002/wdev.319
133
SinyakovA. N.RyabininV. A.KostinaE. V. (2021). Application of array-based oligonucleotides for synthesis of genetic designs. Mol. Biol.55, 487–500. 10.1134/S0026893321030109
134
SongL.-F.DengZ.-H.GongZ.-Y.LiL.-L.LiB.-Z. (2021). Large-scale de novo oligonucleotide synthesis for whole-genome synthesis and data storage: challenges and opportunities. Front. Bioeng. Biotechnol.9, 689797. 10.3389/fbioe.2021.689797
135
StenicoM.LloydA. T.SharpP. M. (1994). Codon usage in Caenorhabditis elegans: delineation of translational selection and mutational biases. Nucleic Acids Res.22, 2437–2446. 10.1093/nar/22.13.2437
136
SunB.ZhangH.FrancoL. M.BrownT.BirdA.SchneiderA.et al (2005). Correction of glycogen storage disease type II by an adeno-associated virus vector containing a muscle-specific promoter. Mol. Ther.11, 889–898. 10.1016/j.ymthe.2005.01.012
137
TanedaA.AsaiK. (2020). COSMO: a dynamic programming algorithm for multicriteria codon optimization. Comput. Struct. Biotechnol. J.18, 1811–1818. 10.1016/j.csbj.2020.06.035
138
ThessA.GrundS.MuiB. L.HopeM. J.BaumhofP.Fotin-MleczekM.et al (2015). Sequence-engineered mRNA without chemical nucleoside modifications enables an effective protein therapy in large animals. Mol. Ther.23, 1456–1464. 10.1038/mt.2015.103
139
ThomasD. R.WalmsleyA. M. (2014). Improved expression of recombinant plant-made hEGF. Plant Cell Rep.33, 1801–1814. 10.1007/s00299-014-1658-8
140
ThomasO. G.BrongeM.TengvallK.AkpinarB.NilssonO. B.HolmgrenE.et al (2023). Cross-reactive EBNA1 immunity targets alpha-crystallin B and is associated with multiple sclerosis. Sci. Adv.9, eadg3032–14. 10.1126/sciadv.adg3032
141
ThulP. J.LindskogC. (2018). The human protein atlas: a spatial map of the human proteome. Protein Sci.27, 233–244. 10.1002/pro.3307
142
VillanuevaE.Martí-SolanoM.FillatC. (2016). Codon optimization of the adenoviral fiber negatively impacts structural protein expression and viral fitness. Sci. Rep.6, 27546. 10.1038/srep27546
143
WanJ.YangJ.WangZ.ShenR.ZhangC.WuY.et al (2023). A single immunization with core–shell structured lipopolyplex mRNA vaccine against rabies induces potent humoral immunity in mice and dogs. Emerg. Microbes Infect.12, 2270081. 10.1080/22221751.2023.2270081
144
WanX.-F.ZhouJ.XuD. (2006). CodonO: a new informatics method for measuring synonymous codon usage bias within and across genomes. Int. J. Gen. Syst.35, 109–125. 10.1080/03081070500502967
145
Wayment-SteeleH. K.KimD. S.ChoeC. A.NicolJ. J.Wellington-OguriR.WatkinsA. M.et al (2021). Theoretical basis for stabilizing messenger RNA through secondary structure design. Nucleic Acids Res.49, 10604–10617. 10.1093/nar/gkab764
146
WeiY.SilkeJ. R.XiaX. (2019). An improved estimation of tRNA expression to better elucidate the coevolution between tRNA abundance and codon usage in bacteria. Sci. Rep.9, 3184. 10.1038/s41598-019-39369-x
147
WelchM.VillalobosA.GustafssonC.MinshullJ. (2009). You’re one in a googol: optimizing genes for protein expression. J. R. Soc. Interface6, S467–S476. 10.1098/rsif.2008.0520.focus
148
WrightF. (1990). The ‘effective number of codons’ used in a gene. Gene87, 23–29. 10.1016/0378-1119(90)90491-9
149
WrightG.RodriguezA.LiJ.MilenkovicT.EmrichS. J.ClarkP. L. (2022). CHARMING: harmonizing synonymous codon usage to replicate a desired codon usage pattern. Protein Sci.31, 221–231. 10.1002/pro.4223
150
WrightJ. F. (2020). Quantification of CpG motifs in rAAV genomes: avoiding the Toll. Mol. Ther.28, 1756–1758. 10.1016/j.ymthe.2020.07.006
151
WuQ.MedinaS. G.KushawahG.DeVoreM. L.CastellanoL. A.HandJ. M.et al (2019). Translation affects mRNA stability in a codon-dependent manner in human cells. Elife8, e45396. 10.7554/eLife.45396
152
WuX.ShanK.ZanF.TangX.QianZ.LuJ. (2023). Optimization and deoptimization of codons in SARS‐CoV‐2 and related implications for vaccine development. Adv. Sci.10, e2205445. 10.1002/advs.202205445
153
XiaX. (2015). A major controversy in codon-anticodon adaptation resolved by a new codon usage index. Genetics199, 573–579. 10.1534/genetics.114.172106
154
XiaX. (2021). Detailed dissection and critical evaluation of the pfizer/BioNTech and Moderna mRNA vaccines. Vaccines9, 734. 10.3390/vaccines9070734
155
XueC.ChuQ.ZhengQ.JiangS.BaoZ.SuY.et al (2022). Role of main RNA modifications in cancer: N6-methyladenosine, 5-methylcytosine, and pseudouridine. Signal Transduct. Target. Ther.7, 142. 10.1038/s41392-022-01003-0
156
YangT. yuanBraunM.LembkeW.McBlaneF.KamerudJ.DeWallS.et al (2022). Immunogenicity assessment of AAV-based gene therapies: an IQ consortium industry white paper. Mol. Ther. - Methods Clin. Dev.26, 471–494. 10.1016/j.omtm.2022.07.018
157
YewN. S.ChengS. H. (2004). Reducing the immunostimulatory activity of CpG-containing plasmid DNA vectors for non-viral gene therapy. Expert Opin. Drug Deliv.1, 115–125. 10.1517/17425247.1.1.115
158
ZhangH.ZhangL.LinA.XuC.LiZ.LiuK.et al (2023). Algorithm for optimized mRNA design improves stability and immunogenicity. Nature621, 396–403. 10.1038/s41586-023-06127-z
159
ZhangZ.LiJ.CuiP.DingF.LiA.TownsendJ. P.et al (2012). Codon Deviation Coefficient: a novel measure for estimating codon usage bias and its statistical significance. BMC Bioinforma.13, 43. 10.1186/1471-2105-13-43
160
ZukerM. (1994). “Prediction of RNA secondary structure by energy minimization,” in Computer analysis of sequence data (Totowa, NJ: Humana Press), 267–294. 10.1385/0-89603-276-0:267
161
ZukerM.StieglerP. (1981). Optimal computer folding of large RNA sequences using thermodynamics and auxiliary information. Nucleic Acids Res.9, 133–148. 10.1093/nar/9.1.133
Summary
Keywords
gene therapy, codon-optimization metrics, mRNA, immunogenicity, clinical trials
Citation
Paremskaia AI, Kogan AA, Murashkina A, Naumova DA, Satish A, Abramov IS, Feoktistova SG, Mityaeva ON, Deviatkin AA and Volchkov PY (2024) Codon-optimization in gene therapy: promises, prospects and challenges. Front. Bioeng. Biotechnol. 12:1371596. doi: 10.3389/fbioe.2024.1371596
Received
16 January 2024
Accepted
19 March 2024
Published
28 March 2024
Volume
12 - 2024
Edited by
Yaroslava G. Yingling, North Carolina State University, United States
Reviewed by
Siguna Mueller, Independent Researcher, Kaernten, Austria
Clement T. Y. Chan, University of North Texas, United States
Updates
Copyright
© 2024 Paremskaia, Kogan, Murashkina, Naumova, Satish, Abramov, Feoktistova, Mityaeva, Deviatkin and Volchkov.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Andrei A. Deviatkin, andreideviatkin@gmail.com; Pavel Yu Volchkov, vpwwww@gmail.com
† These authors have contributed equally to this work and share last authorship
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.