REVIEW article

Front. Bioeng. Biotechnol., 28 March 2024

Sec. Synthetic Biology

Volume 12 - 2024 | https://doi.org/10.3389/fbioe.2024.1371596

Codon-optimization in gene therapy: promises, prospects and challenges

  • 1. Federal Research Center for Innovator and Emerging Biomedical and Pharmaceutical Technologies, Moscow, Russia

  • 2. The MCSC named after A. S. Loginov, Moscow, Russia

Abstract

Codon optimization has evolved to enhance protein expression efficiency by exploiting the genetic code’s redundancy, allowing for multiple codon options for a single amino acid. Initially observed in E. coli, optimal codon usage correlates with high gene expression, which has propelled applications expanding from basic research to biopharmaceuticals and vaccine development. The method is especially valuable for adjusting immune responses in gene therapies and has the potenial to create tissue-specific therapies. However, challenges persist, such as the risk of unintended effects on protein function and the complexity of evaluating optimization effectiveness. Despite these issues, codon optimization is crucial in advancing gene therapeutics. This study provides a comprehensive review of the current metrics for codon-optimization, and its practical usage in research and clinical applications, in the context of gene therapy.

1 Introduction

Codon optimization first appeared due to the search for an approach to increase the efficiency of expression of target proteins in bacterial cultures. The known property of degeneracy of the genetic code allows mRNA to encode the same proteins in different ways since 20 proteinogenic amino acids can be encoded by 61 codons (Welch et al., 2009). This property formed the basis of the codon optimization method, when, with the advent of genetic sequencing, it became evident that the usage of codons is non-random. Bias in codon usage occurs between different organisms, tissues, and sometimes even between parts of the same gene (; Pouyet et al., 2017). Thus, it became clear that the selection of the most common codons deemed suitable for an organism or cell line during genetic engineering research allows significantly changing approaches to conducting experiments.

Escherichia coli was the first organism with an analyzed codon usage system. Knowing the sequences of anticodons and the abundance of various tRNAs in the cell, the authors identified criteria for codon optimality (Ikemura, 1981). The first criterion was high codon recognition, the second was the highest abundance of tRNA. Highly expressed genes had a bias in frequency of use towards optimal codons, while genes with low expression were characterized by high randomness in the choice of codons ().

Currently, codon optimization has found application in a wide range of topics. In addition to fundamental research, control of the efficiency of protein expression through the selection of synonymous codons is also used for the development and production of biotherapies (), most of which are based on the expression of recombinant proteins. The method has become indispensable for molecular pharming on plants, where the problem of low expression efficiency is most pressing (Perlak et al., 1991; ; Thomas and Walmsley, 2014).

Differentiated cells determine the formation of tissues of various types. This complicated process can be modulated at the cellular and molecular level (Simon et al., 2018). At the molecular level, this diversity is reflected in particular in differences in protein expression - proteins that are abundant in one tissue may be absent in another (Thul and Lindskog, 2018). Differences in protein abundance are, in turn, caused by differences in RNA expression. One of the possible factors affecting such patterns is the different frequency of use of synonymous codons encoding the same amino acid during translation (Kames et al., 2020) (Figure 1). Indeed, either the rarity of codon usage (Plotkin et al., 2004) or the frequency of tRNA variants (; ) both vary between tissues. This can potentially be exploited for the construction of tissue-specific gene therapy. At the same time, to our knowledge, there is currently only one paper in peer-reviewed journals that has experimentally tested this hypothesis (Hernandez-Alias et al., 2023). This study is evidence that tissue-specific codon usage can potentially be used to design tissue-specific transgenes. At the same time, this metric is only one additional tool in the gene design toolbox whose implementation needs to be further explored and cannot be considered in isolation from several other indicators discussed below (Hernandez-Alias et al., 2023).

FIGURE 1

One of the most relevant and important areas of codon optimization application is the development of vaccines. The current way to create non-live vaccines is the use of attenuated viruses. Several research groups have experimented with attenuating poliovirus by changing codon bias in the gene encoding the poliovirus capsid protein, which involved replacing more frequent codons with less frequent ones (; Mueller et al., 2006). Moreover, increasing transgene expression in vaccines may improve the effectiveness of immunization and can be achieved through codon optimization (; ). In addition, a new class of vaccines—mRNA vaccines—has recently been introduced into clinical practice in the context of the COVID-19 pandemic (Oliver et al., 2020). Currently, the possibility of a similar approach for the prevention of infectious diseases such as rabies (Wan et al., 2023), influenza virus (Lee et al., 2023), Zika virus (), Lassa virus (Ronk et al., 2023) is the subject of active research and development. Remarkably, codon optimization of mRNA vaccines can significantly improve their stability and immunogenicity (Zhang et al., 2023). Despite the benefits of codon optimization, it is important to maintain a balance in the use of these techniques. Excessive interest in codon optimization can possibly lead to the accumulation of substances that are poorly excreted from the body, such as, for example, modified mRNA and the corresponding antigen (; Röltgen et al., 2022).

Currently, various approaches could be used for the development of gene therapeutics. Control of the immunogenicity of the administered drug is one of the most vital tasks not only in the preparation of vaccines but also for gene therapies. For the drug to work effectively, it is necessary to reduce the viral vector’s immunogenicity. It has been shown that by varying synonymous codons in the transgene and vector, it is possible to increase the effectiveness of therapy by lowering immunogenicity (; ), which provides optimism for simplifying vector selection and expanding the application of this type of therapy.

Regrettably, codon optimization techniques, while widely employed in the development of gene therapies, are far from perfect and are fraught with several challenges. One prominent issue lies in the incomplete synonymy of substitutions. This drawback carries the potential to disrupt natural post-transcriptional modification sites or, alternatively, give rise to novel sites, leading to critical alterations in the final protein’s structure, properties, and functions (; Irimia et al., 2012). Furthermore, overlooking the existence of alternative translation initiation sites (Malarkannan et al., 1999; Matsuda and Mauro, 2010)can lead to the unintended production of new proteins, adding another layer of complexity to the process. Beyond these intrinsic challenges, the selection of an appropriate numerical method for evaluating the effectiveness of codon optimization poses an additional obstacle. The abundance of metrics available complicates the task, requiring careful consideration to ensure a meaningful assessment. Despite the above difficulties, codon optimization approaches are actively used in clinical trials around the world and, furthermore, COVID-19 mRNA vaccines Pfizer/BioNTech and Moderna employ codon optimization.

Codon optimization can be carried out in many different ways today. It is often not clear which of these approaches is best suited to fulfill a particular task. The purpose of this review is to cover the current state of this problem and future directions for codon optimization approaches for gene therapies.

2 The quantitative assessment of codon usage and optimization

2.1 Measures of codon usage

The codon usage bias (CUB), also known as codon usage preferences (CUP), is influenced by a combination of factors that vary among species. Such factors include mutation frequency (Pizzo et al., 2015), selection for translation efficiency (Navon and Pilpel, 2011), and the presence of transfer RNA (tRNA) molecules that recognize specific codons (; Wei et al., 2019), ribosome binding efficiency (Shi et al., 2020), and translation speed and co-translational protein folding (Mitarai et al., 2008; Liu, 2020).

Based on the non-random usage of codons in the genomes of different species and the previously demonstrated positive correlation between codon bias and gene expression efficiency, Sharp and Li developed the relative synonymous codon usage (RSCU) scale (Sharp and Li, 1986). The RSCU value was calculated for a set of genes as the ratio of the observed codon frequency to the expected frequency, assuming equal usage of synonymous codons. This research has made a substantial contribution to the creation of various metrics, including but not limited to codon adaptation index (CAI) (Sharp and Li, 1987), average ratio of RSCU (ARSCU) (), and genetic tRNA adaptation index (gtAI) (). CAI continues to be a widely employed metric in both commercial and academic applications. CAI reflects the level of species-specific codon adaptation and is calculated as the geometric mean of RCSU values for each codon in the gene relative to the value of the most frequently used triplet encoding a single amino acid.

To date, numerous metrics for quantitative assessment of the level sequence optimization have been developed. Table 1 offers concise descriptions of commonly used metrics. To give the readers an idea of the frequency of metric usage, we added the citation rate of the original sources. However, it is important to emphasize that this approach does not reflect the level of usage of optimization tools based on the mentioned metrics.

TABLE 1

IndexMathematical formulaPrincipleInterpretationOriginal sourceCitation count
Frequency of optimal codons (Fop)Fop = Nopt/Ntotal, where Nopt denotes number of optimal codons in a gene, Ntotal indicates number of total codons in sequenceThe rationale behind Fop is that genes with a higher expression level tend to preferentially use certain codons, and this bias can be quantified using this ratioRepresenting the occurrence of the optimal codon within a gene sequence. Values close to 1 indicate a strong bias toward optimal codons, while values closer to 0 suggest a more even or random usage of synonymous codonsIkemura, 1981 (1982)1724 and 768
Relative synonymous codon usage (RSCU) where xij denotes the number of occurrences of the j-th codon for the i-th amino acid, ni represents the degeneracy for the i-th amino acidCalculated as the ratio of the observed codon frequency to the expected frequency assuming equal usage of synonymous codonsThe codon with RSCU of 1 value indicates average (random or equally) synonymous codon usage; RSCU value greater than 1 indicates positive codon usage bias (overrepresented); RSCU value less than 1 indicates negative codon usage bias (underrepresented)Sharp and Li (1986)653
Codon adaptation index (CAI), where L is the length of a gene measured in codons. Wk is the relative adaptiveness value for the k-th codon in the geneQuantifies the geometric mean of the RSCU for each codon with respect to the codon usage of a reference set of highly expressed genes. Genes that possess higher scores are anticipated to demonstrate enhanced efficiency in translation and elevated levels of protein expressionA higher CAI score implies that the gene’s codon usage is better aligned with the preferred codons identified in a highly expressed reference set. A CAI value of 1 suggests that the codon usage in the gene is perfectly adapted to the preferred codons observed in a highly expressed reference set; value of 0 indicates that the codon usage in the gene is not adaptedSharp and Li (1987)4091
Codon pair adaptation index (CPAI), where N − 1 is the number of codon pairs in gene g and w-i, j is the relative adaptiveness of the codon pair (i, j). , where fi, j is the frequency of codon pair (i, j) and f(max(i,j)) is the frequency of the codon pair most often used to code for the amino acid pair (aa(i),aa(j)) in a set of highly expressed genes GThe advantage of the index is the automatic weight selection algorithm . Combined use of CPAI and CAI have been shown to better predict gene expressionSimilar to CAI.43 and 402
Codon bias index (CBI), where Nopt represents the number of preferred optimal codons, Ntot is the total number of codons in the gene, and Nrand denotes the expected number of preferred codons if random codon assignments were made for each amino acidCodon bias refers to the unequal usage of synonymous codons encoding the same amino acid in a DNA sequence. The codon bias index is a measure that quantifies the extent of this biasValues range from 0 to 1: 0 indicates random selection of codons, and 1 indicates a significant bias towards preferred codons1955
Effective number of codons (ENc or Nc), where , where is the degeneracy of the amino acid (number of synonymous codons), is the fraction of each synonymous codon () out of the total codons () for that amino acid, The ENc is a measure designed to assess codon usage bias by evaluating how far the observed usage of synonymous codons deviates from an anticipated equal distributionThe ENc values range from 20 (maximum codon bias, only one codon used for each amino acid) to 61 (no bias, all synonymous codons used equally). If the ENc value is less than 35, it indicates a pronounced biasWright (1990)2306
Effective number of codon-pair (ENcp), where Fm is the average degeneracy of all amino acid pairs with m synonymous codon pairs, and km is the number of amino acid pairs with m synonymous representationsENcp is a metric designed to assess the bias in the usage of codon pairs, defined analogously to ENc, with the addition of a square rootSimilar to ENc100
Codon usage similarity index (COUSIN)COUSIN evaluates the codon CUB of a given query in comparison to a reference, and it standardizes the results by applying a Null Hypothesis that assumes random codon usage. In COUSIN18, all 18 families of synonymous codons contribute equally to the overall index, whereas in COUSIN59, each family contributes proportionally based on the frequency of the corresponding amino acid in the queryA COUSIN score of 1 signifies that the CUB in the query closely resemble those in the reference dataset. A COUSIN score of 0 indicates that the CUB in the query align with those in the Null Hypothesis. For scores exceeding 1, the CUB in the query share similarities with those in the reference but on a larger scale. Scores falling between 0 and 1 suggest that the CUPrefs in the query are akin to those in the reference but with a smaller magnitude. A score below 0 implies that the CUPrefs in the query are opposite to those in the reference51
, where N is the number of amino acids present in both the query and the reference and is the set of synonymous codons coding amino acid a. weight for each codon (Wc,a) in reference and test set is the frequency of the amino acid a in the query
Average ratio of RSCU (ARSCU), where aa is amino acid, a is RSCU of GC end codons and b is RSCU of AT end codons (any a and b with a value of zero is arbitrarily assigned a value of 0.1)Measures the ratio of RSCUs with GC-ending codons to the AT-ending codons for all amino acids in a geneGenes exhibiting ARSCU values surpassing 13 (a subjective threshold) are anticipated to showcase heightened expression levels. Genes with ARSCU values within the range of 9–13 are predicted to have either high or intermediate expression, while genes possessing ARSCU values below 9 are expected to demonstrate low or intermediate expression21
Relative codon bias strength (RCBS), where L is the length, in codons, of the gene, f (x,y,z) the observed frequency of codon xyz and f1(x), f2(y) and f3(z) the observed frequencies of bases x, y and z at, respectively, codon positions 1, 2 and 3It calculates the observed frequency of specific codons relative to the expected frequency, considering biases in base composition at three codon sites, providing a reliable measure of codon preferences while accounting for sequence-specific features like GC contentAn RCBS value near 0 implies an absence of codon usage bias, while a value exceeding 0.5 indicates a notable preference for specific codon usageRoymondal et al. (2009)109
Directional codon bias score (DCBS)DCBS is based on RCBS and allows the measurement of both positive and negative codon usage biasSimilar to RCBS.Sabi and Tuller (2014)116
Modified relative codon bias strength (MRCBS), where fxyz is the normalized codon frequency of a codon xyz and fn(m) is the normalized frequency of base m at codon position n in a gene. RCBSaa, max is the maximum value of RCBS of codon encoding the same amino acid aa in the same reference set, and N is the codon length of the query sequenceRibosomal protein is used as a reference set of genes and the dependence on gene length is overcomeThe score of the modified relative codon bias ranges from 0 to 11
Relative codon adaptation (RCA), where fxyz is the observed relative frequency of codon xyz in any reference gene set, fi(m) is the observed relative frequency of base m at codon position i in the same reference set and L is the length of the query sequenceThe RCA index calculates the anticipated frequency of a codon within a provided reference set by considering the positional base frequencies. Then, it assesses codon adaptation by comparing the observed codon frequency with the anticipated frequencyThe score of the RCA ranges from 0 to 2107
Codon deviation coefficient (CDC)The calculation of the metric can be found in the original sourceThe index takes into account the nucleotide composition of the sequence, GC content, and purine content. CUB is estimated using the cosine distance between the expected and observed codon usage vectors. Assessing the statistical significance of results using bootstrap resamplingValues range between 0 and 1: value 0 - no bias, 1 - maximum biasZhang et al. (2012)62
Index of Translation Elongation (ITE), where is the frequency of codon , is the number of sense codons (excluding those in single-codon families)CAI-like index, but it involves determining a weight for each codon by considering its frequency within the NNR and NNY codon subfamilies in the reference set. Subfamilies are distinguished due to different translation by different tRNAs and susceptibility to different mutational errorsITE values range between 0 and 1. A greater score is assigned to genes containing codons that are more commonly found in highly expressed genes[25]136
Synonymous codon usage order (SCUO) where j is the codon i-th amino acid. SCUOi is the SCUO for i-th amino acid in each sequence and Hi and Hmaxi are the entropy and maximum for an i-th amino acid in a sequenceThe SCUO index assesses how much a sequence deviates from a uniform distribution, using Shannon entropy as a basis. It involves the normalized difference between the maximum entropy and the observed entropySCUO the value varies between 0 and 1, and higher values indicate a stronger codon usage biasWan et al. (2006)21

Metrics for codon optimization with formal definition and description. The number of citations was retrieved from the Scopus database.

TABLE 2

Synonymous codon variants1st2nd3rd4th
Amino acid
AGCTGCCGCAGCG
DGATGAC
GGGTGGCGGAGGG
YTATTAC

Example representation of the 4-letter amino acid sequence ADGY (alanine-aspartic acid-glycine-tyrosine) via synonymous codons. Nucleotide sequence of wild-type GCC-GAT-GGT-TAT. There are 4 codon variants for the first and third amino acids, and 2 variants for the second and fourth amino acids. Total 64 possible variants of nucleotide presentation of this sequence.

Numerous metrics can be easily calculated with a reference set of genes to obtain the codon usage frequency. For example, Fop is calculated as the ratio of optimal codons to the total number of codons, excluding stop codons and codons without alternatives for amino acids (methionine, tryptophan) (Ikemura, 1981; 1982). The index aids in gauging the prevalence of synonymous codon usage. Other metrics are grounded in the assumption that the usage of codons is non-random. The metrics quantify the difference in codon usage frequency from a uniform distribution within the coding sequence. When all codon variants for a specific amino acid are utilized with equal frequency, such difference is minimal. Conversely, the maximum is achieved when only one codon out of the possible ones is utilized. Examples of such indices include ENC, CDC, SCUO, and others.

2.2 Codon adaptation metrics for assessing mRNA properties

Codon optimization is a strategy aimed at increasing the efficiency of mRNA translation and overcoming protein expression limitations. The use of synonymous codons affects the stability of mRNA in human cells (Narula et al., 2019; Wu et al., 2019). The thermodynamic stability of mRNA within a cell significantly influences translation efficiency (; ). mRNA is inherently unstable and can undergo transient states and adopt multiple stable structures. One approach to selecting synonymous amino acids for the purpose of thermodynamic stabilization is aimed at minimizing the free energy ΔG (MFE) released during RNA folding (Zuker and Stiegler, 1981; Zuker, 1994). Ringner and Krogh demonstrated in Saccharomyces cerevisiae that the folding free energy in the vicinity of the 5′-UTR correlates positively with transcription efficiency and mRNA half-life (Ringnér and Krogh, 2005).

An alternative approach suggests that the optimal structure will possess the maximum number of chemical bonds (Wayment-Steele et al., 2021). The AUP (Average Unpaired Probability) and SUP (Sum of Unpaired Probabilities) metrics, employed to assess RNA stability against hydrolytic degradation, operate under the premise that structures formed by paired bases exhibit lower susceptibility to hydrolysis.

Cluster analysis discovered that different mRNAs preferentially use different types of codons. Some mRNAs predominantly use optimal codons, while others prefer non-optimal codons. Furthermore, they observed that mRNAs with a higher proportion of optimal codons tend to be more stable, while those with a lower proportion of optimal codons are more unstable. Based on conducted experimental research, a metric called the codon stability coefficient (CSC) has been proposed. It is calculated as the Pearson correlation coefficient between the frequency of each codon and mRNA half-lives (Presnyak et al., 2015).

In the standard genetic code, the first two positions of a codon play a decisive role in coding an amino acid, while the third position is variable for one amino acid. Collection of metrics developed GC1, GC2, and GC3 represents the frequency of G + C usage at the first, second, and third positions, respectively (Stenico et al., 1994). Another evaluation derived from RSCU is the Average RSCU Ratio (ARSCU) (). Its noteworthy feature involves considering the base at the third position of the codon. The optimization of protein expression often involves the frequent usage of GC content. The model of post-transcriptional mRNA regulation involving P-bodies, 5′-3′ exonuclease XRN1, RNA helicase DDX6, and enhancer of decapping PAT1B shows that GC-rich coding sequences (CDS) result in higher protein production compared to AU-rich ones, and are controlled by a mechanism involving degradation factors DDX6 and XRN1 (). On the contrary, reducing the GC content in the 5′-UTR leads to an increase in free energy and also enhances protein yield, presumably due to mRNA destabilization in the translation initiation region and greater accessibility of the ribosome binding site (). The GC3 content varies depending on the type of tissue but is not an exhaustive characteristic for tissue-specific gene separation (Plotkin et al., 2004). GC3 codons are also associated with a longer half-life of mRNA (Kudla et al., 2006; Hia et al., 2019).

2.3 Metrics for adaptation to tRNA pool

Codon usage bias is closely linked to translational selection, which is the process of selecting codons that match abundant tRNAs, the molecules responsible for carrying amino acids during protein synthesis. Highly expressed genes tend to use such preferred codons, resulting in enhanced translation rates and accuracy. showed that the expression levels of nuclear and mitochondrial tRNAs vary between human tissues, indicating tissue-specific translational selection. However, minor differences in mouse mitochondrial RNA have only been detected for cardiac tissue, while significant differences between the central nervous system and other tissues have been demonstrated at the level of tRNA isodecoders, i.e., transcripts with the same anticodon but encoded by numerous different genes (Pinkard et al., 2020). It is important to note that the strength of translational selection varies across different organisms based on their genome sizes and genomic tRNA content (Reis, 2004).

To account for the role of intracellular tRNA content in translation efficiency, the following indices have been developed: P2index () and tRNA adaptation index (tAI) ().

Initially, tAI was only applicable to S. cerevisiae, but its subsequent modifications, stAI (Sabi et al., 2017) and gtAI ()—overcome this limitation by incorporating species-specific weights through algorithmic approaches to find extrema. gtAI demonstrated greater efficiency by employing a genetic algorithm to identify the optimal set of weights. In its calculation, indices ENc and RSCU are also incorporated. gtAI ranges from 0 to 1, where a higher value implies better adaptation of the codon to the tRNA pool.

The P2 Index is a metric used for the quantitative assessment of the efficiency of interactions between codons and their corresponding anticodons during the translation process. Based on the frequency of specific types of codons, values exceeding 0.5 indicate the presence of translational selection influencing the coding sequence.

2.4 Algorithmic approaches and tools for codon optimization

Currently, various optimization algorithms are utilized, such as the genetic algorithm (), multi-objective artificial bee colony (), Ribotree Monte Carlo (Leppek et al., 2022), and dynamic programming (Pham et al., 2004; Taneda and Asai, 2020), to identify codon combinations with desired characteristics. In several studies, the use of recurrent neural networks for codon optimization in heterologous protein expression has been presented in Chinese hamster (Gricetulus griseus) ovary cells () and E. coli (Jain et al., 2023). The Bidirectional Long Short-Term Memory (LSTM) deep learning model has also been trained for E. coli ().

Other studies applied machine learning methods for mRNA stabilization, such as integrated deep learning-based mRNA optimization (iDRO) (Jain et al., 2023), which provides a two-step optimization for the open reading frame and the untranslated regions. S. Castillo-Hair and G. Seelig trained a model on the 5′UTR polysome profile dataset to predict ribosome loading and protein expression (). The predictive power of such models strongly depends on the quantity and quality of the training datasets. At the same time, the accumulation of experimentally verified data sets is often not as fast as the development of machine learning methods. For example, to date (February 2024) only 6,142, of which 1,416 are human, experimentally validated RNA structures have been deposited in the Protein Data Bank (). This indicates that the high-precision prediction of RNA 3D structures using machine learning methods may be accurate for training data, but not for new data (Sato and Hamada, 2023).

Several software tools that utilize statistical and algorithmic solutions are available for commercial and free use. Here, we present some current tools that can be used for various tasks, including those related to gene therapy: ATGme (), OPTIMIZER (Puigbo et al., 2007), CHARMING (Wright et al., 2022), %MinMax (Rodriguez et al., 2018), JCat (), Optipyzer (LeRoy and Roleck, 2023), IDT (Owczarzy et al., 2008), gtAI ().

3 Codon optimization for gene therapy vectors

Above, the elucidation of metrics and principles related to codon optimization has been expounded. At the same time, it should be noted that the resources required to test the functionality of in silico predicted RNA variants significantly exceed the cost of the prediction itself. For this reason, studies often mainly present unconfirmed hypotheses in in vitro or in vivo experiments. Nevertheless, we present below some examples where codon optimization has been successfully applied in vitro. Proceeding to in vitro studies, it should be noted that gene therapeutics consist of a delivery vector and a therapeutic gene. Currently many types of vectors are used as a transgene vehicle (e.g., lipoplexes (), polyplexes (), virus-like particles (Pitoiset et al., 2017)).

Some of these vectors are a cassette with the selected viral genes, others do not contain nucleic acids. In some cases, wild-type viral genes in the gene therapy vector are not optimized for efficient application (). At the same time, codon-optimized variants of these sequences increase the efficacy of gene therapy, although they may lead to unfavorable results such as undesirable conformational changes and subsequently alterations in protein activity and function. Examples of codon optimization of adenoviral (), retroviral and lentiviral vectors () are discussed below.

Since adeno-associated vectors have recently become the most widely used platform for gene transfer (Mendell et al., 2021) and adenoviruses have long been successfully used to deliver genes (), we will consider the application of optimizations on their example.

It has been shown that in adenoviruses, the genes responsible for highly abundant late structural proteins tend to use codons frequently used in humans (optimal codons), while early regulatory use less optimal codons (Villanueva et al., 2016). However, the adenoviral fiber protein specifically uses suboptimal codons for efficient viral replication. Surprisingly, analysis of transgenes expressed in oncolytic adenoviruses, that are used for the oncoselective expression of a wide range of therapeutic molecules (; Huang et al., 2019) shows that most transgenes also use suboptimal codons. This contradicts the recommendation to use optimal host codons in transgenes to maximize gene expression. The study investigates the impact of transgene codon usage on viral fitness and finds that transgenes with higher GC3 content (optimal codon usage) have higher gene expression and viral replication, while those with lower GC3 content have lower expression and replication (Núñez-Manchón et al., 2021). By tuning the codon usage of transgenes, it is possible to achieve better transgene expression without compromising viral replication, thus optimizing the therapeutic outcome.

In the development of gene therapies, the problem arises of achieving high titers and a high ratio of empty to full capsids in viral vectors. One of the solutions to this obstacle is codon optimization of viral genomes encoding capsid proteins and assembly proteins. Thus, not only transgenes but also the coding sequences of the viral vector itself are subjected to codon optimization. For AAV-based (adeno-associated virus) vectors a novel codon optimization method was presented (Localized Codon-Optimization or LCO) ().

This method aims to preserve functional elements of the capsid genes and improve capsid shuffling efficiency for AAV engineering. The LCO algorithm performs localized optimization of codons at each position independently, based on the usage frequency of codons observed in the input variants of AAV sequences. A codon usage frequency table is generated for each amino acid position, and this table is used to optimize individual sequences (Table 3). The LCO-modified capsid genes showed increased sequence identity between parental AAV capsids and novel AAV capsid variants.

TABLE 3

Probability of finding a codon at a given position
codonPosition 1Position 2Position 3Position 4
GCT0.4
GCC0.17
GCA0.24
GCG0.19
GAT0.68
GAC0.32
GGT0.26
GGC0.22
GGA0.36
GGG0.16
TAT0.39
TAC0.61

An example of how the LCO method works to optimize the four codons of the mRNA encoding ADGY (see Table 2). A probability is calculated for all possible codons for a particular amino acid at a particular position. The most probable codons are marked in bold. Accordingly: GCC-GAT-GGT-TAT (wild-type nucleotide sequence)—would be optimized to GCT-GAT-GGA-TAC (final LCO-optimised sequence).

Functionality tests demonstrated that the optimized capsids retained their function, and transduction efficiency was similar to unoptimized counterparts. The LCO method also improved the efficiency of capsid shuffling, resulting in a highly shuffled library with increased complexity and reduced size of donor sequence segments. The shuffled clones generated using LCO-encoded capsids demonstrated successful transduction, indicating the effectiveness of LCO in generating novel AAV variants.

Ironically, the extensive use of codon optimization occurred simultaneously with abundant research findings that revealed the impact of synonymous mutations on protein function. This has been shown on a variety of proteins (; Kirchner et al., 2017).

The mechanism being discussed involves the comparison between codon-optimized (CO) and wild-type (WT) variants of a protein named FIX (coagulation factor IX). The results highlight that the CO and WT FIX variants exhibit distinct conformations, suggesting that the codon optimization process has influenced the protein’s structure. Ribosome profiling analyses uncover altered ribosomal distribution patterns and local translational kinetics in the CO variant when compared to the WT variant. Notably, these differences are unique to the CO FIX variant, as control genes demonstrate comparable ribosome distribution profiles ().

Despite the observed differences in translational kinetics, the overall efficiency of protein synthesis between the CO and WT variants remained similar. This finding is consistent with previous studies conducted in vitro (outside of a living organism) and suggests that the rate of protein synthesis is comparable between the two variants. The researchers propose that differences in translational kinetics within these domains may contribute to the observed conformational differences between the CO and WT FIX variants.

Codon optimization can be approached not only by a global view of codon usage in general, but also by a local optimization for each individual position in a particular amino acid. Moreover, it is also important to check that the functions of the essential elements and the optimized protein of interest remain unchanged.

4 The effect of codon optimization on immunogenicity

The immune response to an administered foreign substance or molecule can be defined as immunogenicity. It should be noted that higher immunogenicity increases the efficacy of the drug in some cases, but decreases it in others (Figure 2). For example, the purpose of immunization is to generate an immune response against a pathogen. In this case, methods should be used to increase the immunogenicity of the drug. It should be noted that in the development of mRNA vaccines, an excessive overreaction of the immune system is undesirable due to possible damage to the human organism (Igyártó and Qin, 2024) and should be taken into account during codon optimization. On the other hand, if a transgene introduced into the organism is intended to lead to the production of the corresponding protein, any degree of immunogenicity will reduce the effectiveness of the therapy. The innate and adaptive immune response to gene therapy may vary depending on the source of immunogenicity. These may be factors related to the capsid of the virion or to the viral genome. In relation to the capsid, binding of TLR2 or TLR9 can potentially activate the innate immune response and initiate the MyD88 signaling cascade, which in turn stimulates the production of proinflammatory cytokines such as TNF-alpha or induces the synthesis of IFN-gamma (Yang et al., 2022). Depending on the composition of the viral vector, the innate immune response can lead to enhanced adaptive immune responses. For example, AAVs, which are often used as gene therapy vectors, circulate naturally between humans. As a result, most people develop antibodies against natural AAV serotypes due to previous exposure. These antibodies are even known to cross-react with engineered vectors (). As a result, these antibodies can lead to either complement activation or neutralization of the capsid. The adaptive immune response is characterized by the degradation of the capsid protein by the proteasome and peptide presentation on MHC class I molecules. CD8+ cytotoxic T-cell lymphocytes can bind to the MHC, which leads to cell death (Martino et al., 2013). Peptide presentation on MHC class II molecules after phagocytosis and proteolysis can be recognized by CD4+ T lymphocytes, which can then stimulate the proliferation of B cells and the production of capsid-specific antibodies (Li et al., 2013). Studies have shown that plasmacytoid dendritic cells (pDCs) and conventional dendritic cells (cDCs) co-operate to achieve cross-priming of CD8+ T cells (Rogers et al., 2017). pDCs recognize the AAV genome via TLR9, while cDCs present the antigen on MHC I. The binding of cytokine-produced IFN to its receptor on cDCs is necessary for this process, indicating a direct relationship between pDC-produced cytokines and the activation of cDCs. Cross-priming of CD8+ T cells against AAV capsids requires CD40CD40L co-stimulation, which is performed in addition to T1 IFN from CD4+ Th cells (Shirley et al., 2020b).

FIGURE 2

After viral uncoating, TLR9 receptors can recognize unmethylated CpG motifs in the released single-stranded DNA, which also leads to activation of the innate immune system and stimulates cytokine production. The humoral and cellular innate immune responses described above for AAV capsids also occur for the transgene protein. The adaptive immune response can depend on various factors such as the target tissue, vector design and dose. Depending on the specificity of the promoter, there is a potential risk of immunogenicity (Shirley et al., 2020a). For example, a ubiquitous promoter can increase the risk of an adaptive cellular immune response of target and non-target cells (Sun et al., 2005).

It should be noted that the appearance of a foreign protein in the human organism is associated with the development of autoimmune diseases due to the similarity of individual epitopes of foreign and self proteins (Rojas et al., 2018). For example, it was recently shown that the same antibodies cross-react with the Epstein-Barr virus protein and the human alpha-crystallin B protein (Thomas et al., 2023). This phenomenon of molecular mimicry could be associated with the development of multiple sclerosis. The possibility of molecular mimicry of proteins resulting from the translation of the nucleic acids used must therefore be taken into account in the development of gene therapeutics. As already mentioned, codon optimization of the RNA can influence the structure of the translated protein (). As a result, depending on the different variants of the synonymous substitutions, the presentation of different epitopes of the same protein is possible.

It is of interest to reduce these CpG motifs to circumvent the possible human immune response, which can be achieved by codon optimization. For example, various elements of an AAV vector such as the CMV enhancer and promoter, ITR regions, UTR regions and the therapeutic transgene itself may contain CpG motifs. The CpGs within the promoter sequence can be removed, but with unpredictable effects on the activity and specificity of the promoter. For example, the authors have shown that the removal of CpGs within the CMV promoter gene significantly reduces its activity (Yew and Cheng, 2004). Although CpGs can be removed from the expression cassette, as in the case of human coagulation factor IX (hFIX) (), this does not always increase efficiency—CpG elimination had only reduced antibody formation against the transgene and not against the capsid itself. There are several studies in which this strategy was used, but mostly with a modification of the transgene. They have shown that the elimination of CpG motifs may lead to a significant reduction in the CD8+ T cell response (Yew and Cheng, 2004; ; Herzog et al., 2019; Wright, 2020; ; Konkle et al., 2021).

Several codon optimization strategies, including the chemical modification of nucleosides (Karikó et al., 2005) and the incorporation of pseudouridine (Karikó et al., 2008; ; Thess et al., 2015), have been shown to improve translation and reduce the immune response to mRNAs. pDCs exposed to such modified RNA exhibit a significant reduction in cytokines and activation markers. Nucleoside modification at a single position in a chemically synthesized oligoribonucleotide (ORN) is sufficient to abrogate TLR activation. In addition, the incorporation of pseudouridine in particular has been shown to facilitate evasion of recognition by Toll-like receptors (Karikó et al., 2005), although the molecular differences contributing to this mechanism has not yet been elucidated. Although the implementation of pseudouridine increases the stability of the mRNA and its translational capacity, it is important to note the disadvantages of replacing uridine with pseudouridine (Xia, 2021; Mueller, 2023). A recent study has shown that the presence of pseudouridine in IVT mRNA increases ribosomal + 1 frameshifting during mRNA translation. In addition, new peptides were generated that triggered an immune response (Mulroney et al., 2024). The presence of pseudouridine in the stop codon region suppresses translation termination and allows non-canonical base pairing, which is particularly detrimental for in vitro transcribed mRNAs (Loomis et al., 2016). The negative effects of pseudouridine synthases have been associated with various cancers (Xue et al., 2022) and autoimmune diseases (). This strongly suggests that the influence of codon optimization and pseudouridine incorporation on mRNA expression needs to be further investigated. A limitation of the present review is that it does not focus on a detailed description of the specific effects of codon optimization on the mRNA vaccines against COVID-19 per se that have been introduced into clinical practice (reviewed in Xia, 2021), but aims to discuss the advantages and disadvantages of the different options for the use of codon optimization in gene therapy in general.

To summarize, a common strategy to avoid immunogenicity is to eliminate redundant CpG motifs, implement chemical modifications of ORNs and replace uridine with pseudouridine. However, it should be noted that the implementation of codon optimization to eliminate CpG motifs and pseudouridine modification must be performed strategically to avoid the negative consequences of both approaches. Given the various unresolved factors leading to potential immunogenicity as a consequence of gene therapy, developing metrics for prediction is a complicated task. Nevertheless, a recent report (Wright, 2020) proposed a metric for prediction focusing exclusively on CpG motifs and their potential immunogenicity. Three formulas were developed that take into account the amount of unmethylated CpG motifs in the vector sequence. Known immunostimulatory sequences commonly used in DNA vaccines were also considered in the development of the formulae (). Although these formulae still need to be improved for full validation and accurate prediction, they reflect the beginning of a deeper understanding of how codon optimization can contribute to the reduction of immunogenicity.

5 Experimental testing of codon optimized sequences

There are numerous strategies for optimizing codons in nucleic acids. The methods mentioned above enable the creation of numerous optimized sequence variants. However, experimental verification of properties such as mRNA stability and protein expression levels is necessary before further experimentation can be conducted. Depending on the goals and available resources, it may be possible to select the best candidates based on chosen criteria from the range of design variants. These candidates can then be examined using routine laboratory methods. Alternatively, a pool of hundreds of sequences can be studied, in which case high-throughput protocols must be developed (Figure 3).

FIGURE 3

When studying a small number of variants, it is possible to determine the expression level separately for each construct after transfecting the cells. To quantify transgene expression in this case, the most common method is to use target-specific primers with cDNA obtained from RNA by reverse transcription as a matrix and perform qPCR (Leppek et al., 2022). Expression can be quantified at both the transcriptional and translational levels. The latter involves the analysis of synthesized proteins and can be performed using antibodies specific to the target protein. For instance, Zhang (Zhang et al., 2023) described the properties of the optimized structure of the SARS-CoV-2 virus S protein using flow cytometry. A possible alternative method for determining protein concentrations is to use SDS-PAGE gels for Western blot analysis, along with specific antibodies (Raab et al., 2010; ).

Although codon optimization of the target sequence can provide certain benefits, it may also result in reduced mRNA stability in solution, which impairs its functionality. Therefore, it is necessary to experimentally confirm the stability of the structure of optimized nucleic acids. The stability of mRNA molecules is inversely proportional to their degradation rate in solution. To determine the degradation rate, mRNAs are incubated in PBS buffer containing Mg2+ ions. Samples are collected at various time intervals of 1–2 h, and the number of fragments produced is estimated using capillary electrophoresis (Zhang et al., 2023) or polyacrylamide gel electrophoresis with urea. Therefore, the RNA is less stable if it degrades more quickly after being incubated in solution.

However, the laboratory approaches described above are time-consuming when testing multiple variants of codon-optimized sequences. In light of this, there is a great need to create high-throughput methods for studying many sequences simultaneously.

Most methods that allow mass screening of sequences follow a general principle: a unique barcode, a sequence of several nucleotides, is inserted into each variant. All the sequences to be tested can then be pooled and processed in a multiplex format. The presence of the barcode makes it possible to identify a variant using high-throughput sequencing platforms after all the necessary protocol steps have been completed.

Massively parallel variant analysis requires the synthesis of a library of DNA templates. The next steps in the study can be performed in two ways. The first involves transcription and modification (3′ polyA tail and 5′ m7G capping) in vitro, followed by transfection of the resulting mRNA pool into cells for further experiments. The “PERSIST-seq” method was developed based on this approach. It enables the simultaneous evaluation of stability and translation efficiency of over 200 mRNA molecules, making it a convenient tool for messenger RNA development (Leppek et al., 2022). In this case, the design of the DNA must take into account the presence of a promoter in the initial sequence. The second approach involves creating a vector library with cassettes that contain the sequence under study and regions of homology. The cells are then transfected with the library, and the sequences are integrated into the genome using CRISPR/Cas. This process enables the direct synthesis of mRNA within the cells. A study of the motifs that cause ribosome slowdown in a yeast model system describes a similar approach (). The next steps for experimental validation in both cases involve isolating RNA from cell culture, analyzing it through high-throughput sequencing, and quantifying the results. To identify inserts in the pool of isolated nucleic acids, unique barcodes are introduced into the library construct, which is a common aspect of the described strategies.

The presence of unique barcodes in the original DNA matrices allows quantitative assessment of the expression level for each individual variant using high-throughput RNA sequencing.

Translation of sequence variants has been demonstrated to be a crucial determinant in mammalian gene expression (). However, genomic expression profiling alone cannot reveal the precise regulation provided by post-transcriptional mechanisms, such as 5′ capping, splicing, polyadenylation, nuclear export, translation, and decay. To overcome this limitation, a polysome profiling method can be used to isolate ribosome-free and polysome-associated RNAs for further independent analysis (Pereira et al., 2018) This method involves separating mRNA in a sucrose gradient into two fractions: polysome-bound and polysome-free. The mRNA is then isolated from both fractions and sequenced using one of the available high-throughput platforms.

When studying multiple variants, stability assessment is also important. To identify full-length molecules that have not degraded, it is necessary to amplify the cDNA that was reverse transcribed from the RNA and then sequence it to quantify the amount of intact mRNA at each time point. This method can evaluate mRNA stability in both solution and cells. The solution replicates the conditions in which the molecules may be present during therapy, typically high pH and positively charged media. It is important to note that the outcomes obtained after incubation in solution differ significantly from those obtained after isolation from cells. This is likely due to cellular mechanisms of RNA degradation (Leppek et al., 2022).

Therefore, there are approaches that allow for the evaluation of the efficiency and stability of nucleic acid sequences obtained during codon optimization. The choice of a particular method depends on the number of variants to be analyzed. If there are only a few variants, it is possible to describe the properties of each variant separately, providing a fairly accurate understanding of its characteristics. When dealing with hundreds or thousands of variants, high-throughput methods are necessary. This allows for a pool of samples to be tested instead of individual samples, greatly increasing the productivity of experimental work. It is important to note that massively parallel sequencing methods provide high accuracy analysis, while polysome profiling can offer additional insights into the impact of codon optimization on the final product’s quality.

6 Future directions

Currently, there are some gene therapies that use different codon optimization metrics and are approved by the FDA (). To analyse other therapies that are in clinical trials and where codon optimization has been used, we conducted a thorough examination of the data available on ClinicalTrials.gov () until December 2023. A systematic search strategy was devised using the keyword “gene therapy” in the Condition/disease field. In addition to the specified search criteria, it is important to note that the term “vector” was included in the “Other terms” considered in the search. The algorithm did not include any specified values for the “Intervention/treatment” and “Location” categories in the search process. After searching, the algorithm automatically incorporated synonyms for the given query: gene: “Genes,” gene therapy: “Gene transfer”; “Gene Transfer Procedure,”, therapy: “treatment”; “Therapeutic”; “therapeutics”.

Furthermore, a comprehensive search was conducted using the specific only Condition/disease of “codon optimized” and excluded any specified values for the “Other terms,” “Intervention/treatment” and “Location” categories in the search process. However, it is crucial to mention that studies explicitly referring to monoclonal antibodies and enzymes as drugs in the Study URL and Brief Summary columns were manually excluded from the sample. This careful exclusion strategy ensured that the selected studies focused specifically on codon optimization. The search was conducted over a period of 20 years to capture an extensive range of relevant clinical studies.

Of the 395 clinical studies analyzed, only 12 contained information on codon optimization (Figure 4).

FIGURE 4

Prior to experimental testing of codon-optimized sequences using any of the aforementioned methods, it is essential to synthesize these sequences, often in large quantities. The most widely used method currently is phosphoramidite synthesis, which involves the interaction of nucleotide phosphoramidite monomers protected by acid-labile groups with an activating agent, binding to the growing oligonucleotide (Sinyakov et al., 2021). There are two main types of implementation for this approach, depending on the equipment used: synthesis on columns or on microarrays. The former option allows for the synthesis of oligonucleotides at a relatively low cost and with an error rate of 1 per 600 base pairs or less on average. However, it does not provide sufficient throughput for mass synthesis of oligonucleotides (Ma et al., 2012). Furthermore, if the sequence of interest exceeds 200 base pairs (some estimates suggest 300 (Palluk et al., 2018)), an additional assembly step via molecular cloning is required (). These factors significantly limit the speed of testing and represent the primary bottleneck in experimental design.

This problem can be solved by integrating higher-throughput oligonucleotide microarray synthesisers into laboratory practice (Song et al., 2021). Commercially available technologies are also based on phosphoramidite synthesis, albeit with slight modifications. Although microarray-based nucleotide synthesis is more error-prone due to heterogeneity and edge effects, it enables the synthesis of oligonucleotide pools and also reduces the cost per nucleotide by 2–4 orders of magnitude compared to column synthesis (Kosuri and Church, 2014). This suggests that advances in de novo DNA synthesis and experimental verification of codon-optimized sequences are likely to be associated with the microarray approach.

Since 2020, a trend towards an increase in the proportion of codon-optimized studies has been observed. In 2020, 1 in 34 (2.9%) clinical trials used codon optimization, compared to 4 in 42 (9.5%) in the first 11 months of 2023 (Figure 4). The main aim of codon optimization was to increase the level of transgene expression and the stability of the mRNA. In addition, a study using codon optimization to reduce immunogenicity was reported in 2021.

To effectively achieve the goals of codon optimization in research, it is important to follow established metrics. However, today there is no single generally accepted standard for codon optimization. Therefore, it is possible to use a large number of combinations of the methods described above to create optimal RNA variants. Some of these approaches significantly increase the efficacy of gene therapeutics. Therefore, several drug options have been registered in clinical trials, for example.

Codon optimization has played an important role in the development of RNA-based COVID-19 vaccines. Current research efforts are focused on further advancing the field of codon optimization for COVID-19 vaccines to address new strains of the coronavirus (Wu et al., 2023). Unfortunately, it was not possible to provide here the specific metrics used for codon optimization in the above-mentioned studies for commercial product development. This limitation results from the intellectual property of the original codon-optimized constructs. In this article, we have explored various metrics for assessing codon usage, based on both the composition of the coding sequence and the composition of a reference set of genes. One widely used metric is the Codon Adaptation Index (CAI). Although these measures provide useful information about adaptation to the host organism, they do not necessarily indicate an increase in translational efficiency due to selection pressure (Rahman et al., 2018; ). Furthermore, CAI is also interpreted as an indicator of the speed of translational elongation (Kudla et al., 2009). In turn, an increase in translation speed may not necessarily result in the production of a protein with similar properties in greater quantities.

Apparently, during translation, the most important regions for codon optimization are the areas around the start codon. This is supported by work demonstrating the contribution of the CDS position near the start codon (Höllerer and Jeschek, 2023; Nieuwkoop et al., 2023) and the 5′UTR sequence region (). The efficiency of translation is significantly dependent on the energy of mRNA folding, particularly in the vicinity of the start codon (). This is associated with the fact that unfolding more stable RNA secondary structures require greater energy before the initiation of translation (Figure 5). Additionally, the presence of hairpin, stem-loop, and pseudoknot structures in mRNA can hinder ribosome translocation and tRNA binding, thus impeding translation elongation (Kozak, 2005; ).

FIGURE 5

Thus, advancements in gene therapy could be directed towards a more comprehensive exploration of the impact of codon optimization on the characteristics and secondary structure of mRNA.Also, it is possible to apply optimization metrics locally to the start region, but there are limitations since many of them are based on codon usage frequency without taking into account the features of untranslated regions.

In addition, consideration of local codon optimization is a critical aspect that must be taken into account during codon optimization for a particular protein of interest. Furthermore, essential protein functions may change due to the possible influence of codon optimization on the conformation of the resulting protein, which should also be taken into account.

Statements

Author contributions

AP: Writing–original draft, Writing–review and editing. AK: Writing–original draft, Writing–review and editing. AM: Writing–original draft, Writing–review and editing. DN: Writing–original draft, Writing–review and editing. AS: Writing–original draft, Writing–review and editing. IA: Writing–original draft, Writing–review and editing. SF: Writing–original draft, Writing–review and editing. OM: Writing–original draft, Writing–review and editing. AD: Writing–original draft, Writing–review and editing. PV: Writing–original draft.

Funding

The author(s) declare that financial support was received for the research, authorship, and/or publication of this article. This study was supported by the Russian Science Foundation (Grant No. 23-64-00002).

Conflict of interest

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

  • 1

    AlexakiA.HettiarachchiG. K.AtheyJ. C.KatneniU. K.SimhadriV.Hamasaki-KatagiriN.et al (2019a). Effects of codon optimization on coagulation factor IX translation and structure: implications for protein and gene therapies. Sci. Rep.9, 15449. 10.1038/s41598-019-51984-2

  • 2

    AlexakiA.KamesJ.HolcombD. D.AtheyJ.Santana-QuinteroL. V.LamP. V. N.et al (2019b). Codon and codon-pair usage tables (CoCoPUTs): facilitating genetic variation analyses and recombinant gene design. J. Mol. Biol.431, 24342441. 10.1016/j.jmb.2019.04.021

  • 3

    AndersonB. R.MuramatsuH.NallagatlaS. R.BevilacquaP. C.SansingL. H.WeissmanD.et al (2010). Incorporation of pseudouridine into mRNA enhances translation by diminishing PKR activation. Nucleic Acids Res.38, 58845892. 10.1093/nar/gkq347

  • 4

    AnwarA. M.KhodaryS. M.AhmedE. A.OsamaA.EzzeldinS.TaniosA.et al (2023). gtAI: an improved species-specific tRNA adaptation index using the genetic algorithm. Front. Mol. Biosci.10, 1218518. 10.3389/fmolb.2023.1218518

  • 5

    AthanasopoulosT.FosterH.FosterK.DicksonG. (2011). Codon optimization of the microdystrophin gene for Duchene muscular dystrophy gene therapy. Gene Ther.709, 2137. 10.1007/978-1-61737-982-6_2

  • 6

    AtheyJ.AlexakiA.OsipovaE.RostovtsevA.Santana-QuinteroL. V.KatneniU.et al (2017). A new and updated resource for codon usage tables. BMC Bioinforma.18, 391. 10.1186/s12859-017-1793-7

  • 7

    AyyarB. V.AroraS.RaviS. S. (2017). Optimizing antibody expression: the nuts and bolts. Methods116, 5162. 10.1016/j.ymeth.2017.01.009

  • 8

    BainbridgeJ. W. B.SmithA. J.BarkerS. S.RobbieS.HendersonR.BalagganK.et al (2008). Effect of gene therapy on visual function in leber’s congenital amaurosis. N. Engl. J. Med.358, 22312239. 10.1056/NEJMoa0802268

  • 9

    BansalS.PerincheriS.FlemingT.PoulsonC.TiffanyB.BremnerR. M.et al (2021). Cutting edge: circulating exosomes with covid spike protein are induced by BNT162b2 (Pfizer–BioNTech) vaccination prior to development of antibodies: a novel mechanism for immune activation by mRNA vaccines. J. Immunol.207, 24052410. 10.4049/jimmunol.2100637

  • 10

    BaoC.LoerchS.LingC.KorostelevA. A.GrigorieffN.ErmolenkoD. N. (2020). mRNA stem-loops can pause the ribosome by hindering A-site tRNA binding. Elife9, e55799. 10.7554/eLife.55799

  • 11

    BellP.WangL.ChenS.-J.YuH.ZhuY.NayalM.et al (2016). Effects of self-complementarity, codon optimization, transgene, and dose on liver transduction with AAV8. Hum. Gene Ther. Methods27, 228237. 10.1089/hgtb.2016.039

  • 12

    BennetzenJ. L.HallB. D. (1982). Codon selection in yeast. J. Biol. Chem.257, 30263031. 10.1016/S0021-9258(19)81068-2

  • 13

    BermanH. M. (2000). The protein Data Bank. Nucleic Acids Res.28, 235242. 10.1093/nar/28.1.235

  • 14

    BertoliniT. B.ShirleyJ. L.ZolotukhinI.LiX.KaishoT.XiaoW.et al (2021). Effect of CpG depletion of vector genome on CD8+ T cell responses in AAV gene therapy. Front. Immunol.12, 672449. 10.3389/fimmu.2021.672449

  • 15

    BłażejP.WnętrzakM.MackiewiczD.MackiewiczP. (2018). Optimization of the standard genetic code according to three codon positions using an evolutionary algorithm. PLoS One13, e0201715. 10.1371/journal.pone.0201715

  • 16

    BodeC.ZhaoG.SteinhagenF.KinjoT.KlinmanD. M. (2011). CpG DNA as a vaccine adjuvant. Expert Rev. Vaccines10, 499511. 10.1586/erv.10.174

  • 17

    BollmanB.NunnaN.BahlK.HsiaoC. J.BennettH.ButlerS.et al (2023). An optimized messenger RNA vaccine candidate protects non-human primates from Zika virus infection. npj Vaccines8, 58. 10.1038/s41541-023-00656-4

  • 18

    BourretJ.AlizonS.BravoI. G. (2019). COUSIN (COdon usage similarity INdex): a normalized measure of codon usage preferences. Genome Biol. Evol.11, 35233528. 10.1093/gbe/evz262

  • 19

    BoutinS.MonteilhetV.VeronP.LeborgneC.BenvenisteO.MontusM. F.et al (2010). Prevalence of serum IgG and neutralizing factors against adeno-associated virus (AAV) types 1, 2, 5, 6, 8, and 9 in the healthy population: implications for gene therapy using AAV vectors. Hum. Gene Ther.21, 704712. 10.1089/hum.2009.182

  • 20

    BreckpotK.EscorsD.ArceF.LopesL.KarwaczK.Van LintS.et al (2010). HIV-1 lentiviral vector immunogenicity is mediated by toll-like receptor 3 (TLR3) and TLR7. J. Virol.84, 56275636. 10.1128/JVI.00014-10

  • 21

    BuchanJ. R. (2006). tRNA properties help shape codon pair preferences in open reading frames. Nucleic Acids Res.34, 10151027. 10.1093/nar/gkj488

  • 22

    BuhrF.JhaS.ThommenM.MittelstaetJ.KutzF.SchwalbeH.et al (2016). Synonymous codons direct cotranslational folding toward different protein conformations. Mol. Cell61, 341351. 10.1016/j.molcel.2016.01.008

  • 23

    BulchaJ. T.WangY.MaH.TaiP. W. L.GaoG. (2021). Viral vector platforms within the gene therapy landscape. Signal Transduct. Target. Ther.6, 53. 10.1038/s41392-021-00487-6

  • 24

    BurkeP. C.ParkH.SubramaniamA. R. (2022). A nascent peptide code for translational control of mRNA stability in human cells. Nat. Commun.13, 6829. 10.1038/s41467-022-34664-0

  • 25

    BurnsC. C.ShawJ.CampagnoliR.JorbaJ.VincentA.QuayJ.et al (2006). Modulation of poliovirus replicative fitness in HeLa cells by deoptimization of synonymous codon usage in the capsid region. J. Virol.80, 32593272. 10.1128/JVI.80.7.3259-3272.2006

  • 26

    Cabanes-CreusM.GinnS. L.AmayaA. K.LiaoS. H. Y.WesthausA.HallwirthC. V.et al (2019). Codon-optimization of wild-type adeno-associated virus capsid sequences enhances DNA family shuffling while conserving functionality. Mol. Ther. - Methods Clin. Dev.12, 7184. 10.1016/j.omtm.2018.10.016

  • 27

    CapellA.FellererK.HaassC. (2014). Progranulin transcripts with Short and long 5′ untranslated regions (UTRs) are differentially expressed via posttranscriptional and translational repression. J. Biol. Chem.289, 2587925889. 10.1074/jbc.M114.560128

  • 28

    CarboneA.ZinovyevA.KépèsF. (2003). Codon adaptation index as a measure of dominating codon bias. Bioinformatics19, 20052015. 10.1093/bioinformatics/btg272

  • 29

    CasiniA.StorchM.BaldwinG. S.EllisT. (2015). Bricks and blueprints: methods and standards for DNA assembly. Nat. Rev. Mol. Cell Biol.16, 568576. 10.1038/nrm4014

  • 30

    Castillo-HairS. M.SeeligG. (2022). Machine learning for designing next-generation mRNA therapeutics. Acc. Chem. Res.55, 2434. 10.1021/acs.accounts.1c00621

  • 31

    Chamani MohassesF.SoloukiM.GhareyazieB.FahmidehL.MohsenpourM. (2020). Correlation between gene expression levels under drought stress and synonymous codon usage in rice plant by in-silico study. PLoS One15, e0237334. 10.1371/journal.pone.0237334

  • 32

    ChenK. Y.ParkH.SubramaniamA. R. (2023). Massively parallel identification of sequence motifs triggering ribosome-associated mRNA quality control. bioRxiv, 2023.09.27.2023.09.27.559793. 10.1101/2023.09.27.559793

  • 33

    ChenM.-W.ChengT.-J. R.HuangY.JanJ.-T.MaS.-H.YuA. L.et al (2008). A consensus–hemagglutinin-based DNA vaccine that protects mice against divergent H5N1 influenza viruses. Proc. Natl. Acad. Sci.105, 1353813543. 10.1073/pnas.0806901105

  • 34

    ChenW.LiH.LiuZ.YuanW. (2016). Lipopolyplex for therapeutic gene delivery and its application for the treatment of Parkinson’s disease. Front. Aging Neurosci.8, 68. 10.3389/fnagi.2016.00068

  • 35

    ClinicalTrials.govClinicalTrials.gov (2024).

  • 36

    CoughlanL. (2020). Factors which contribute to the immunogenicity of non-replicating adenoviral vectored vaccines. Front. Immunol.11, 909. 10.3389/fimmu.2020.00909

  • 37

    CourelM.ClémentY.BossevainC.ForetekD.Vidal CruchezO.YiZ.et al (2019). GC content shapes mRNA storage and decay in human cells. Elife8, e49708. 10.7554/eLife.49708

  • 38

    DanielE.OnwukweG. U.WierengaR. K.QuagginS. E.VainioS. J.KrauseM. (2015). ATGme: open-source web application for rare codon identification and custom DNA sequence optimization. BMC Bioinforma.16, 303. 10.1186/s12859-015-0743-5

  • 39

    DasS. (2017). Analysis of gene expression using modified relative codon bias strength in nanoarchaeum equitans. Biosci. Biotechnol. Res. Asia14, 793799. 10.13005/bbra/2510

  • 40

    DesaiP. N.ShrivastavaN.PadhH. (2010). Production of heterologous proteins in plants: strategies for optimal expression. Biotechnol. Adv.28, 427435. 10.1016/j.biotechadv.2010.01.005

  • 41

    de SostoaJ.FajardoC. A.MorenoR.RamosM. D.Farrera-SalM.AlemanyR. (2019). Targeting the tumor stroma with an oncolytic adenovirus secreting a fibroblast activation protein-targeted bispecific T-cell engager. J. Immunother. Cancer7, 19. 10.1186/s40425-019-0505-4

  • 42

    DewiK. S.FuadA. M. (2020). Improving the expression of human granulocyte colony stimulating factor in Escherichia coli by reducing the GC-content and increasing mRNA folding free energy at 5’-terminal end. Adv. Pharm. Bull.10, 610616. 10.34172/apb.2020.073

  • 43

    DiezM.Medina-MuñozS. G.CastellanoL. A.da Silva PescadorG.WuQ.BazziniA. A. (2022). iCodon customizes gene expression based on the codon composition. Sci. Rep.12, 12126. 10.1038/s41598-022-15526-7

  • 44

    DittmarK. A.GoodenbourJ. M.PanT. (2006). Tissue-specific differences in human transfer RNA expression. PLoS Genet.2, e221. 10.1371/journal.pgen.0020221

  • 45

    dos ReisM. (2003). Unexpected correlations between gene expression and codon usage bias from microarray data for the whole Escherichia coli K-12 genome. Nucleic Acids Res.31, 69766985. 10.1093/nar/gkg897

  • 46

    FathS.BauerA. P.LissM.SpriestersbachA.MaertensB.HahnP.et al (2011). Multiparameter RNA and codon optimization: a standardized tool to assess and enhance autologous mammalian gene expression. PLoS One6, e17596. 10.1371/journal.pone.0017596

  • 47

    FaustS. M.BellP.CutlerB. J.AshleyS. N.ZhuY.RabinowitzJ. E.et al (2013). CpG-depleted adeno-associated virus vectors evade immune detection. J. Clin. Invest.123, 29943001. 10.1172/JCI68205

  • 48

  • 49

    FengH.SegalésJ.WangF.JinQ.WangA.ZhangG.et al (2022). Comprehensive analysis of codon usage patterns in Chinese porcine circoviruses based on their major protein-coding sequences. Viruses14, 81. 10.3390/v14010081

  • 50

    FestenE. A. M.GoyetteP.GreenT.BoucherG.BeauchampC.TrynkaG.et al (2011). A meta-analysis of genome-wide association scans identifies IL18RAP, PTPN2, TAGAP, and PUS10 as shared risk loci for crohn’s disease and celiac disease. PLoS Genet.7, e1001283. 10.1371/journal.pgen.1001283

  • 51

    FoxJ. M.ErillI. (2010). Relative codon adaptation: a generic codon bias index for prediction of gene expression. DNA Res.17, 185196. 10.1093/dnares/dsq012

  • 52

    FribergM.von RohrP.GonnetG. (2004). Limitations of codon adaptation index and other coding DNA‐based features for prediction of protein expression in Saccharomyces cerevisiae. Yeast21, 10831093. 10.1002/yea.1150

  • 53

    FuH.LiangY.ZhongX.PanZ.HuangL.ZhangH.et al (2020). Codon optimization with deep learning to enhance protein expression. Sci. Rep.10, 1761717619. 10.1038/s41598-020-74091-z

  • 54

    GaoW.Gallardo-DoddC. J.KutterC. (2022). Cell type–specific analysis by single-cell profiling identifies a stable mammalian tRNA–mRNA interface and increased translation efficiency in neurons. Genome Res.32, 97110. 10.1101/gr.275944.121

  • 55

    Godfried SieC.HeslerS.MaasS.KuchkaM. (2012). IGFBP7’s susceptibility to proteolysis is altered by A‐to‐I RNA editing of its transcript. FEBS Lett.586, 23132317. 10.1016/j.febslet.2012.06.037

  • 56

    Gonzalez-SanchezB.Vega-RodríguezM. A.Santander-JiménezS.Granado-CriadoJ. M. (2019). Multi-Objective Artificial Bee Colony for designing multiple genes encoding the same protein. Appl. Soft Comput.74, 9098. 10.1016/j.asoc.2018.10.023

  • 57

    GouletD. R.YanY.AgrawalP.WaightA. B.MakA. N.ZhuY. (2023). Codon optimization using a recurrent neural network. J. Comput. Biol.30, 7081. 10.1089/cmb.2021.0458

  • 58

    GouyM.GautierC. (1982). Codon usage in bacteria: correlation with gene expressivity. Nucleic Acids Res.10, 70557074. 10.1093/nar/10.22.7055

  • 59

    GroteA.HillerK.ScheerM.MunchR.NortemannB.HempelD. C.et al (2005). JCat: a novel tool to adapt codon usage of a target gene to its potential expression host. Nucleic Acids Res.33, W526W531. 10.1093/nar/gki376

  • 60

    GuW.ZhouT.WilkeC. O. (2010). A universal trend of reduced mRNA stability near the translation-initiation site in prokaryotes and eukaryotes. PLoS Comput. Biol.6, e1000664. 10.1371/journal.pcbi.1000664

  • 61

    HansonG.CollerJ. (2018). Codon optimality, bias and usage in translation and mRNA decay. Nat. Rev. Mol. Cell Biol.19, 2030. 10.1038/nrm.2017.91

  • 62

    HayatS. M. G.FarahaniN.SafdarianE.RoointanA.SahebkarA. (2019). Gene delivery using lipoplexes and polyplexes: principles, limitations and solutions. Crit. Rev. Eukaryot. Gene Expr.29, 2936. 10.1615/CritRevEukaryotGeneExpr.2018025132

  • 63

    Hernandez-AliasX.BenistyH.RaduskyL. G.SerranoL.SchaeferM. H. (2023). Using protein-per-mRNA differences among human tissues in codon optimization. Genome Biol.24, 3420. 10.1186/s13059-023-02868-2

  • 64

    HerzogR. W.CooperM.PerrinG. Q.BiswasM.MartinoA. T.MorelL.et al (2019). Regulatory T cells and TLR9 activation shape antibody formation to a secreted transgene product in AAV muscle gene transfer. Cell. Immunol.342, 103682. 10.1016/j.cellimm.2017.07.012

  • 65

    HiaF.YangS. F.ShichinoY.YoshinagaM.MurakawaY.VandenbonA.et al (2019). Codon bias confers stability to human mRNA s. EMBO Rep.20, e48220. 10.15252/embr.201948220

  • 66

    HöllererS.JeschekM. (2023). Ultradeep characterisation of translational sequence determinants refutes rare-codon hypothesis and unveils quadruplet base pairing of initiator tRNA and transcript. Nucleic Acids Res.51, 23772396. 10.1093/nar/gkad040

  • 67

    HuangH.LiuY.LiaoW.CaoY.LiuQ.GuoY.et al (2019). Oncolytic adenovirus programmed by synthetic gene circuit for cancer immunotherapy. Nat. Commun.10, 4801. 10.1038/s41467-019-12794-2

  • 68

    IgyártóB. Z.QinZ. (2024). The mRNA-LNP vaccines – the good, the bad and the ugly?Front. Immunol.15, 1336906. 10.3389/fimmu.2024.1336906

  • 69

    IkemuraT. (1981). Correlation between the abundance of Escherichia coli transfer RNAs and the occurrence of the respective codons in its protein genes: a proposal for a synonymous codon choice that is optimal for the E. coli translational system. J. Mol. Biol.151, 389409. 10.1016/0022-2836(81)90003-6

  • 70

    IkemuraT. (1982). Correlation between the abundance of yeast transfer RNAs and the occurrence of the respective codons in protein genes. J. Mol. Biol.158, 573597. 10.1016/0022-2836(82)90250-9

  • 71

    IrimiaM.DenucA.FerranJ. L.PernauteB.PuellesL.RoyS. W.et al (2012). Evolutionarily conserved A-to-I editing increases protein stability of the alternative splicing factor Nova1. RNA Biol.9, 1221. 10.4161/rna.9.1.18387

  • 72

    JainR.JainA.MauroE.LeShaneK.DensmoreD. (2023). ICOR: improving codon optimization with recurrent neural networks. BMC Bioinforma.24, 132. 10.1186/s12859-023-05246-8

  • 73

    KamesJ.AlexakiA.HolcombD. D.Santana-QuinteroL. V.AtheyJ. C.Hamasaki-KatagiriN.et al (2020). TissueCoCoPUTs: novel human tissue-specific codon and codon-pair usage tables based on differential tissue gene expression. J. Mol. Biol.432, 33693378. 10.1016/j.jmb.2020.01.011

  • 74

    KarikóK.BucksteinM.NiH.WeissmanD. (2005). Suppression of RNA recognition by toll-like receptors: the impact of nucleoside modification and the evolutionary origin of RNA. Immunity23, 165175. 10.1016/j.immuni.2005.06.008

  • 75

    KarikóK.MuramatsuH.WelshF. A.LudwigJ.KatoH.AkiraS.et al (2008). Incorporation of pseudouridine into mRNA yields superior nonimmunogenic vector with increased translational capacity and biological stability. Mol. Ther.16, 18331840. 10.1038/mt.2008.200

  • 76

    KirchnerS.CaiZ.RauscherR.KastelicN.AndingM.CzechA.et al (2017). Alteration of protein function by a silent polymorphism linked to tRNA abundance. PLOS Biol.15, e2000779. 10.1371/journal.pbio.2000779

  • 77

    KonkleB. A.WalshC. E.EscobarM. A.JosephsonN. C.YoungG.von DrygalskiA.et al (2021). BAX 335 hemophilia B gene therapy clinical trial results: potential impact of CpG sequences on gene expression. Blood137, 763774. 10.1182/blood.2019004625

  • 78

    KosuriS.ChurchG. M. (2014). Large-scale de novo DNA synthesis: technologies and applications. Nat. Methods11, 499507. 10.1038/nmeth.2918

  • 79

    KozakM. (2005). Regulation of translation via mRNA structure in prokaryotes and eukaryotes. Gene361, 1337. 10.1016/j.gene.2005.06.037

  • 80

    KudlaG.LipinskiL.CaffinF.HelwakA.ZyliczM. (2006). High guanine and cytosine content increases mRNA levels in mammalian cells. PLoS Biol.4, e180. 10.1371/journal.pbio.0040180

  • 81

    KudlaG.MurrayA. W.TollerveyD.PlotkinJ. B. (2009). Coding-sequence determinants of gene expression in Escherichia coli. Science324, 255258. 10.1126/science.1170160

  • 82

    LeeI. T.NachbagauerR.EnszD.SchwartzH.CarmonaL.SchaefersK.et al (2023). Safety and immunogenicity of a phase 1/2 randomized clinical trial of a quadrivalent, mRNA-based seasonal influenza vaccine (mRNA-1010) in healthy adults: interim analysis. Nat. Commun.14, 3631. 10.1038/s41467-023-39376-7

  • 83

    LeppekK.ByeonG. W.KladwangW.Wayment-SteeleH. K.KerrC. H.XuA. F.et al (2022). Combinatorial optimization of mRNA structure, stability, and translation for RNA-based therapeutics. Nat. Commun.13, 1536. 10.1038/s41467-022-28776-w

  • 84

    LeRoyN.RoleckC. (2023). Optipyzer: a fast and flexible multi-species codon optimization server. bioRxiv, 2023.05.22.541759. 10.1101/2023.05.22.541759

  • 85

    LiC.HeY.NicolsonS.HirschM.WeinbergM. S.ZhangP.et al (2013). Adeno-associated virus capsid antigen presentation is dependent on endosomal escape. J. Clin. Invest.123, 13901401. 10.1172/JCI66611

  • 86

    LiuY. (2020). A code within the genetic code: codon usage regulates co-translational protein folding. Cell Commun. Signal.18, 145. 10.1186/s12964-020-00642-6

  • 87

    LoomisK. H.KirschmanJ. L.BhosleS.BellamkondaR. V.SantangeloP. J. (2016). Strategies for modulating innate immune activation and protein production of in vitro transcribed mRNAs. J. Mater. Chem. B4, 16191632. 10.1039/C5TB01753J

  • 88

    MaS.TangN.TianJ. (2012). DNA synthesis, assembly and applications in synthetic biology. Curr. Opin. Chem. Biol.16, 260267. 10.1016/j.cbpa.2012.05.001

  • 89

    MalarkannanS.HorngT.ShihP. P.SchwabS.ShastriN. (1999). Presentation of out-of-frame peptide/MHC class I complexes by a novel translation initiation mechanism. Immunity10, 681690. 10.1016/S1074-7613(00)80067-9

  • 90

    MartinoA. T.Basner-TschakarjanE.MarkusicD. M.FinnJ. D.HindererC.ZhouS.et al (2013). Engineered AAV vector minimizes in vivo targeting of transduced hepatocytes by capsid-specific CD8+ T cells. Blood121, 22242233. 10.1182/blood-2012-10-460733

  • 91

    MatsudaD.MauroV. P. (2010). Determinants of initiation codon selection during translation in mammalian cells. PLoS One5, e15057. 10.1371/journal.pone.0015057

  • 92

    MendellJ. R.Al-ZaidyS. A.Rodino-KlapacL. R.GoodspeedK.GrayS. J.KayC. N.et al (2021). Current clinical applications of in vivo gene therapy with AAVs. Mol. Ther.29, 464488. 10.1016/j.ymthe.2020.12.007

  • 93

    MitaraiN.SneppenK.PedersenS. (2008). Ribosome collisions and translation efficiency: optimization by codon usage and mRNA destabilization. J. Mol. Biol.382, 236245. 10.1016/j.jmb.2008.06.068

  • 94

    MuellerS. (2023). Challenges and opportunities of mRNA vaccines against SARS-CoV-2. Cham: Springer International Publishing. 10.1007/978-3-031-18903-6

  • 95

    MuellerS.PapamichailD.ColemanJ. R.SkienaS.WimmerE. (2006). Reduction of the rate of poliovirus protein synthesis through large-scale codon deoptimization causes attenuation of viral virulence by lowering specific infectivity. J. Virol.80, 96879696. 10.1128/JVI.00738-06

  • 96

    MulroneyT. E.PöyryT.Yam-PucJ. C.RustM.HarveyR. F.KalmarL.et al (2024). N1-methylpseudouridylation of mRNA causes +1 ribosomal frameshifting. Nature625, 189194. 10.1038/s41586-023-06800-3

  • 97

    NarulaA.EllisJ.TaliaferroJ. M.RisslandO. S. (2019). Coding regions affect mRNA stability in human cells. RNA25, 17511764. 10.1261/rna.073239.119

  • 98

    NavonS.PilpelY. (2011). The role of codon selection in regulation of translation efficiency deduced from synthetic libraries. Genome Biol.12, R12. 10.1186/gb-2011-12-2-r12

  • 99

    NieuwkoopT.TerlouwB. R.StevensK. G.ScheltemaR. A.de RidderD.van der OostJ.et al (2023). Revealing determinants of translation efficiency via whole-gene codon randomization and machine learning. Nucleic Acids Res.51, 23632376. 10.1093/nar/gkad035

  • 100

    Núñez-ManchónE.Farrera-SalM.Otero-MateoM.CastellanoG.MorenoR.MedelD.et al (2021). Transgene codon usage drives viral fitness and therapeutic efficacy in oncolytic adenoviruses. Nar. Cancer3, zcab015. 10.1093/narcan/zcab015

  • 101

    OliverS. E.GarganoJ. W.MarinM.WallaceM.CurranK. G.ChamberlandM.et al (2020). The advisory committee on immunization practices’ interim recommendation for use of pfizer-BioNTech COVID-19 vaccine — United States, december 2020. MMWR. Morb. Mortal. Wkly. Rep.69, 19221924. 10.15585/mmwr.mm6950e2

  • 102

    OwczarzyR.TataurovA. V.WuY.MantheyJ. A.McQuistenK. A.AlmabraziH. G.et al (2008). IDT SciTools: a suite for analysis and design of nucleic acid oligomers. Nucleic Acids Res.36, W163W169. 10.1093/nar/gkn198

  • 103

    PallukS.ArlowD. H.de RondT.BarthelS.KangJ. S.BectorR.et al (2018). De novo DNA synthesis using polymerase-nucleotide conjugates. Nat. Biotechnol.36, 645650. 10.1038/nbt.4173

  • 104

    PereiraI. T.SpangenbergL.RobertA. W.AmorínR.StimamiglioM. A.NayaH.et al (2018). Polysome profiling followed by RNA-seq of cardiac differentiation stages in hESCs. Sci. Data5, 180287. 10.1038/sdata.2018.287

  • 105

    PerlakF. J.FuchsR. L.DeanD. A.McPhersonS. L.FischhoffD. A. (1991). Modification of the coding sequence enhances plant expression of insect control protein genes. Proc. Natl. Acad. Sci.88, 33243328. 10.1073/pnas.88.8.3324

  • 106

    PhamT. D.O’ConnellJ.CraneD. I. (2004). “Constrained codon optimization by dynamic programming,” in Proceedings of 2004 International Symposium on Intelligent Multimedia, Video and Speech Processing, 2004 (IEEE), 153156. 10.1109/ISIMP.2004.1434023

  • 107

    PinkardO.McFarlandS.SweetT.CollerJ. (2020). Quantitative tRNA-sequencing uncovers metazoan tissue-specific tRNA regulation. Nat. Commun.11, 4104. 10.1038/s41467-020-17879-x

  • 108

    PitoisetF.VazquezT.LevacherB.Nehar-BelaidD.DérianN.VigneronJ.et al (2017). Retrovirus-based virus-like particle immunogenicity and its modulation by toll-like receptor activation. J. Virol.91, e01230-17. 10.1128/JVI.01230-17

  • 109

    PizzoL.IriarteA.Alvarez-ValinF.MarínM. (2015). Conservation of CFTR codon frequency through primates suggests synonymous mutations could have a functional effect. Mutat. Res. Mol. Mech. Mutagen.775, 1925. 10.1016/j.mrfmmm.2015.03.005

  • 110

    PlotkinJ. B.RobinsH.LevineA. J. (2004). Tissue-specific codon usage and the expression of human genes. Proc. Natl. Acad. Sci.101, 1258812591. 10.1073/pnas.0404957101

  • 111

    PouyetF.MouchiroudD.DuretL.SémonM. (2017). Recombination, meiotic expression and human codon usage. Elife6, e27344. 10.7554/eLife.27344

  • 112

    PresnyakV.AlhusainiN.ChenY.-H.MartinS.MorrisN.KlineN.et al (2015). Codon optimality is a major determinant of mRNA stability. Cell160, 11111124. 10.1016/j.cell.2015.02.029

  • 113

    PuigboP.GuzmanE.RomeuA.Garcia-VallveS. (2007). OPTIMIZER: a web server for optimizing the codon usage of DNA sequences. Nucleic Acids Res.35, W126W131. 10.1093/nar/gkm219

  • 114

    RaabA. M.GebhardtG.BolotinaN.Weuster-BotzD.LangC. (2010). Metabolic engineering of Saccharomyces cerevisiae for the biotechnological production of succinic acid. Metab. Eng.12, 518525. 10.1016/j.ymben.2010.08.005

  • 115

    RahmanS. U.YaoX.LiX.ChenD.TaoS. (2018). Analysis of codon usage bias of Crimean-Congo hemorrhagic fever virus and its adaptation to hosts. Infect. Genet. Evol.58, 116. 10.1016/j.meegid.2017.11.027

  • 116

    ReisM. d. (2004). Solving the riddle of codon usage preferences: a test for translational selection. Nucleic Acids Res.32, 50365044. 10.1093/nar/gkh834

  • 117

    RingnérM.KroghM. (2005). Folding free energies of 5′-UTRs impact post-transcriptional regulation on a genomic scale in yeast. PLoS Comput. Biol.1, e72. 10.1371/journal.pcbi.0010072

  • 118

    RodriguezA.WrightG.EmrichS.ClarkP. L. (2018). %MinMax: a versatile tool for calculating and comparing synonymous codon usage and its impact on protein folding. Protein Sci.27, 356362. 10.1002/pro.3336

  • 119

    RogersG. L.ShirleyJ. L.ZolotukhinI.KumarS. R. P.ShermanA.PerrinG. Q.et al (2017). Plasmacytoid and conventional dendritic cells cooperate in crosspriming AAV capsid-specific CD8+ T cells. Blood129, 31843195. 10.1182/blood-2016-11-751040

  • 120

    RojasM.Restrepo-JiménezP.MonsalveD. M.PachecoY.Acosta-AmpudiaY.Ramírez-SantanaC.et al (2018). Molecular mimicry and autoimmunity. J. Autoimmun.95, 100123. 10.1016/j.jaut.2018.10.012

  • 121

    RöltgenK.NielsenS. C. A.SilvaO.YounesS. F.ZaslavskyM.CostalesC.et al (2022). Immune imprinting, breadth of variant recognition, and germinal center response in human SARS-CoV-2 infection and vaccination. Cell185, 10251040.e14. 10.1016/j.cell.2022.01.018

  • 122

    RonkA. J.LloydN. M.ZhangM.AtyeoC.PerrettH. R.MireC. E.et al (2023). A Lassa virus mRNA vaccine confers protection but does not require neutralizing antibody in a Guinea pig model of infection. Nat. Commun.14, 5603. 10.1038/s41467-023-41376-6

  • 123

    RoymondalU.DasS.SahooS. (2009). Predicting gene expression level from relative codon usage bias: an application to Escherichia coli genome. DNA Res.16, 1330. 10.1093/dnares/dsn029

  • 124

    SabiR.TullerT. (2014). Modelling the efficiency of codon–tRNA interactions based on codon usage bias. DNA Res.21, 511526. 10.1093/dnares/dsu017

  • 125

    SabiR.Volvovitch DanielR.TullerT. (2017). stAIcalc: tRNA adaptation index calculator based on species-specific weights. Bioinformatics33, 589591. 10.1093/bioinformatics/btw647

  • 126

    SatoK.HamadaM. (2023). Recent trends in RNA informatics: a review of machine learning and deep learning for RNA secondary structure prediction and RNA drug discovery. Brief. Bioinform.24, bbad186. 10.1093/bib/bbad186

  • 127

    SharpP.LiW.-H. (1986). Codon usage in regulatory genes in Escherichia coli does not reflect selection for ‘rare’ codons. Nucleic Acids Res.14, 77377749. 10.1093/nar/14.19.7737

  • 128

    SharpP. M.LiW.-H. (1987). The codon adaptation index-a measure of directional synonymous codon usage bias, and its potential applications. Nucleic Acids Res.15, 12811295. 10.1093/nar/15.3.1281

  • 129

    ShiF.FanZ.ZhangS.WangY.TanS.LiY. (2020). Optimization of ribosomal binding site sequences for gene expression and 4-hydroxyisoleucine biosynthesis in recombinant corynebacterium glutamicum. Enzyme Microb. Technol.140, 109622. 10.1016/j.enzmictec.2020.109622

  • 130

    ShirleyJ. L.de JongY. P.TerhorstC.HerzogR. W. (2020a). Immune responses to viral gene therapy vectors. Mol. Ther.28, 709722. 10.1016/j.ymthe.2020.01.001

  • 131

    ShirleyJ. L.KeelerG. D.ShermanA.ZolotukhinI.MarkusicD. M.HoffmanB. E.et al (2020b). Type I IFN sensing by cDCs and CD4+ T cell help are both requisite for cross-priming of AAV capsid-specific CD8+ T cells. Mol. Ther.28, 758770. 10.1016/j.ymthe.2019.11.011

  • 132

    SimonC. S.HadjantonakisA.SchröterC. (2018). Making lineage decisions with biological noise: lessons from the early mouse embryo. WIREs Dev. Biol.7, e319. 10.1002/wdev.319

  • 133

    SinyakovA. N.RyabininV. A.KostinaE. V. (2021). Application of array-based oligonucleotides for synthesis of genetic designs. Mol. Biol.55, 487500. 10.1134/S0026893321030109

  • 134

    SongL.-F.DengZ.-H.GongZ.-Y.LiL.-L.LiB.-Z. (2021). Large-scale de novo oligonucleotide synthesis for whole-genome synthesis and data storage: challenges and opportunities. Front. Bioeng. Biotechnol.9, 689797. 10.3389/fbioe.2021.689797

  • 135

    StenicoM.LloydA. T.SharpP. M. (1994). Codon usage in Caenorhabditis elegans: delineation of translational selection and mutational biases. Nucleic Acids Res.22, 24372446. 10.1093/nar/22.13.2437

  • 136

    SunB.ZhangH.FrancoL. M.BrownT.BirdA.SchneiderA.et al (2005). Correction of glycogen storage disease type II by an adeno-associated virus vector containing a muscle-specific promoter. Mol. Ther.11, 889898. 10.1016/j.ymthe.2005.01.012

  • 137

    TanedaA.AsaiK. (2020). COSMO: a dynamic programming algorithm for multicriteria codon optimization. Comput. Struct. Biotechnol. J.18, 18111818. 10.1016/j.csbj.2020.06.035

  • 138

    ThessA.GrundS.MuiB. L.HopeM. J.BaumhofP.Fotin-MleczekM.et al (2015). Sequence-engineered mRNA without chemical nucleoside modifications enables an effective protein therapy in large animals. Mol. Ther.23, 14561464. 10.1038/mt.2015.103

  • 139

    ThomasD. R.WalmsleyA. M. (2014). Improved expression of recombinant plant-made hEGF. Plant Cell Rep.33, 18011814. 10.1007/s00299-014-1658-8

  • 140

    ThomasO. G.BrongeM.TengvallK.AkpinarB.NilssonO. B.HolmgrenE.et al (2023). Cross-reactive EBNA1 immunity targets alpha-crystallin B and is associated with multiple sclerosis. Sci. Adv.9, eadg303214. 10.1126/sciadv.adg3032

  • 141

    ThulP. J.LindskogC. (2018). The human protein atlas: a spatial map of the human proteome. Protein Sci.27, 233244. 10.1002/pro.3307

  • 142

    VillanuevaE.Martí-SolanoM.FillatC. (2016). Codon optimization of the adenoviral fiber negatively impacts structural protein expression and viral fitness. Sci. Rep.6, 27546. 10.1038/srep27546

  • 143

    WanJ.YangJ.WangZ.ShenR.ZhangC.WuY.et al (2023). A single immunization with core–shell structured lipopolyplex mRNA vaccine against rabies induces potent humoral immunity in mice and dogs. Emerg. Microbes Infect.12, 2270081. 10.1080/22221751.2023.2270081

  • 144

    WanX.-F.ZhouJ.XuD. (2006). CodonO: a new informatics method for measuring synonymous codon usage bias within and across genomes. Int. J. Gen. Syst.35, 109125. 10.1080/03081070500502967

  • 145

    Wayment-SteeleH. K.KimD. S.ChoeC. A.NicolJ. J.Wellington-OguriR.WatkinsA. M.et al (2021). Theoretical basis for stabilizing messenger RNA through secondary structure design. Nucleic Acids Res.49, 1060410617. 10.1093/nar/gkab764

  • 146

    WeiY.SilkeJ. R.XiaX. (2019). An improved estimation of tRNA expression to better elucidate the coevolution between tRNA abundance and codon usage in bacteria. Sci. Rep.9, 3184. 10.1038/s41598-019-39369-x

  • 147

    WelchM.VillalobosA.GustafssonC.MinshullJ. (2009). You’re one in a googol: optimizing genes for protein expression. J. R. Soc. Interface6, S467S476. 10.1098/rsif.2008.0520.focus

  • 148

    WrightF. (1990). The ‘effective number of codons’ used in a gene. Gene87, 2329. 10.1016/0378-1119(90)90491-9

  • 149

    WrightG.RodriguezA.LiJ.MilenkovicT.EmrichS. J.ClarkP. L. (2022). CHARMING: harmonizing synonymous codon usage to replicate a desired codon usage pattern. Protein Sci.31, 221231. 10.1002/pro.4223

  • 150

    WrightJ. F. (2020). Quantification of CpG motifs in rAAV genomes: avoiding the Toll. Mol. Ther.28, 17561758. 10.1016/j.ymthe.2020.07.006

  • 151

    WuQ.MedinaS. G.KushawahG.DeVoreM. L.CastellanoL. A.HandJ. M.et al (2019). Translation affects mRNA stability in a codon-dependent manner in human cells. Elife8, e45396. 10.7554/eLife.45396

  • 152

    WuX.ShanK.ZanF.TangX.QianZ.LuJ. (2023). Optimization and deoptimization of codons in SARS‐CoV‐2 and related implications for vaccine development. Adv. Sci.10, e2205445. 10.1002/advs.202205445

  • 153

    XiaX. (2015). A major controversy in codon-anticodon adaptation resolved by a new codon usage index. Genetics199, 573579. 10.1534/genetics.114.172106

  • 154

    XiaX. (2021). Detailed dissection and critical evaluation of the pfizer/BioNTech and Moderna mRNA vaccines. Vaccines9, 734. 10.3390/vaccines9070734

  • 155

    XueC.ChuQ.ZhengQ.JiangS.BaoZ.SuY.et al (2022). Role of main RNA modifications in cancer: N6-methyladenosine, 5-methylcytosine, and pseudouridine. Signal Transduct. Target. Ther.7, 142. 10.1038/s41392-022-01003-0

  • 156

    YangT. yuanBraunM.LembkeW.McBlaneF.KamerudJ.DeWallS.et al (2022). Immunogenicity assessment of AAV-based gene therapies: an IQ consortium industry white paper. Mol. Ther. - Methods Clin. Dev.26, 471494. 10.1016/j.omtm.2022.07.018

  • 157

    YewN. S.ChengS. H. (2004). Reducing the immunostimulatory activity of CpG-containing plasmid DNA vectors for non-viral gene therapy. Expert Opin. Drug Deliv.1, 115125. 10.1517/17425247.1.1.115

  • 158

    ZhangH.ZhangL.LinA.XuC.LiZ.LiuK.et al (2023). Algorithm for optimized mRNA design improves stability and immunogenicity. Nature621, 396403. 10.1038/s41586-023-06127-z

  • 159

    ZhangZ.LiJ.CuiP.DingF.LiA.TownsendJ. P.et al (2012). Codon Deviation Coefficient: a novel measure for estimating codon usage bias and its statistical significance. BMC Bioinforma.13, 43. 10.1186/1471-2105-13-43

  • 160

    ZukerM. (1994). “Prediction of RNA secondary structure by energy minimization,” in Computer analysis of sequence data (Totowa, NJ: Humana Press), 267294. 10.1385/0-89603-276-0:267

  • 161

    ZukerM.StieglerP. (1981). Optimal computer folding of large RNA sequences using thermodynamics and auxiliary information. Nucleic Acids Res.9, 133148. 10.1093/nar/9.1.133

Summary

Keywords

gene therapy, codon-optimization metrics, mRNA, immunogenicity, clinical trials

Citation

Paremskaia AI, Kogan AA, Murashkina A, Naumova DA, Satish A, Abramov IS, Feoktistova SG, Mityaeva ON, Deviatkin AA and Volchkov PY (2024) Codon-optimization in gene therapy: promises, prospects and challenges. Front. Bioeng. Biotechnol. 12:1371596. doi: 10.3389/fbioe.2024.1371596

Received

16 January 2024

Accepted

19 March 2024

Published

28 March 2024

Volume

12 - 2024

Edited by

Yaroslava G. Yingling, North Carolina State University, United States

Reviewed by

Siguna Mueller, Independent Researcher, Kaernten, Austria

Clement T. Y. Chan, University of North Texas, United States

Updates

Copyright

*Correspondence: Andrei A. Deviatkin, ; Pavel Yu Volchkov,

† These authors have contributed equally to this work and share last authorship

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics