Genetic code expansion and recoding are two halves of one project: changing the rules by which a cell translates DNA into protein. Expansion adds a new amino acid to an organism's repertoire by assigning it a codon; recoding clears that codon first, by replacing every natural instance of it across the genome and deleting the machinery that reads it. The result is a cell that builds proteins from a chemistry unavailable to any natural organism — and, as a side effect that has proved at least as interesting, one that viruses cannot read.
How it works
Changing the genetic code means changing two things that normally match: the meaning a cell assigns to a codon, and the sequences written in the old meaning. Expansion handles the first and is comparatively easy. Recoding handles the second and is the reason the field depends on genome-scale DNA synthesis.
Orthogonal translation
Protein synthesis assigns amino acids to codons through aminoacyl-tRNA synthetases, each of which charges a particular transfer RNA with a particular amino acid. To add a twenty-first amino acid, an engineer supplies a synthetase–tRNA pair that ignores, and is ignored by, everything already in the cell — an orthogonal pair. The synthetase is then evolved in the laboratory to accept the desired non-standard amino acid, and the tRNA is given an anticodon matching whichever codon has been set aside.
The first such system, reported in 2001, used a pair borrowed from an archaeon and assigned a modified tyrosine to the amber stop codon.1 The pair most used since derives from the pyrrolysine system of methanogenic archaea, which is orthogonal in bacterial, yeast and mammalian cells alike and tolerates an unusually wide range of substrates.
Freeing a codon
Assigning a stop codon to a new amino acid works, but imperfectly: the cell's release factor still competes for the same codon, so incorporation is inefficient and native proteins are occasionally read through. The clean solution is to remove every instance of the codon from the genome and then delete the release factor. That is genome surgery at thousands of positions, well beyond the reach of one-site-at-a-time tools such as CRISPR–Cas9 or Prime editing, and it is why the field is inseparable from Synthetic genomes.
The first genomically recoded organism, reported in 2013 by a group including George Church, replaced all 321 amber stop codons in Escherichia coli with an alternative stop and deleted the corresponding release factor.2 A later effort rebuilt the entire four-megabase genome from synthetic fragments to remove two serine codons as well as the amber codon, then deleted the transfer RNAs that read them, freeing three codons at once.34
Why viruses cannot read a recoded genomeA bacteriophage genome is written in the standard code. Injected into a cell that has deleted the tRNAs or release factor for particular codons, the viral messenger RNAs stall or are mistranslated, and no functional virions are produced. Resistance is therefore structural rather than a defence the virus can evolve around, since escaping it would require the phage to rewrite its own coding sequence throughout.
-
2001First expanded codeSchultz's group incorporates a non-standard amino acid into a protein in living E. coli using an orthogonal synthetase–tRNA pair assigned to the amber codon.
-
2010Orthogonal ribosomesEngineered ribosomes that read four-base codons allow multiple distinct non-standard amino acids to be encoded in one protein.
-
2013First genomically recoded organismAll 321 amber codons are removed from the E. coli genome and the release factor deleted, giving efficient incorporation and partial phage resistance.
-
2015Synthetic auxotrophyEssential proteins are redesigned to require a synthetic amino acid, making the strain dependent on a compound available only in the laboratory.
-
2019–2021Three codons freedA fully synthetic 4-megabase E. coli genome removes two serine codons and the amber codon; deleting the corresponding tRNAs yields broad resistance to bacteriophages.
-
2023Code swappingReassigning freed codons to new meanings is shown to block viral propagation and to prevent engineered genes from functioning if transferred to natural organisms.
Current state
Recoding at genome scale exists only in bacteria. Code expansion — adding one or two non-standard amino acids without clearing their codons — works routinely in yeast, in cultured mammalian cells, and in whole animals including mice, where it is used to make proteins that can be switched on with light or cross-linked to their binding partners on command. No recoded mammalian genome exists, and none is close: the recoding of a bacterial genome required total synthesis of four megabases, and a human genome is close to a thousand times that size.
The commercial application that reached medicine first is protein conjugation. Placing a non-standard amino acid with a chemically distinct handle at a chosen position in an antibody allows a drug payload to be attached at exactly that site and nowhere else, giving a homogeneous product rather than the mixture produced by conventional chemistry. Antibody–drug conjugates built this way have advanced through clinical development, and the approach is now one of the more mature strands of Targeted drug delivery.
Limitations
Recoded strains generally grow more slowly than their parents, and each additional freed codon compounds the fitness cost. Incorporation efficiency falls as more non-standard amino acids are encoded in the same protein, and the amino acids themselves must be supplied in the medium or synthesized by an engineered pathway. Orthogonality is never perfect: engineered synthetases retain some activity on natural substrates, and engineered tRNAs are occasionally charged by native enzymes.
The deeper constraint is that recoding requires writing a genome, so the technology inherits every limit of DNA synthesis and assembly. Until mammalian chromosome construction is routine, extending recoding beyond microbes is a proposal rather than a programme.
Biocontainment and risk
The most developed safety application inverts the usual worry about engineered organisms. Rather than relying on physical containment, essential proteins in a recoded strain are redesigned so that they fold only when a synthetic amino acid is incorporated. The organism then cannot survive outside a laboratory that supplies the compound, and measured escape rates in the founding studies fell below the limit of detection.5 Code swapping adds a second layer: engineered genes written in a non-standard code are mistranslated if they are transferred horizontally into a natural organism, so the construct does not function even if the DNA escapes.6
These are the strongest existing demonstrations that engineered biology can be made intrinsically, rather than procedurally, contained — a contrast with the self-spreading elements described in Gene drives, where containment is the unsolved problem. They remain laboratory results at modest scale, and no regulator currently treats synthetic auxotrophy as sufficient grounds for reduced physical containment.
The risks run in the other direction too. A cell immune to all natural viruses is also a cell that natural ecosystems have no mechanism to control, which is one of the concerns that motivated the warnings collected in Mirror life about organisms placed outside the reach of existing biology. Recoding is discussed in Dual-use research of concern assessments mainly as a capability multiplier rather than a direct hazard: the same synthesis and assembly infrastructure serves both, and the Precautionary principle arguments applied to genome writing apply here with the same force.
Outlook
Two goals define the near-term agenda. The first is bacterial: strains with several freed codons that can polymerise non-biological monomers, turning cells into programmable chemical factories for molecules the ribosome was never built to make. The second is mammalian, and much further off — a virus-resistant human cell line for biomanufacturing, proposed as the flagship of human genome writing, which would remove the contamination risk that shuts down production runs. A related and more tractable version of the same logic is already in clinical use in Xenotransplantation, where donor pigs have had endogenous retroviral sequences inactivated wholesale rather than recoded.
Whether recoding ever reaches a human organism rather than a human cell line is a separate question and not an active one. It would require rewriting every cell of a person, which means doing it at the embryo stage, and so falls squarely within the prohibitions on Human germline editing surveyed in Governance of human genome editing. The proposal occasionally surfaces in discussions of engineered virus immunity; nobody has proposed a credible route, and the mismatch between a laboratory strain that grows slowly on a defined medium and a functioning mammal is the reason.
See also
- Synthetic genomes
- CRISPR–Cas9
- Mirror life
- Gene drives
- Dual-use research of concern
- George Church
- Targeted drug delivery
- Xenotransplantation
References
Footnotes
-
paperWang, L., Brock, A., Herberich, B., Schultz, P. G. "Expanding the genetic code of Escherichia coli." Science, 2001. ↩
-
paperLajoie, M. J. et al. "Genomically recoded organisms expand biological functions." Science, 2013.↩The codons were replaced by editing a living strain rather than by synthesizing the genome, and the resulting resistance was to particular phages, not the broad resistance reported later.
-
paperFredens, J. et al. "Total synthesis of Escherichia coli with a recoded genome." Nature, 2019. ↩
-
paperRobertson, W. E. et al. "Sense codon reassignment enables viral resistance and encoded polymer synthesis." Science, 2021. ↩
-
paperMandell, D. J. et al. "Biocontainment of genetically modified organisms by synthetic protein design." Nature, 2015.↩A detection-limit result rather than a measured escape rate: the assay bounds escape at the population sizes tested in laboratory culture.
-
paperNyerges, A. et al. "A swapped genetic code prevents viral infections and gene transfer." Nature, 2023. ↩