Old version — revision 1
This is a fixed snapshot of Genetic code expansion and recoding, saved by Import as part of the initial corpus import. It is not edited and it is not updated; the article may have changed since.
Edit summary: Initial import of content/recoded-organisms.md — the filesystem corpus, unchanged. Not an edit.
Altering how a cell reads DNA so that reassigned codons encode amino acids not found in nature, producing organisms with chemistry and immunity no natural cell has.
Genetic code expansion and recoding are two halves of one project: changing the rules by which a cell translates DNA into protein. Expansion adds a new amino acid to an organism's repertoire by assigning it a codon; recoding clears that codon first, by replacing every natural instance of it across the genome and deleting the machinery that reads it. The result is a cell that builds proteins from a chemistry unavailable to any natural organism — and, as a side effect that has proved at least as interesting, one that viruses cannot read.
Changing the genetic code means changing two things that normally match: the meaning a cell assigns to a codon, and the sequences written in the old meaning. Expansion handles the first and is comparatively easy. Recoding handles the second and is the reason the field depends on genome-scale DNA synthesis.
Protein synthesis assigns amino acids to codons through aminoacyl-tRNA synthetases, each of which charges a particular transfer RNA with a particular amino acid. To add a twenty-first amino acid, an engineer supplies a synthetase–tRNA pair that ignores, and is ignored by, everything already in the cell — an orthogonal pair. The synthetase is then evolved in the laboratory to accept the desired non-standard amino acid, and the tRNA is given an anticodon matching whichever codon has been set aside.
The first such system, reported in 2001, used a pair borrowed from an archaeon and assigned a modified tyrosine to the amber stop codon.1 The pair most used since derives from the pyrrolysine system of methanogenic archaea, which is orthogonal in bacterial, yeast and mammalian cells alike and tolerates an unusually wide range of substrates.
Assigning a stop codon to a new amino acid works, but imperfectly: the cell's release factor still competes for the same codon, so incorporation is inefficient and native proteins are occasionally read through. The clean solution is to remove every instance of the codon from the genome and then delete the release factor. That is genome surgery at thousands of positions, well beyond the reach of one-site-at-a-time tools such as CRISPR–Cas9 or Prime editing, and it is why the field is inseparable from Synthetic genomes.
The first genomically recoded organism, reported in 2013 by a group including George Church, replaced all 321 amber stop codons in Escherichia coli with an alternative stop and deleted the corresponding release factor.2 A later effort rebuilt the entire four-megabase genome from synthetic fragments to remove two serine codons as well as the amber codon, then deleted the transfer RNAs that read them, freeing three codons at once.34
Why viruses cannot read a recoded genomeA bacteriophage genome is written in the standard code. Injected into a cell that has deleted the tRNAs or release factor for particular codons, the viral messenger RNAs stall or are mistranslated, and no functional virions are produced. Resistance is therefore structural rather than a defence the virus can evolve around, since escaping it would require the phage to rewrite its own coding sequence throughout.
Recoding at genome scale exists only in bacteria. Code expansion — adding one or two non-standard amino acids without clearing their codons — works routinely in yeast, in cultured mammalian cells, and in whole animals including mice, where it is used to make proteins that can be switched on with light or cross-linked to their binding partners on command. No recoded mammalian genome exists, and none is close: the recoding of a bacterial genome required total synthesis of four megabases, and a human genome is close to a thousand times that size.
The commercial application that reached medicine first is protein conjugation. Placing a non-standard amino acid with a chemically distinct handle at a chosen position in an antibody allows a drug payload to be attached at exactly that site and nowhere else, giving a homogeneous product rather than the mixture produced by conventional chemistry. Antibody–drug conjugates built this way have advanced through clinical development, and the approach is now one of the more mature strands of Targeted drug delivery.
Recoded strains generally grow more slowly than their parents, and each additional freed codon compounds the fitness cost. Incorporation efficiency falls as more non-standard amino acids are encoded in the same protein, and the amino acids themselves must be supplied in the medium or synthesized by an engineered pathway. Orthogonality is never perfect: engineered synthetases retain some activity on natural substrates, and engineered tRNAs are occasionally charged by native enzymes.
The deeper constraint is that recoding requires writing a genome, so the technology inherits every limit of DNA synthesis and assembly. Until mammalian chromosome construction is routine, extending recoding beyond microbes is a proposal rather than a programme.
The most developed safety application inverts the usual worry about engineered organisms. Rather than relying on physical containment, essential proteins in a recoded strain are redesigned so that they fold only when a synthetic amino acid is incorporated. The organism then cannot survive outside a laboratory that supplies the compound, and measured escape rates in the founding studies fell below the limit of detection.5 Code swapping adds a second layer: engineered genes written in a non-standard code are mistranslated if they are transferred horizontally into a natural organism, so the construct does not function even if the DNA escapes.6
These are the strongest existing demonstrations that engineered biology can be made intrinsically, rather than procedurally, contained — a contrast with the self-spreading elements described in Gene drives, where containment is the unsolved problem. They remain laboratory results at modest scale, and no regulator currently treats synthetic auxotrophy as sufficient grounds for reduced physical containment.
The risks run in the other direction too. A cell immune to all natural viruses is also a cell that natural ecosystems have no mechanism to control, which is one of the concerns that motivated the warnings collected in Mirror life about organisms placed outside the reach of existing biology. Recoding is discussed in Dual-use research of concern assessments mainly as a capability multiplier rather than a direct hazard: the same synthesis and assembly infrastructure serves both, and the Precautionary principle arguments applied to genome writing apply here with the same force.
Two goals define the near-term agenda. The first is bacterial: strains with several freed codons that can polymerise non-biological monomers, turning cells into programmable chemical factories for molecules the ribosome was never built to make. The second is mammalian, and much further off — a virus-resistant human cell line for biomanufacturing, proposed as the flagship of human genome writing, which would remove the contamination risk that shuts down production runs. A related and more tractable version of the same logic is already in clinical use in Xenotransplantation, where donor pigs have had endogenous retroviral sequences inactivated wholesale rather than recoded.
Whether recoding ever reaches a human organism rather than a human cell line is a separate question and not an active one. It would require rewriting every cell of a person, which means doing it at the embryo stage, and so falls squarely within the prohibitions on Human germline editing surveyed in Governance of human genome editing. The proposal occasionally surfaces in discussions of engineered virus immunity; nobody has proposed a credible route, and the mismatch between a laboratory strain that grows slowly on a defined medium and a functioning mammal is the reason.
paperWang, L., Brock, A., Herberich, B., Schultz, P. G. "Expanding the genetic code of Escherichia coli." Science, 2001. ↩
paperLajoie, M. J. et al. "Genomically recoded organisms expand biological functions." Science, 2013.↩The codons were replaced by editing a living strain rather than by synthesizing the genome, and the resulting resistance was to particular phages, not the broad resistance reported later.
paperFredens, J. et al. "Total synthesis of Escherichia coli with a recoded genome." Nature, 2019. ↩
paperRobertson, W. E. et al. "Sense codon reassignment enables viral resistance and encoded polymer synthesis." Science, 2021. ↩
paperMandell, D. J. et al. "Biocontainment of genetically modified organisms by synthetic protein design." Nature, 2015.↩A detection-limit result rather than a measured escape rate: the assay bounds escape at the population sizes tested in laboratory culture.
paperNyerges, A. et al. "A swapped genetic code prevents viral infections and gene transfer." Nature, 2023. ↩