AI-Generated · xiaomi/mimo-v2.5-pro

E. coli's Transcription Machinery Reads an Eight-Letter Genetic Alphabet

A UC San Diego study shows that E. coli RNA polymerase faithfully transcribes an eight-letter genetic alphabet, with cryo-EM structures revealing the enzyme handles synthetic base pairs through the same mechanisms it uses for natural ones.

E. coli's Transcription Machinery Reads an Eight-Letter Genetic Alphabet
Domain organization of σ70 in an E. coli RNA polymerase transcription initiation complex (PDB: 4YLN). Illustration by Mark S. Paget, 2015.
Photo: Mark S. Paget, CC BY 4.0

Every known organism on Earth runs its genetic information through the same four-letter alphabet — A, T, C, G — and has for roughly four billion years. A peer-reviewed study published September 2 in Nature Communications shows that the molecular machine responsible for reading that alphabet, at least in E. coli, can handle a much wider one. Using cryo-EM imaging at resolutions between 2.42 and 2.75 angstroms, a UC San Diego–led team demonstrated that E. coli RNA polymerase faithfully transcribes an eight-letter genetic alphabet that includes two pairs of synthetic bases — P:Z and B:S — orthogonally placed alongside the natural ones. The enzyme doesn’t tolerate them as compromises; it processes them through the same recognition mechanisms it uses for A, T, C, and G.

The expanded alphabet belongs to a system called Hachimoji DNAhachi being Japanese for eight — and it has been a research goal for years in synthetic biology. What made the UC San Diego work significant was showing that a natural, unmodified polymerase could handle the job. RNA polymerase is the enzyme that transcribes DNA into RNA; if the transcription machinery chokes on synthetic bases, any expanded-code design stays trapped at the bench-top chemistry stage. The study’s cryo-EM structures revealed that the enzyme stabilizes the unnatural base pairs using hydrogen-bonding networks similar to those formed by the natural pairs, which goes a long way toward explaining the transcription fidelity.

Independent coverage published September 5 by The Economic Times confirmed the core finding and pointed toward the downstream applications: an eight-letter code could carry significantly more genetic information per unit of sequence length, which matters for diagnostics that depend on distinctive molecular signatures and for targeted therapies that need to avoid off-target binding. More letters also mean more capacity for encoding non-standard amino acids or molecular switches directly into a genetic payload.

The implication extends beyond any single enzyme. If E. coli’s transcription machinery — evolved entirely within a four-letter constraint — reads an eight-letter alphabet accurately, the barrier to distributing expanded genetic codes across biological systems is lower than it might have seemed. The cell’s own machinery does not inherently reject information it was never designed to carry, provided the synthetic base pairs are geometrically and chemically compatible with the active site.

That does not mean hachimoji organisms are around the corner. Transcription is one step; translation — the process by which the cell reads RNA and builds proteins — introduces a separate set of recognition problems, and nothing in the current study addresses ribosomal fidelity on eight-letter codons. But as a proof-of-concept at the transcription level, the result is clean: a natural polymerase, a synthetic alphabet, and high-resolution structural evidence for why it works.

Sources