The genetic code is the correspondence between sequences of nucleotides in messenger RNA (mRNA) and the amino acids assembled into a protein. During translation, cellular machinery reads three nucleotides at a time; each triplet, or codon, specifies an amino acid or a signal to terminate synthesis. The code is therefore a decoding system, not the particular genetic sequence of an individual or species. Most organisms share the standard genetic code, although several natural variants exist. (ncbi.nlm.nih.gov)
Organization of the standard code
RNA contains four principal nucleotide bases: adenine, uracil, guanine, and cytosine, abbreviated A, U, G, and C. Their possible three-base combinations produce 4³, or 64, codons. In the standard code, 61 specify the 20 canonical amino acids, while UAA, UAG, and UGA serve as termination signals. Codons are conventionally written in the 5′-to-3′ direction of the mRNA. DNA sequence tables may instead use T, representing thymine, in place of U. (ncbi.nlm.nih.gov)
The code is degenerate, meaning that multiple codons can specify the same amino acid. Leucine, serine, and arginine each have six codons; methionine and tryptophan each have only one in the standard code. Codons with the same amino-acid assignment are called synonymous codons. Degeneracy does not ordinarily mean ambiguity: within a given code and decoding context, a codon has a defined assignment. (ncbi.nlm.nih.gov)
AUG specifies methionine and is the most common initiation codon. Its role as a start signal depends on surrounding sequences and initiation machinery, rather than on the triplet alone. Alternative initiation codons occur, particularly in bacteria and organelles; their initiation assignment can differ from their meaning during elongation. (ncbi.nlm.nih.gov)
Reading frames and molecular decoding
For a protein-coding gene, transcription produces an RNA copy of genetic information. Translation then proceeds along the mRNA in a particular reading frame, the grouping of successive nucleotides into triplets. Within an ordinary coding region, codons are read consecutively without intervening punctuation and without overlap. Moving the starting position by one nucleotide changes the grouping and usually the resulting protein sequence. (ncbi.nlm.nih.gov)
The ribosome coordinates decoding, while transfer RNAs (tRNAs) act as adaptors. Each tRNA carries an amino acid and contains an anticodon that pairs with an mRNA codon. Aminoacyl-tRNA synthetases are enzymes that attach appropriate amino acids to their corresponding tRNAs. The relationship between codons and amino acids is thus implemented by the translation apparatus, rather than by direct recognition between an amino acid and its codon. (ncbi.nlm.nih.gov)
Pairing at the third codon position can accommodate certain nonstandard base interactions. This wobble base pairing allows one tRNA to recognize more than one synonymous codon. Modified bases in tRNA further influence recognition. Standard stop codons are normally recognized by release factors rather than amino-acid-bearing tRNAs, leading to release of the completed polypeptide. (ncbi.nlm.nih.gov)
Experimental decipherment
The code was deciphered through genetic experiments and biochemical systems that synthesized proteins outside intact cells. In 1961, Marshall Nirenberg and Heinrich Matthaei demonstrated that synthetic RNA consisting of repeated uracil residues directed the formation of polyphenylalanine. This established UUU as a phenylalanine codon. (njc.rockefeller.edu)
Subsequent experiments using synthetic RNAs and codon-recognition assays established the remaining assignments. Har Gobind Khorana and other researchers contributed methods using defined RNA sequences, and the standard code had been deciphered by 1966. Nirenberg, Khorana, and Robert Holley received the 1968 Nobel Prize in Physiology or Medicine for interpreting the genetic code and its function in protein synthesis. (genome.gov)
Universality and natural exceptions
The broad conservation of the code across bacteria, archaea, and eukaryotes is consistent with shared ancestry. Nevertheless, “universal code” is an approximation. Different translation tables are required for some nuclear and organellar genomes. (ncbi.nlm.nih.gov)
Well-established exceptions occur in mitochondria. In the vertebrate mitochondrial code, UGA specifies tryptophan rather than termination, and AUA specifies methionine rather than isoleucine. Some ciliates assign UAA and UAG to glutamine. Such reassignments differ from merely using synonymous codons at different frequencies. (ncbi.nlm.nih.gov)
Additional amino acids can also be incorporated through specialized decoding systems. Selenocysteine is usually inserted at selected UGA codons with the help of specialized tRNA, translation factors, and RNA signals. Pyrrolysine can be inserted at UAG codons in certain microorganisms. These systems extend the ordinary twenty-amino-acid repertoire without introducing a new basic nucleotide alphabet. (pmc.ncbi.nlm.nih.gov)
Mutation and sequence interpretation
The code determines how a coding-sequence mutation affects the amino-acid sequence. A synonymous substitution preserves the encoded amino acid; a missense substitution changes it; a nonsense substitution creates a premature termination codon. Insertions or deletions whose lengths are not multiples of three can shift the reading frame, altering downstream codons. (ncbi.nlm.nih.gov)
Synonymous changes are not necessarily functionally neutral: they can influence RNA processing or translation. Conversely, an amino-acid substitution does not by itself establish its effect on protein function. When interpreting results from DNA sequencing, researchers must distinguish the codon assignment from the biological consequences of a sequence change and select the appropriate translation table for the organism and cellular compartment. (genome.gov)