Biology › Nucleic acids, genomes and protein synthesis › The genetic code, and getting a gene out of the nucleus
The genetic code, and getting a gene out of the nucleus
Four bases have to specify twenty amino acids, and they do it in threes. The rules that follow — 64 triplets for 20 amino acids, no overlaps, and almost the same table in every organism alive — explain a great deal about how mutations behave and why genetic engineering works at all.
Before this DNA and RNA structure · Complementary base pairing · The nuclear envelope and nuclear pores
Before you start
The genetic code is degenerate — more than one triplet codes for the same amino acid — and since that means a lot of duplication, it is a wasteful or inefficient system that evolution has not yet tidied up. Degeneracy is not waste. It is a safety margin. Because 61 of the 64 triplets code for amino acids and only 20 amino acids need coding, most amino acids have several triplets, and those triplets usually differ only in the third base. A copying error in that third position therefore very often produces exactly the same amino acid and exactly the same protein. Far from being untidy, the spare capacity absorbs a substantial fraction of all mutations before they can do anything at all.
What you should be able to do
- State the four properties of the genetic code and explain a consequence of each.
- Explain, with numbers, why the code has to be read in threes rather than ones or twos.
- Distinguish exons from introns, and a gene from the genome.
- Describe transcription in sequence, naming the enzyme, the strand used and the bonds involved.
- Explain what splicing does and why one gene can give rise to more than one polypeptide.
- Explain why many base substitutions change nothing about the protein.
Four properties, and what each one buys you
The genetic code is the relationship between triplets of bases and amino acids. Four statements describe it, and each has a consequence you can be asked about directly.
| Property | What it means | What follows |
|---|---|---|
| Triplet | Three bases specify one amino acid | 64 possible triplets for 20 amino acids, so there is capacity to spare |
| Degenerate | Most amino acids have more than one triplet | Many base substitutions produce no change in the protein |
| Non-overlapping | Each base belongs to one triplet only | An insertion or deletion shifts every triplet after it |
| Universal | Almost all organisms read the same triplets the same way | A human gene inserted into a bacterium makes the human protein |
Why three and not two
Show, using numbers, why a code based on pairs of bases could not specify twenty amino acids, and state how many triplets are spare once every amino acid has one.
There are four bases. A code of one base gives 4¹ = 4 combinations, which is nowhere near enough. A code of two bases gives 4² = 16 combinations, still four short of the twenty amino acids that need specifying.
A code of three bases gives 4³ = 64 combinations, which is comfortably enough. Three of the 64 are stop signals, leaving 61 that specify amino acids.
Sixty-one triplets sharing twenty amino acids means 41 more than the minimum, and that surplus is what degeneracy consists of. Leucine, serine and arginine have six triplets each; methionine and tryptophan have only one.
Universality is worth stating carefully. The code is very nearly universal rather than absolutely so — mitochondria and a few protoctists read a handful of triplets differently. Say 'universal' if a mark scheme asks for the property; say 'almost universal' if you are asked to be precise. Either way the consequence is the point: because a bacterium reads human triplets the same way a human cell does, human insulin can be manufactured in E. coli, and it has been since 1982.
- Codon
- A sequence of three bases in mRNA that specifies one amino acid or a stop signal.
- Degenerate
- Describing a code in which most amino acids are specified by more than one triplet.
- Non-overlapping
- Describing a code in which each base is part of one triplet only and triplets are read in sequence.
Genes, exons and introns
A gene is a sequence of DNA bases that codes for a polypeptide, or for a functional RNA molecule such as a tRNA. Its position on a chromosome is its locus. The genome is the complete set of DNA in a cell, and the proteome is the full range of proteins that cell can produce — a smaller and more changeable thing, since not every gene is expressed at once.
In eukaryotes a gene is not one continuous coding sequence. It is broken into exons, which are the parts that end up represented in the finished mRNA, separated by introns, which do not. Introns can be long: some human genes are more than 95% intron by length. Prokaryotic genes have no introns, which is one of several reasons a human gene cannot simply be pasted into a bacterium and expected to work.
- Gene
- A sequence of DNA bases coding for a polypeptide or a functional RNA.
- Exon
- A section of a gene that is represented in the mature mRNA.
- Intron
- A non-coding section of a gene, transcribed into pre-mRNA and then removed before the mRNA leaves the nucleus.
- Genome
- The complete set of DNA in a cell, including the DNA that does not code for polypeptides.
A large proportion of eukaryotic DNA is non-coding: introns, repeated sequences between genes, and regions involved in switching genes on and off. Calling all of it useless was a fashion that did not survive contact with the evidence, and questions increasingly expect you to know that non-coding does not mean functionless.
Transcription, step by step
Transcription is the copying of one gene into mRNA, and it happens in the nucleus because that is where the DNA is. The enzyme is RNA polymerase — not DNA polymerase, which does a different job in a different process, and confusing the two is one of the quickest ways to lose a whole answer.
In order, then. RNA polymerase binds to a promoter region at the start of the gene and unwinds a short section of the double helix, breaking the hydrogen bonds between the paired bases. Only one of the two exposed strands is used: the template strand. Free RNA nucleotides in the nucleoplasm align against it by complementary base pairing, with uracil pairing to adenine wherever thymine would have gone. RNA polymerase joins the aligned nucleotides by condensation, forming phosphodiester bonds, and moves along the gene as it works. Behind it, the two DNA strands re-form their hydrogen bonds and the helix rewinds. When the enzyme reaches a stop signal at the end of the gene it detaches, and the finished pre-mRNA is released.
The strand that is not read is the coding strand, and it earns its name because its sequence is the same as the mRNA's, with T wherever the mRNA has U. That is useful in exam questions: if you are given the coding strand, you can write the mRNA by changing every T to a U. If you are given the template strand, you have to take the complement.
Splicing: cutting out what is not needed
The pre-mRNA that RNA polymerase produces is a copy of the whole gene, introns included. Before it can be used it goes through splicing: structures called spliceosomes cut the introns out and join the exons together, producing mature mRNA. Only then does the molecule leave the nucleus through a nuclear pore and travel to a ribosome.
Splicing is not always done the same way on the same transcript. Under alternative splicing, different combinations of exons are retained, so one gene gives rise to several different mRNA molecules and therefore several different polypeptides. This is a large part of why humans manage a proteome of well over 100,000 proteins from around 20,000 protein-coding genes.
- Pre-mRNA
- The immediate product of transcription in a eukaryote, containing both exons and introns.
- Splicing
- The removal of introns from pre-mRNA and the joining of exons to produce mature mRNA.
Why degeneracy makes some mutations silent
Take the mRNA codon GAA, which specifies glutamic acid. Change its third base to G and you have GAG — which also specifies glutamic acid. The DNA has changed, the mRNA has changed, and the protein has not. A substitution of that kind is a silent mutation, and it is common precisely because the triplets sharing an amino acid tend to differ in the third base.
Compare that with a deletion. Remove one base and every triplet downstream is read from a new starting point, because the code is non-overlapping and read strictly in threes from a fixed start. A single deletion near the beginning of a gene can therefore change every amino acid after it and usually produces a stop codon early, giving a short and useless polypeptide. The contrast between a substitution that changes nothing and a deletion that changes everything is a standard exam comparison, and it comes straight out of two properties of the code.
TRY IT — Two mutations, two very different outcomes
A section of the coding strand of a gene reads TTA CGA GAA CCT. One sample has the fourth base changed from C to A. Another sample has the fourth base deleted altogether. Count along the sequence before you start. Explain why the second is likely to be far more damaging than the first, without needing a codon table.
Check your answer
The fourth base is the C at the start of the second triplet, so the substitution changes one triplet and nothing else. CGA becomes AGA, so at worst one amino acid in the finished polypeptide is different, and because the code is degenerate the new triplet may well specify the same amino acid anyway. Even if it does not, a single amino acid change in a region away from the active site or binding site often leaves the protein working normally.
The deletion is different in kind. The code is non-overlapping and read in threes from a fixed start, so removing one base pulls every subsequent base one place forward. Every triplet from that point on is read differently — a frame shift — and the amino acid sequence after the deletion bears no relation to the original.
A shifted reading frame also produces stop codons at random positions, so the polypeptide is usually cut short as well as scrambled. The protein almost never folds into a working shape. One base removed, an entire protein lost.
In the exam
- Say 'triplet', 'degenerate', 'non-overlapping' and 'universal' by name. Each is a mark, and vague paraphrases such as 'three letters at a time' often are not.
- The enzyme for transcription is RNA polymerase. Writing DNA polymerase here usually costs the mark for the whole step.
- Hydrogen bonds are broken to expose the template and phosphodiester bonds are formed along the new RNA. Both belong in a full description.
- mRNA carries U, never T. Check any sequence you write before moving on.
- If a question gives you the coding strand, swap T for U. If it gives you the template strand, take the complement first. Read which one you have been given.
- Silent mutations happen because the code is degenerate. Frame shifts happen because it is non-overlapping. Name the property, not just the effect.
Check yourself
A biotechnology company wants a human gene expressed in a bacterium. They isolate the gene directly from a human chromosome and insert it into E. coli, but no functional human protein is produced. Suggest why, and what they should have used instead.
Answer
A gene taken straight from a human chromosome contains introns. Bacteria do not carry out splicing, because their own genes have no introns and they have no spliceosomes.
The bacterium therefore transcribes the whole inserted sequence, introns included, and translates it as though every base were coding. The introns are read in the reading frame that happens to follow, producing wrong amino acids and almost certainly premature stop codons, so no functional human protein appears.
The fix is to start from mature mRNA rather than from the chromosome. Isolate the mRNA from a human cell that expresses the gene heavily, then use reverse transcriptase to make a DNA copy of it. Because the introns have already been spliced out of that mRNA, the resulting DNA is a continuous coding sequence the bacterium can handle.
Note what this does not challenge: the code is still universal, so the bacterium reads each triplet exactly as a human cell would. The problem was never the code — it was the introns.
Questions
Question 15 marks
Describe how a molecule of pre-mRNA is produced from a gene in the nucleus of a eukaryotic cell.
Mark scheme
- B1 RNA polymerase binds to a promoter region at the start of the gene
- B1 the enzyme unwinds a short section of the double helix, breaking the hydrogen bonds between the paired bases
- B1 only one of the two exposed strands, the template strand, is read
- B1 free RNA nucleotides align against the template by complementary base pairing, with uracil pairing to adenine wherever thymine would have gone
- B1 RNA polymerase joins the aligned nucleotides by condensation into phosphodiester bonds, moving along the gene until it reaches a stop signal and releases the pre-mRNA
Question 24 marks
Explain why a mature mRNA molecule is shorter than the gene it was transcribed from, and explain how one gene can give rise to more than one polypeptide.
Mark scheme
- B1 the pre-mRNA is a copy of the whole gene, introns included
- B1 spliceosomes cut the introns out and join the exons together, so the introns are not represented in the mature mRNA
- B1 in alternative splicing, different combinations of exons are retained in the mature mRNA
- B1 so several different mRNA molecules, and therefore several different polypeptides, can come from the same gene
Question 34 marks
A single base is inserted into the coding sequence of a gene close to its start. The polypeptide produced is much shorter than normal and its amino acid sequence bears no relation to the original beyond the point of insertion. Suggest why, referring to two properties of the genetic code.
Mark scheme
- B1 the code is read in threes from a fixed starting point and is non-overlapping, so each base belongs to one triplet only
- B1 inserting a base pushes every subsequent base one place along, so every triplet after the insertion is read differently, which is a frame shift
- B1 the amino acid sequence after that point is therefore completely altered rather than altered in one place
- B1 a shifted reading frame produces stop codons at positions where there were none, so translation ends early and the polypeptide is short
Question 43 marks
The coding strand of a short section of a gene reads GAT CCA TTG. Give the base sequence of the template strand, and give the base sequence of the mRNA transcribed from this section.
Mark scheme
- B1 template strand CTA GGT AAC, each base complementary to the coding strand
- A1 mRNA GAU CCA UUG
- B1 the mRNA matches the coding strand, with uracil in place of every thymine
Question 52 marks
State what is meant by saying that the genetic code is degenerate, and state what is meant by saying that it is non-overlapping.
Mark scheme
- B1 degenerate: most amino acids are specified by more than one triplet
- B1 non-overlapping: each base belongs to one triplet only, and the triplets are read one after another
Worth remembering
- Triplet, degenerate, non-overlapping, universal — and a consequence for each.
- 4³ = 64 triplets, 61 coding and 3 stop, for 20 amino acids.
- RNA polymerase transcribes; only the template strand is read.
- The mRNA matches the coding strand with U in place of T.
- Splicing removes introns and joins exons; alternative splicing lets one gene give several polypeptides.