27 Alleles and Phenotypes
Learning Objectives
- Distinguish between gene and allele, and genotype and phenotype
- Explain how different gene sequences (alleles, arising from mutations) encode proteins with different phenotypes and functions
- Explain how the presence of certain alleles (genotype) affects cell phenotypes
- Learning objectives from prior class days on gene expression (Gene expression I & II, Regulation of gene expression I & II) and DNA replication
Throughout this text, we have considered many different phenotypes of cells—the molecules they contain and produce, how cells organize themselves, how cells obtain and use energy, how cells divide, what genes are expressed in cells. Anything cells do and are, anything that can observed about cells, is a phenotype. These phenotypes are influenced by the environment, but are also due to the specific DNA sequences in cells (genotypes). Therefore, the key to altering cell phenotypes is to alter cell genotypes.
The study of gene mutations and any phenotypes that result is crucial for the study of biology, as they help us understand how proteins function and how processes work. For example, if 100 individuals with a specific disease have the same mutation in the same gene, it’s a strong hypothesis that the protein encoded that gene is linked to a process that is altered to cause the disease.
In this chapter, we will combine all the information from previous chapters to predict how changes to DNA sequences will affect cellular phenotypes. In so doing, we will review the major cellular processes of gene expression, gene regulation, and DNA replication, and identify mutations that might affect each of these. Finally, to tie everything together, we will focus on a specific organismal phenotype that you have likely heard of—blood types.
Chapter Overview
Section 27.1 Genes and Alleles
Section 27.2 Mutation Locations and Phenotypes
Section 27.3 Mutation Case Study
Section 27.4 Blood Type Alleles and Phenotypes
Section 27.1 Genes and Alleles
As introduced in Chapter 18, the DNA of eukaryotes is organized into linear chromosomes, and these are often found in pairs called homologous chromosomes. In Figure 27.1, we see two chromosomes of the same length, with the centromere in the same location on each, and with the same two genes labeled, Gene S and Gene P.

These are common conventions for depicting homologous chromosomes in models, but remember that homologous chromosomes should have about 95% similarity in overall sequence. That leaves some room for sequence variation, which is indicated by the slightly different patterns in the gene labeling between chromosomes. The models we will build in class will show the sequence differences, but if you see different shading or patterns as shown in Figure 27.1, you can assume these are indicating different alleles, or versions, of the same gene.
The right half of Figure 27.1 recaps another important term from Chapter 18: sister chromatids. DNA replication should result in a complete copy of the DNA sequence. If there was one homologous pair before replication, there is still one homologous pair after replication, only now each chromosome has two identical sister chromatids (which should really be called “twin chromatids”). In the model, this is indicated by the same shading on each chromatid for each of the two genes depicted. Each sister chromatid will become a chromosome in a daughter cell, thus resulting in daughter cells that are genetically identical to the original parent cell.
The different alleles depicted in Figure 27.1 indicate differences in gene sequences, or mutations. As discussed in Chapter 21, mutations are heritable changes in the genome that typically result from errors or DNA damage that was not correctly repaired before DNA replication. We have only recently developed the technology to intentionally change DNA sequences, so most of the genetic variation that is observed is spontaneous and random in origin. Importantly, because mutations change the DNA sequence, all mutations create new alleles. However, not all mutations create new phenotypes, as we will explore in the next section.
Section 27.2 Mutation Locations and Phenotypes
If you were asked to predict the effect of a mutation in one of your cells, which choice below best describes your prediction? How confident are you in your prediction?
- No impact on phenotypes
- A minor change to a phenotype
- A major change to a phenotype
- All options are possible, and more information is needed
Since you’ve been asked a similar question before, you may feel confident in picking the last option, and more information is usually helpful to have. Generally, though, most mutations will have no phenotypic effect for several reasons summarized in Figure 27.2.

First, most of your cells have two chromosomes of each type for chromosomes 1-22 (Figure 27.2a). Even if a mutation disrupts the function of one allele, the presence of another functional allele is often sufficient (genes that encode tumor suppressors are a notable exception). Another important fact is that to change protein function usually requires changing the shape of the protein by altering its amino acid sequence. Yet, only 2% of the human genome is sequence that encodes proteins, so a random mutation is less likely to be found in that subset of the genome (Figure 27.2b). Finally, even if a protein-coding sequence is altered, there are many mutations that would not alter the amino acid sequence due to the nature of the genetic code. If we focus on the small portion in Figure 27.2c, we can see that there are six different codons for Leu alone.
However, for the subset of mutations that appear in protein-coding sequences or other sequences of known importance, we should be able to predict whether phenotypes are likely affected and why. Importantly, even a small change can be the cause of disease. Figure 27.3 illustrates a well-known example that we examined in Chapter 6, the mutation in the beta-globin gene linked to sickle-cell disease.

Here, a single change in the DNA sequence of the A allele creates the S allele. The change to the DNA sequence alters the mRNA codon from 5’-GAG-3’ to 5’-GUG-3’, which in turn changes the amino acid added in translation from Glu to Val. You can see from the amino acid side chains that Glu and Val have very different chemical properties and therefore form different interactions with other amino acids and the surrounding environment when assembled into the hemoglobin protein. Review Chapters 2, 6, and 7 to further strengthen the explanation. As a result of this change, the hemoglobin protein structure is sufficiently altered in individuals with only the S allele such that their red blood cells, which are full of hemoglobin protein and not much else, have a distinct sickle-shape. This altered cellular shape adversely affects red blood cell function, leading to the painful symptoms of the disease.
To fully uncover the specific role that mutation plays in sickle cell disease requires some specialized knowledge, but we can make predictions based on the type of amino acid change that occurred. A general guideline is that mutations that change an amino acid to one with significantly different properties (e.g., from a nonpolar amino acid to a polar one or vice versa) are more likely to create phenotypes than a mutation that changes the amino acid to one with similar properties as before. This is because the former change is more likely to alter the structure of the protein.
We can also make predictions about phenotypes based on what we know about the gene expression process. As summarized in Figure 27.4, the process of transcription requires promoter sequences in DNA, where RNA polymerase and general transcription factors bind to initiate transcription that produces the RNA.

Although not shown in Figure 27.4, enhancer sequences are also part of the transcription initiation process for many genes; see Figure 24.7 for a reminder of their role. Mutations that alter these sequences could affect the gene expression process, changing how much transcription takes place, and therefore the number of mRNA copies that are detected in cells, which is a phenotype.
Additionally, the start and stop codons in the RNA are required for the ribosome to start and end translation that produces the protein, while the sequence between the start and stop determines the specific amino acid sequence of the protein. Changes to any coding sequence have the potential to alter the protein produced in translation, changing the amino acid sequence, the number of amino acids in the sequence, or both. We previously discussed specific types of mutations to coding sequences in Chapter 21—what length or specific protein sequence phenotypes would you predict that synonymous, missense, nonsense, or frameshift mutations would have?
Section 27.3 Mutation Case Study
A portion of the in-class activity will be focused on the gene model introduced in Chapter 24, shown again in Figure 27.5. This model shows key features in the DNA that relate to gene expression, along with three separate mutations into the sequence.

The sequences written out represent the DNA with each mutation (the original sequence is not shown). Assume that the bottom strand in each sequence is the template DNA strand. Based on your knowledge of this model and the gene expression process, which mutation, if any, would alter the sequence of the mRNA and or the sequence of the protein produced in gene expression?
First, remember that transcription starts at the start site indicated in the model, and all sequences until the termination sequence are transcribed by RNA polymerase. All mutations shown are between the transcription start site and the termination sequence, so all will be transcribed and affect the sequence of the primary transcript. The sequence produced in transcription is complementary to the template strand (3’-ATC-5’), so each mutation produces the sequence 5’-UAG-3’ when transcription occurs. However, recall that some sequences (introns) in the primary transcript are spliced out during RNA processing, and only the exon sequences remain in the mature mRNA. Since Mutation Y is in intron 3, this will not affect the mature mRNA in any way. Meanwhile, Mutations X and Z are both in exons and thus will change the sequence of the mature mRNA.
Additionally, both Mutations X and Z also have the potential to affect the expression of the protein, although in different ways. Mutation X is in the shaded region of an exon, while Mutation Z is in an unshaded exon. Remember that shaded regions of exons indicate coding sequences that will be translated, so only Mutation X can affect the amino acid sequence of the protein. If we consult the genetic code table in Figure 27.6, you will see that 5’-UAG-3’ is a stop codon.

If this codon is in the reading frame, this is a nonsense mutation that ends translation prematurely. This means that the protein will be shorter than usual due to Mutation X. Whether or not the function of the protein will be affected depends on the role of the part of the protein that is now missing, but it certainly seems plausible to hypothesize that the protein will have an altered function due to the change in structure.
In contrast, the same sequence change in the unshaded exon 4 is not a stop codon, since these sequences are not translated. The amino acid sequence of the protein would not be altered due to Mutation Z, but the mRNA stability or other aspects of translation might be altered by the sequence change in this untranslated region of the mRNA. Depending on the specific effect, the number of protein copies produced in translation might be higher or lower, though the sequence itself would be unchanged.
To summarize the gene expression phenotypes, all three mutations affect the sequence of the primary transcript, Mutations X and Z affect the mRNA sequence, and only mutation X affects the amino acid sequence and will shorten the protein produced in translation. Mutation Z may also alter the number of protein copies, which is another measurable phenotype. Importantly, while each mutation creates a stop codon sequence, only Mutation X, within the coding region to be translated, is a nonsense mutation. Mutations Y and Z are simply point mutations. Furthermore, an in-frame stop codon only stops translation. Transcription is a carried out by RNA polymerase, which does not read codons, so none of the mutations affect the length of the mRNA.
Finally, what is the effect of each mutation on DNA replication? Like RNA polymerase, DNA polymerase does not read codons, so a stop codon sequence anywhere in the genome does not stop DNA polymerase from replicating sequences. However, mutations are permanent changes to the DNA sequence, and DNA replication will accurately copy the mutated sequences with every cell cycle.
Section 27.4 Blood Type Alleles and Phenotypes
To practice more with predicting and explaining phenotypes (traits) based on genotypes (specific DNA sequences), let’s explore a phenotype we all have—blood type. Each drop of your blood contains hundreds of millions of red blood cells, and embedded in the plasma membranes of these cells are proteins with attached sugars, called glycoproteins (Figure 27.7a). These proteins are produced by translation at the endoplasmic reticulum, then travel through the Golgi Apparatus on their way to the plasma membrane, where sugars are covalently attached to them by the process of glycosylation (Figure 27.7b). In red blood cells, as proteins pass through the Golgi, an enzyme called glycosyltransferase (GT) adds specific sugars to the proteins, and a person’s blood type refers to those sugars.

There are three common alleles for the gene that encodes the GT enzyme, which we will refer to as blood type (BT) alleles: A, B, and O. Every person should have two BT alleles, inheriting one allele from each parent, which means there are six different combinations (AA, BB, AO, BO, AB, and OO). However, only four common blood types are observed (Type A, Type B, Type AB, and Type O). Note—if you have heard of blood types that are positive (+) or negative (-), this refers to Rh blood groups that are also very important but controlled by a different gene and thus will not be considered here.
The presence of different alleles and different blood types might raise some questions: how do these alleles differ, how are the different alleles related to different blood types, why are there fewer blood type phenotypes than there are blood type genotypes, to name a few. These are questions that we can use our knowledge of gene expression, mutations, and protein structure-function to answer in the in-class activity!
Let’s work through some examples to introduce the models for the in-class activity. Let’s assume a person has two B alleles for the BT gene. If we were to draw the genotype of this person, we would draw two homologous chromosomes and label the B allele of the BT gene on each (Figure 27.8a). There are other genes present on chromosome 9, but we will just show the BT gene to keep our model as simple as possible. When the BT gene is expressed, translation of the mRNA produced creates the B version of the glycosyltransferase enzyme, or GT B. This enzyme adds a specific type of sugar, which we will call the B sugar, to proteins in the Golgi (Figure 27.8b). The red blood cells of this person all have the B type sugar added (Figure 27.8c), thus this person has the blood type phenotype of Type B.

The sequences of the BT A and BT O alleles are very similar to the BT B sequence, yet are associated with different blood type phenotypes. The glycosyltransferase enzyme encoded by the A allele adds a different sugar to proteins, which we will call the A sugar (Figure 27.9).

Some people have one A and one B allele, and thus have both types of sugars on their red blood cells. Can you sketch models of the red blood cells of individuals with two A or one A and one B alleles? Which blood type phenotype does each have?
In contrast, the enzyme encoded by the BT O allele does not add any sugars. In class, we will examine specific differences in the DNA sequences and what happens in gene expression and protein folding to explain why the A and O enzymes function differently than the B enzyme. This will allow us to connect genotype to phenotype, using the knowledge collected throughout the semester.