24 Regulation of Gene Expression I: Eukaryotic Gene Regulation

Learning Objectives

  • Explain what it means for a gene to be “expressed”
  • Use a model to explain the role of DNA sequences, DNA-binding proteins (transcription factors), and their interactions in controlling gene expression.
  • Explain the role of epigenetic modifications in the regulation of gene expression.
  • Use a model to explain how differences in gene expression account for the diverse cell types present in multicellular organisms.

In your body, how many different types of cells can you name? Maybe you thought of liver, skin, muscle, neuron, or other types of cells? In fact, there are over 200 different cell types in the human body, each with distinct features and functions that are related to the unique set of proteins found in different cell types. Yet all cells in your body came from one cell, a fertilized egg, that divided trillions of times by mitosis. These cells all have the same genetic information (DNA sequence), so what accounts for their differences? The answer is gene regulation: different types of cells express (transcribe and translate) a different subset of specific genes to produce the combination of proteins that define the features of that cell type.

Beyond controlling which genes are expressed in different cell types, gene regulation also includes dynamic expression events: within a cell, some genes are “turned on” and “turned off” in response to signaling. Think back to the previous unit, when we learned about cyclins that are expressed by actively dividing cells to promote cell cycle events. Cells also respond to DNA damage, low oxygen, and changes in nutrient availability with changes in gene expression. Thus, the characteristics and functions of a cell are controlled by the information in its DNA in combination with input from the environment.

Some distinctive phenotypes are caused by mutations that do not alter the protein produced in gene expression but instead affect when, where, or how much the gene is expressed. A familiar phenotype affected by a gene regulation mutation is blue eye color. Blue-eyed individuals have the typical version of a protein for producing melanin, which are responsible for pigmentation. However, a common mutation in individuals with blue eye color reduces expression of this gene specifically in eyes, leading to reduced eye pigmentation (Figure 24.1a). Another example is polydactyly, which is the growth of extra digits in development. The number of digits and their pattern is determined by the Sonic hedgehog protein (yes, named after the video game character). A mutation located 1 million base pairs away from the Sonic hedgehog gene causes expression of the protein in an area of the developing limb that normally doesn’t express it, causing the extra digits to form (Figure 24.1b).

Eye color and the number of digits is affected by mutations to gene regulation processes.
Figure 24.1. Two phenotypes affected by mutations affecting gene regulation. a, The range of human eye color phenotypes is due to several genes, but a single common variant in a gene regulatory region commonly leads to blue eyes. b, A cat displaying polydactyly, which is caused by a single mutation in a distant gene regulatory region. Eye color images by By Silar, licensed under the Creative Commons Attribution-Share Alike 4.0 International license. Cat images by B, flickr.com, licensed under CC BY-NC-SA 2.0.

Here, we will explore many ways in which cells regulate the gene expression process to fine-tune the location, amount, and function of the molecules produced in gene expression in eukaryotic cells (gene regulation in bacterial cells will be discussed in subsequent chapters). Figure 24.2 is an overview of these levels, and we take each of these in turn in the subsequent sections. We will revisit these ideas in the final chapter, using our knowledge of gene expression and other biological processes we have learned about to help us predict the effects of different mutations on phenotypes.

In eukaryotic cells, many different aspects of gene expression can be modulated to optimize the amount of active molecules produced.
Figure 24.2. Gene regulation occurs at multiple levels, ranging from DNA structure through protein modifications and stability. Created using BioRender.

Chapter Outline

Section 24.1 Epigenetic Modifications and DNA Structure

Section 24.2 Regulation of Transcription

Section 24.3 RNA Processing, Translational Regulation, and Post-Translational Regulation

Section 24.4 Gene Model for the In-Class Activity

Section 24.1 Epigenetic Modifications and DNA Structure

The first level of gene regulation occurs at the level of the DNA structure itself. DNA in the cell is associated with many proteins that organize and protect genetic information in ways that either relax or compact the DNA structure (Figure 24.3).

DNA structure varies from naked and not compacted to very compacted due to histones and scaffolding proteins.
Figure 24.3. DNA in cells is found in varying levels of compaction and condensation. The approximate diameters of DNA structures are reported in nanometers (nm). Modified from OpenStax Biology 2e, licensed under a Creative Commons Attribution 4.0 International (CC BY) license.

Chromosomes in mitosis are the most compacted versions of cellular DNA, and this compaction reduces the chance of breakage while dividing the DNA among daughter cells. During interphase, DNA is more relaxed but still highly organized, wrapped around complexes of histone proteins that interact with other proteins to control compaction (Figure 24.4a). Since proteins such as transcription factors and RNA polymerase need to interact with DNA bases for gene expression, a relaxed configuration is more permissive to transcription, while compacted DNA is inhibitory to transcription (Figure 24.5a).

Histone modifications that open the chromatin and unmethylated promoter sequences are associated with more transcription.
Figure 24.4. Epigenetic modifications affect transcription. a, DNA is wrapped around complexes of histone proteins (blue shapes), and histone modifications can influence DNA organization. Modifications that lead to an open DNA structure (light purple groups) are permissive for transcription factor and RNA polymerase binding and transcription, while modifications that lead to a closed DNA structure (dark red groups) reduce accessibility to proteins and transcription. b, Methylation of cytosine bases in promoter regions reduces transcription of the regulated gene. Created using BioRender.

The structure of DNA is dynamically regulated in ways that do not change the DNA sequence (the genetic information). Processes that alter the structure of DNA in ways that may influence gene expression are called epigenetic modifications. For example, the histone proteins can be chemically modified by methylation, acetylation, or phosphorylation (see the section on protein modifications of Figure 24.2 for examples), which in turn affects the association of histones with DNA and with other proteins that affect DNA structure. Certain histone modifications are associated with closed chromatin that is not accessible by transcription, while a different modification opens the chromatin to allow proteins associated with transcription to bind (Figure 24.4a). Another epigenetic modification is to the DNA itself: cytosine bases can be modified by the addition of a methyl group, creating 5-methylcytosine (Figure 24.4b). While this doesn’t affect base pairing (and therefore doesn’t affect the genetic information), methylation can lead to further chromatin restructuring that inhibits transcription, particularly methylation that occurs in gene promoters.

Within an organism, different types of cells have different patterns of epigenetic marks that emerge in development and contribute to unique gene expression patterns across cell types. Interestingly, even though epigenetic changes do not alter the DNA sequence, they are heritable—when a cell divides, the resulting daughter cells have the same epigenetic features. And while epigenetic changes are generally stable over time, they can be reversed—the histone modifications can be changed and the methyl groups on 5-methylcytosines can be removed. Epigenetic modifications are also implicated in disease. In several types of cancer, the promoter sequences of tumor suppressors show greater methylation than the same sequences in non-cancerous cells. Based on what you know about tumor suppressors, why do you think this change is associated with cancer?

Section 24.2 Regulation of Transcription

Revisit the gene expression simulation presented in Chapter 4. For transcription to occur, you first needed to move a positive transcription factor to the regulatory region, followed by RNA polymerase. This is illustrative of a general theme in eukaryotic transcription—the transcription factor binding to the DNA aids the function of RNA polymerase. Transcription factors therefore need to have the shape and chemical properties to allow for binding to DNA. Some transcription factors recognize specific DNA sequences, so they have a shape that allows for many interactions with a particular sequence of DNA. For example, transcription factors called GATA factors bind to the sequence 5′-GATA-3′. To recognize a specific sequence, the structure of a transcription factor allows the protein to nestle into the groove of the DNA helix, creating many opportunities for noncovalent interactions with the bases (Figure 24.5a). One such interaction between an asparagine amino acid in the transcription factor protein and the adenine base in the DNA is shown in Figure 24.5b.

Transcription factor proteins can interact with specific DNA sequences and increase or decrease transcription.
Figure 24.5. Transcription factors bind to DNA sequences to influence transcription. a, Model of a transcription factor protein interacting with bases in the DNA double helix. b, Zoomed in view of a single amino acid’s side chain (asparagine) on a transcription factor protein interacting with an adenine base in DNA Red dashed lines indicate hydrogen bond interactions between asparagine and adenine. c, Several DNA sequences are bound by different types of proteins to regulate gene expression. Created using BioRender.

The process of transcription requires many transcription factor proteins interacting with specific DNA sequences (Figure 24.5c). Adjacent to the DNA sequences to be transcribed, a core promoter sequence attracts general transcription factors, which then recruit RNA polymerase. As the name implies, general transcription factors are present in many cell types and aid in transcription initiation across for many genes. In contrast, some transcription factor proteins are only present in certain cell types and influence transcription of a subset of genes by binding to specific sequences that can be far away from the sequence to be transcribed. These regulatory transcription factors can act in ways that positively or negatively influence transcription, and are termed positive and negative transcription factors, respectively. The DNA sequences that are recognized by positive and negative transcription factors are called enhancers and silencers, respectively.

A complex of many proteins, interacting with both promoter and enhancer sequences, is required for transcription initiation.
Figure 24.6. A model of transcription initiation. General transcription factors and RNA polymerase assemble at the core promoter sequence, while a positive transcription factor binds to an enhancer sequence elsewhere on the chromosome. The Mediator protein complex bridges the protein interactions and stabilizes the looping of the DNA. Created using BioRender.

While the details of transcription initiation can vary significantly from gene to gene, Figure 24.6 shows one model of transcription initiation. Here, many different general transcription factors form a complex with RNA polymerase at the core promoter sequence, while a positive transcription factor binds to an enhancer sequence located farther away. An intermediary protein complex called Mediator bridges these proteins, which are brought together by the folding over or looping of the DNA in between the promoter and the enhancer. The interaction of the positive transcription factor with general transcription factors and RNA polymerase is the cue for transcription to begin.

This suggests at least two ways in which cells can regulate which genes are expressed in which cell type. One is by activating an existing transcription factor through phosphorylation. Recall from unit 3 that cells expressed Cyclin D in response to a growth factor signal, caused by signaling by the receptor that led to phosphorylation of positive transcription factor proteins already present in the cell. Phosphorylation changes the shape of the protein so that the transcription factor can bind to DNA and recruit RNA polymerase. However, another approach is to regulate the transcription of the positive transcription factor itself. If Gene A is needed for the function of neurons but not other cell types, we may observe that only neurons express the positive transcription factor to activate the expression of Gene A. Even though the DNA sequence of Gene A is present in all cells of a given organism, the protein encoded by Gene A would only be present in neurons due to the positive transcription factor also only present in neurons.

In contrast, negative transcription factors bind to silencer sequences and have a negative impact on transcription. Once again, the specifics vary from gene to gene, but one model of transcriptional repression shows that negative transcription factor binding disrupts the assembly of proteins required for transcription initiation (Figure 24.7). Negative transcription factors may also be expressed in specific cell types and not others, helping to create specific patterns of gene expression in different cell types. Overall, though, there are far fewer studies of silencers as compared to enhancers, despite the important role of silencers in the regulation of gene expression.

Negative transcription factors prevent transcription by interfering with assembly of proteins needed for transcription.
Figure 24.7. A model of transcription repression. A negative transcription factor interacting with a silencer sequence can interfere with transcription initiation complex assembly and other protein binding needed for transcription to occur. Created using BioRender.

Section 24.3 RNA Processing, Translational Regulation, and Post-Translational Regulation

Once transcription occurs, there are still ample opportunities for the cell to modulate the specific protein that is produced. Initially, transcription of a gene produces a primary transcript that is complementary to (base pairs with) the template DNA sequence (represented schematically in Figure 24.8). However, this primary transcript is altered by three processing steps before leaving the nucleus to be translated: capping, splicing, and polyadenylation.

After transcription, the primary transcript is modified to create the mature mRNA to be translated.
Figure 24.8. Transcription and RNA processing create the mRNA. The template DNA strand is transcribed to make a single-stranded primary transcript that includes both exons and introns. RNA processing of the primary transcript add a 5′ cap, removes introns and joins exons via splicing, and adds a poly(A) tail. RNA processing steps to create the mRNA occur in the nucleus.

The first step, capping, occurs shortly after the 5’ end of the RNA emerges from RNA polymerase, during the process of transcription. Next, splicing removes some sequences from the RNA, called introns, and connects the remaining sequences together. The sequences in the primary transcript that remain in the processed RNA are called exons, since they contain sequences that will be expressed. The removed introns will remain in the nucleus, while the exons will eventually exit the nucleus for translation. The final modification is the addition of ~200 adenine nucleotides to the 3’ end of the processed RNA, the poly(A) tail. Once these three processes have occurred, the resulting RNA molecule is termed an mRNA and can leave the nucleus for translation.

As you can see in Figure 24.8, the processed mRNA is shorter in length as compared to the primary transcript. Because there is an energy cost to synthesizing RNA, it may seem wasteful to create sequences that do not contribute to the molecule that will be translated. Given that, introns are surprisingly commonly in multicellular organisms, and most human genes contain long introns that need to be spliced out during RNA processing. This suggests the ability to splice genes presents an advantage that outweighs the energy cost of transcribing long introns. Indeed, while two cells may express the same gene, the proteins produced can be different, due to different processed mRNAs being produced (Figure 24.9).

Primary transcripts from some genes can be spliced in different ways to create different mRNA and proteins.
Figure 24.9.Alternative splicing creates different mRNAs and proteins from the same gene. Transcription of a gene produces a single type of primary transcript that can be spliced differently in different cell types, leading to different gene expression products. Created using BioRender.

Cell type A splices the primary transcript to include exons 1, 2, and 4, leaving out exon 3, while cell type B splices the transcript to include exons 1, 3, and 4, leaving out exon 2. Because the sequence of the mRNA is different, the proteins produced in translation of each mRNA may also be different. Thus, this process of alternative splicing allows cells to produce multiple types of proteins from the same gene, greatly increasing the diversity of proteins and protein functions across the organism.

After RNA processing, the mRNA transits from the nucleus into the cytosol for translation. The translation machinery first assembles at the 5’ end of the mRNA, at the 5’ cap (Figure 24.10). Then the ribosome moves towards the 3’ end of the mRNA, looking for the first 5’-AUG-3’ sequence, which is the start codon for initiation of translation.

The mature mRNA has a 5' cap, 5' UTR, coding sequence, 3' UTR, and a poly(A) tail.
Figure 24.10. Model of mRNA. The first start codon (5′-AUG-3′) and the in-frame stop codon define the coding region of the mRNA that will be translated. Sequences before the start codon and after the stop codon are untranslated regions (5′ and 3′, respectively). Sequences outside the coding region contribute to translation efficiency and RNA stability. Created using BioRender.

As depicted in Figure 24.10, the start codon is typically not at the beginning of the mRNA, meaning there is some sequence in the mRNA between the 5’ cap and the start codon that is not translated, called the 5’ untranslated region (5’ UTR). After the start codon, the ribosome reads codons (groups of three nucleotides) and adds the corresponding amino acid acids until a stop codon is reached and translation ends (see the codon table in Figure 5.2 for the three stop codon sequences). The sequence between the start and stop codons is the coding sequence that will be translated to the amino acid sequence of the protein. As shown in Figure 24.10, the stop codon is not at the very end of the mRNA, meaning a 3’ UTR is present between the stop codon and the poly(A) tail.

While the UTR sequences do not contribute to the sequence of the protein, they play important regulatory roles that affect the amount of protein that is made. The sequence of the 5’ UTR influences how efficiently the ribosome translates the mRNA, while sequences in the 3’ UTR influence how long the mRNA is available to be translated. In the gene expression simulation mentioned earlier, one option in the Biomolecule Toolbox is the “mRNA destroyer.” Revisit the simulation to see how mRNA degradation by the mRNA destroyer affects the amount of a particular protein that you can make.

Once a protein is produced in translation, it may be modified in several ways that affect the protein’s function. We have seen many examples of the effects of phosphorylation on protein shape and function, but many other types of modifications are possible and influence protein function. Lastly, some proteins are more stable than others, and cells can dynamically regulate the amount of protein by controlling when proteins are degraded. This regulation is essential for processes that rely on specific protein amounts to control cellular functions. Remember that cyclins are made and degraded in conjunction with different cell cycle phases, and all cyclins are degraded by the end of M phase, forcing new cyclins to be produced to start a new cell cycle.

Section 24.4 Gene Model for the In-Class Activity

To help us explore the many levels of gene regulation and the impact of changes on the gene expression process, we will use the gene model shown in Figure 24.11. This model represents a gene, which is a double-stranded DNA sequence. Even though we don’t see a double helix or letters that would represent the bases, there are some clues that remind us that we are looking at DNA. To start, there are two lines to remind us of the double-stranded nature of DNA. We also see a promoter sequence that is labeled, which we know from previous chapters is a sequence found in DNA. Finally, we see the labels for the transcription start and the termination sites, and we know that transcription is a process that uses the information in the DNA to make an RNA.

A gene model shows the promoter, several exons and introns, and shaded regions to indicate coding sequences.
Figure 24.11. Model of gene showing key features related to gene expression. Shaded regions indicate coding sequences.

This model also has features that help us understand how the DNA sequences relate to the products of gene expression. Based on the transcription start site label, we can infer that sequences before that site won’t affect the sequence of the RNA made from expression of this gene, and apply similar logic to the transcription termination site. We also see some sequences labeled “exon” and some labeled “intron.” Since transcription begins at the transcription start site and ends at the termination site, all exons and introns will be present in the primary transcript. Then, the primary transcript will be processed (review Section 24.3 for these steps) and the introns will be spliced out as part of this processing. This means that only the exon sequences potentially affect the sequence of the mRNA and the protein. Although these are labeled, we further differentiate exons by drawing them as boxes, while the introns are depicted as lines, similar to the sequences outside the gene that also do not affect the specific sequence of the mRNA. Using Figure 24.11 and comparing with Figure 24.8, can you draw the sequence of the primary transcript and the processed mRNA?

The final feature to note in this model is the shading, which is used to indicate DNA sequences that correspond to sequences in the mRNA that code for the protein. Compare Figure 24.12 with Figure 24.11—if the shaded region is the coding sequence, what does that mean for the sequences at the start and end of the shaded region in the DNA or the mRNA? Can you label the mRNA with the start codon, stop codon (pick any one of the three stop codons), 5’ UTR, and 3’ UTR?

As with many other models, we will revisit the model in Figure 24.11 or a similar variation many times during this unit. Familiarize yourself with the model and the conventions used to communicate key information about gene expression so you can confidently make predictions about gene expression based on this model!

License

Icon for the Creative Commons Attribution-NonCommercial 4.0 International License

Cells and Molecules Copyright © by Katherine Krueger is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License, except where otherwise noted.