20 DNA Replication

Learning Objectives

  • Describe the extracellular and intracellular requirements for DNA replication and identify when DNA replication occurs in the cell cycle
  • Construct models to demonstrate the process and outcome of DNA replication
  • Use a model of DNA replication to explain why continuous synthesis of both DNA strands is not possible

Now that we’ve explored key events in the cell cycle and their regulation, we can investigate more thoroughly one essential step in the process: DNA replication. Remember that cells use DNA to encode proteins that are needed to carry out many functions in the cell. When cells divide, an accurate copy of the cell’s DNA needs to be provided to each daughter cell so that each cell has the required genetic information. Thus, prior to cell division, all of the DNA in the parent cell’s genome must be accurately replicated (Figure 20.1). This chapter will focus on mechanics and the process of DNA replication. In the subsequent chapter on DNA damage and mutations, we will discuss how the cells increase the accuracy of DNA replication and thus maintain the integrity of genetic information.

The order of cell cycle phases is G1, S, G2, and M. DNA replication (DNA synthesis) occurs in S phase.
Figure 20.1. Overview of key cell cycle events. DNA replication occurs during S phase of the cell cycle. Created using BioRender.

Chapter Outline

Section 20.1 Comparison of Transcription to DNA Replication

Section 20.2 Semiconservative DNA Replication

Section 20.3 Initiation of DNA Replication

Section 20.4 Elongation and Termination in DNA Replication

Section 20.5 The Telomere Problem

Section 20.1 Comparison of Transcription to DNA Replication

Before diving into the details of DNA replication, you might find it helpful to review the process of transcription, from Chapter 4. Luckily, several key “rules” from transcription are relevant for DNA replication as well: both processes use base pairing to a DNA template to determine the sequence of the new strand, which is produced in the 5’ to 3’ orientation while reading the DNA template in the 3’ to 5’ direction (Figure 20.2). Both processes also use polymerases to synthesize new nucleic acid strands by joining nucleotides. In each case, covalent phosphodiester linkages form between the 3’ hydroxyl group on the nucleotide’s sugar and the 5’ phosphate group on the incoming nucleotide triphosphate (Figure 20.2b).

Transcription by RNA polymerase and DNA replication by DNA polymerase have many similarities.
Figure 20.2. Comparison of transcription and DNA replication. a, Model of transcription, in which RNA polymerase uses the DNA template strand to create a strand of RNA. b, Formation of a phosphodiester linkage by DNA polymerase (yellow square), in which the 3′ hydroxyl on an existing DNA strand (black color) reacts with the phosphate on the incoming triphosphate deoxyribonucleotide (blue color), creating an extended DNA strand and diphosphate as products. The template strand is not shown. Compare with Figure 4.7. c, Model of replication, in which one DNA polymerase on each strand works to replicate both strands simultaneously. Transcription model modified from OpenStax Biology 2e, licensed under a Creative Commons Attribution 4.0 International (CC BY) license. Phosphodiester linkage formation by Michał Sobkowski, licensed under the Creative Commons Attribution 3.0 Unported license. DNA replication model modified from an image by Christinelmiller, licensed under a Creative Commons Attribution-Share Alike 4.0 International license.

However, because transcription creates RNA and DNA replication creates DNA, different types of polymerases and nucleotides are used for each process. Also, in transcription, only one DNA strand is used as a template when a gene is transcribed, and only a small subset of the genome is read in transcription (Figure 20.2a). In contrast, DNA replication copies the entirety of both DNA strands (Figure 20.2c), which is a significant undertaking–consider that a typical human cell has 6 billion base pairs of DNA! Additionally, while cells are constantly expressing genes (and thus frequently carrying out transcription), the timing of DNA replication is more restricted, occurring mainly during S phase of the cell cycle in eukaryotic cells (Figure 20.1). The production of new DNA, therefore, is another way to distinguish actively dividing cells from cells in G0. Finally, we shall see in Section 20.3 that the enzyme that carries out most DNA replication, DNA polymerase, has unique requirements that further differentiate DNA replication from transcription carried out by RNA polymerase.

Section 20.2 Semiconservative DNA Replication

Let’s take a moment to review the structure of DNA (Figure 20.3). DNA is comprised of deoxyribonucleotides containing a phosphate, sugar, and nitrogenous base. In a cell, DNA is typically double-stranded, and each strand contains nucleotide monomers joined by phosphodiester linkages. When two strands base pair, they are antiparallel, meaning the 3’ end of one strand pairs with the 5’ end of the other strand.

The DNA double helix consists of two antiparallel DNA strands that interact with each other via base pairing.
Figure 20.3. A model of the DNA double helix. Modified from Madeleine Price Ball, licensed under CC0.

Between strands, specific bases form hydrogen bond interactions with other specific bases, A with T and G with C. This means that if you know the sequence of one DNA strand, you can predict the sequence of the other strand in the helix. James Watson and Francis Crick, upon discovering the structure of DNA, observed that this structural feature suggested how genetic information could be copied and passed down from one generation to the next: if the two strands of the double helix were unwound, each strand could serve as the template for producing new strands. Thus, in each double-helix following replication, one strand would be the original, or parent strand, and the other would be a newly produced, or daughter, strand (Figure 20.4). This method of replication is called semiconservative, since half of the DNA is retained, or conserved, from the prior generation.

In semiconservative replication, each strand of the parent DNA is a template for new DNA synthesis.
Figure 20.4. Model of semiconservative DNA replication. Blue strands indicate parental DNA, while newly synthesized DNA is shown with green strands. Due to the antiparallel nature of DNA strands, replication of one parental strand occurs in the opposite direction as replication of the other parental strand.  Modified from Madeleine Price Ball, licensed under CC0.

To test this prediction of the DNA replication mechanism, scientists conducted experiments with two different isotopes of nitrogen, which varied in the number of neutrons. Remember that nitrogenous bases contain several atoms of nitrogen (Figure 20.5a). E. coli bacteria were grown for several generations in media with 15N, which has a higher mass than typical 14N isotope, so that all nitrogen-containing biological molecules would contain the “heavy” 15N. When the DNA from these bacteria is isolated and separated by density, the DNA appeared to be uniformly dense (Figure 20.5b, top figure). These bacteria were then switched to media only containing the lighter 14N isotope, so that all new biological molecules made, including new DNA, would contain 14N. If the semiconservative molecule of DNA was correct, each double helix would contain one heavy and one light strand, with an overall lower density than DNA with only the heavy 15N isotope. Indeed, this is what was observed (Figure 20.5b, middle figure): based on the density, the double helix from this generation was half old (parental DNA with 15N) and half new (daughter DNA with 14N).

Experiments with nitrogen isotopes support the semiconservative model of DNA replication.
Figure 20.5. Using nitrogen isotopes to test the semiconservative model of replication. a, Each of the four nitrogenous bases in DNA contains several nitrogen atoms. b, Summary of experiments to test the semiconservative model using E. coli. Bacterial were grown in media containing only heavy 15N (light blue color) or only light 14N (red color) nitrogen isotopes. The density of DNA after centrifugation indicates whether DNA has only 15N, only 14N, or is mixture of the two isotopes. Nitrogenous base structures from OpenStax Biology 2e, licensed under a Creative Commons Attribution 4.0 International (CC BY) license. Experimental schematic from OpenStax Microbiology, licensed under a Creative Commons Attribution 4.0 International license.

To further test the semiconservative model, bacteria from generation 1 were grown for one additional generation in 14N media, once again ensuring that all new DNA contained the lighter isotope. When the DNA from this second generation was analyzed, two distinct samples were observed: one sample with the same density as observed in the first generation, and a second sample of lower density. When the bacteria from the first generation replicated their DNA, each strand served as the template for new DNA synthesis. When the strand with heavy 15N was used as the template, a new strand containing 14N was produced, resulting in a helix of 50% heavy and 50% light nitrogen isotopes. However, when the parental strand of 14N was used as the template, a helix of entirely 14N was produced. This result was further support of the semiconservative model of DNA replication.

You are encouraged to draw this out for yourself, using lines to represent DNA strands and different colors to represent the different nitrogen isotopes. You might even extend this experiment to a third generation, and consider whether different DNA products would be observed, and the ratios of each product. The key is to remember that each DNA strand in one generation is a template for further DNA replication.

Section 20.3 Initiation of DNA Replication

Before we delve into the mechanics of replication, let’s keep in mind the big picture. If a parent cell has the sequence shown at the top of Figure 20.6, cell division should produce two daughter cells with DNA sequences that are identical to each other and to the parent cell. Moreover, each double helix in each daughter cell should contain one strand from the parent cell (shown as dashed black lines) and one newly synthesized strand (shown as solid blue lines), since DNA replication is semiconservative.

Replication of parent DNA results in G2 phase chromatids that are half parent, half new DNA, which are then divided among daughter cells.
Figure 20.6. Overview of semiconservative DNA replication. The parent cell template DNA is indicated with black font and a dashed line for the DNA backbone. Newly synthesized DNA is shown with blue font and a solid line for the DNA backbone. After S phase, the amount of DNA is doubled so that M phase produces two daughter cells with the same DNA sequence and amount as each other and as the parent cell before S phase. Dotted black lines indicate the phosphate-sugar backbone of the parent strands, while solid lines indicate the backbones of newly synthesized strands.

In the parent cell, after S phase and completion of DNA replication, two double helices would be observed–one for each daughter cell. Each double helix is produced by separating the DNA strands in the parent cell, creating a single-stranded template that can be used in replication. As we add more details about the steps in DNA replication, return to this big picture and note how each step and each enzyme helps to achieve the overall goals in the process.

Just like transcription has a promoter sequence and translation has a start codon to signal where these processes start, DNA replication also has a sequence that indicates where new synthesis will begin, called the origin. This is where the replication machinery, which includes many enzymes, will assemble. There are two sets of these enzymes at each origin that will move outward from the origin, replicating as they go. Unlike transcription and translation, there is no sequence to indicate where replication ends–because the entire sequence needs to be replicated, synthesis simply continues until there is no more template that can be replicated.

Imagine we start with the parent cell DNA as shown at the top of Figure 20.6. In the first step of replication, helicase enzymes at the origin start breaking the hydrogen bonds between strands, moving outward from the origin in both directions to separate the double helix. The area that is now single-stranded is called the replication bubble, and only this single-stranded sequence is a template for replication (Figure 20.7). Outside of the bubble, DNA is double-stranded and DNA replication cannot occur until further unwinding occurs. Enzymes at the replication fork, which is the junction between single- and double-stranded DNA, will further unwind the DNA during the replication process.

Helicase separates the double helix to create a replication bubble. Replication complexes then move outward from the origin.
Figure 20.7. Creation of the replication bubble and movement of replication complexes. a, Helicase enzymes separate the two strands of the double helix, producing a replication bubble of single-stranded DNA that can be replicated. b, Two replication complexes (symbolized by the blue and green rectangles) move outward from the origin. Helicase enzymes (not shown here, see Figure 20.2c) at each replication fork further unwind DNA to expand the bubble. Each complex also contains two DNA polymerase enzymes (pac-man shapes), one to replicate the top strand and one to replicate the bottom strand. Dotted black lines indicate the phosphate-sugar backbone of the parent strands.

Replication will begin at the first bases exposed by helicase (shown in bold in Figure 20.7a). Both strands are replicated simultaneously (Figure 20.7b), but let’s focus on the bottom template strand to start. Notice that this strand has its 3’ end on the left and 5’ end on the right. Nucleic acid synthesis always occurs in the same direction, 5’ to 3’, reading the template strand 3’ to 5’. Therefore, the replication machinery will move from left to right along the bottom template strand, starting at the origin (bolded C), and will synthesize the new strand antiparallel to the template, adding new nucleotides to the 3’ end of the strand by creating phosphodiester linkages. So far, this is very similar to the process of transcription. However, the main enzyme for DNA replication, DNA polymerase, is unable to start DNA synthesis under these conditions. Unlike RNA polymerase, DNA polymerase requires a double-stranded 3’ end to build off of, and thus cannot start new strands (there are evolutionary reasons for this limitation that we won’t get into here). Instead, another enzyme, called primase, must first create a short primer sequence for DNA polymerase. Primase is a type of RNA polymerase, so the primer is built of ribonucleotides. As shown in Figure 20.8, with the template sequence of 3’-CTATA-5’, primase creates the primer sequence 5’-GAUAU-3’ (red sequence), working outwards from the origin towards the right replication fork.

At the origin, the enzyme primase creates an RNA primer complementary to the template DNA.
Figure 20.8. Creation of an RNA primer. The enzyme primase uses base pairing with the DNA template to produce an RNA primer. The first nucleotide is placed at the origin (bolded C on the bottom strand). As shown in the inset, new nucleic acid synthesis (solid line) occurs in the 5′ to 3′ direction to be antiparallel to the 3′ to 5′ template, so the overall direction of replication is from left to right along the bottom template strand, towards the right replication fork. Dotted black lines indicate the phosphate-sugar backbone of the parent strands, while solid lines indicate the backbones of newly synthesized strands.

The last U of the primer has a 3’ end that DNA polymerase can build off of, so DNA polymerase takes over and adds additional deoxyribonucleotides to the primer sequence. To pair with 3’-TAATA-5’ in the template, DNA polymerase adds 5’-ATTAT-3’ (Figure 20.9, blue sequence).

DNA polymerase uses the 3' OH on the last primer nucleotide to add complementary DNA sequence, building towards the replication fork.
Figure 20.9. Extension off the RNA primer. The active site of DNA polymerase is located at the 3′ hydroxyl on the U and creates a phosphodiester linkage with the 5′ phosphate of the incoming A nucleotide. DNA polymerase advances towards the right replication fork, placing the 3′ hydroxyl of the A in the active site, and the process repeats to extend off this nucleotide. This builds the new nucleic acid strand in the 5′ to 3′ direction while moving along the template in the 3′ to 5′ direction. Dotted black lines indicate the phosphate-sugar backbone of the parent strands, while solid lines indicate the backbones of newly synthesized strands.

Synthesis continues towards the right replication fork, where more single-stranded DNA will be exposed as the replication machinery continues outward. But what about the bottom left side of the replication bubble? Once again, consider the orientation of the bottom template strand, 3’ on the left and 5’ on the right, and the rule that all nucleic acid strands are made 5’ to 3’. Replication on this bottom strand will need to occur from left to right, but initially there is no 3’ end for DNA polymerase to build off of on the left side the replication bubble. Once again, primase acts first, building an RNA primer near the left replication fork (Figure 20.10).

As the replication machinery moves outward, primase synthesizes a new primer at the left replication fork.
Figure 20.10. Creation of an additional RNA primer. Another set of replication enzymes (blue rectangle in Figure 20.7b) is moving outward from the origin to replicate sequences 5′ to 3′ on the left side of the replication bubble. After helicase has unwound DNA adjacent to the left replication fork, primase builds a primer near the left replication fork, working 5′ to 3′ (left to right along the bottom strand). Dotted black lines indicate the phosphate-sugar backbone of the parent strands, while solid lines indicate the backbones of newly synthesized strands.

Once the primer is built, DNA polymerase can add DNA sequences to the primer, building towards the origin (Figure 20.11). In this case, DNA polymerase can only add a few nucleotides before the enzyme runs out of single-stranded template and encounters the first RNA primer. Thus, unlike the DNA sequence that builds towards the fork, which is long and continuously produced, synthesis towards the origin produces short strands, and this synthesis “lags” behind the other strand. For this reason, the strand built towards the origin is referred to as the lagging strand, while the strand built towards the replication fork is the leading strand.

DNA polymerase near the left replication fork adds DNA to the RNA primer, building towards the origin.
Figure 20.11. Extension off the second RNA primer. DNA polymerase uses the 3′ hydroxyl on the A of the primer to extend the strand, creating phosphodiester linkages between deoxyribonucleotides, synthesizing the strand 5′ to 3′ along the bottom template. The bottom left strand that is synthesized towards the origin is the lagging strand, while the bottom right strand that is synthesized towards the right replication fork is the leading strand. On the top strand, DNA replication will begin at the origin, with C being the first nucleotide added. Dotted black lines indicate the phosphate-sugar backbone of the parent strands, while solid lines indicate the backbones of newly synthesized strands.

As mentioned previously, both strands are replicated simultaneously by replication complexes moving outward in both directions from the origin. On your own, draw out the replication of the top strand, remembering that replication starts at the origin and to apply the rules of DNA replication covered above. Also, try to identify the leading and lagging strands in the synthesis of the top strand. As you are working, compare your top strand synthesis with the provided figures of the bottom strand synthesis, paying close attention to the patterns at the left and right replication forks. What similarities and differences do you see between the top and bottom strands and between the two replication forks?

Section 20.4 Elongation and Termination in DNA Replication

Let’s recap the steps thus far: helicase enzymes unwind DNA to form the replication bubble, then primase builds RNA primers for DNA polymerase to build off of. This replication bubble will expand in both directions to allow synthesis to continue, with primase adding more primers as needed, until the entire sequence is replicated. Figure 20.6 shows only DNA bases, whereas the process we have described thus far includes both RNA and DNA nucleotides. As part of elongation in DNA replication, these RNA sequences need to be removed and replaced, which requires the actions of several more enzymes.

Figure 20.11 shows the lagging strand created by DNA polymerase as it extends off the RNA primer and builds towards the origin. At this point, another multifunctional polymerase removes the RNA primer while continuing to build off of the 3′ end of the DNA. This replaces the RNA primer with DNA. The replacement DNA sequence is shown in bold purple in Figure 20.12.

RNA primers are eventually removed and replaced with DNA.
Figure 20.12. Removal and replacement of the RNA primer. As the replication machinery moves towards the origin, RNA primers are removed and replaced by building off the previous DNA fragment. Synthesis stops when the DNA polymerase reaches the 5′ end of the next DNA fragment. Dotted black lines indicate the phosphate-sugar backbone of the parent strands, while solid lines indicate the backbones of newly synthesized strands.

However, this process leaves a gap between DNA fragments, since DNA polymerase is unable to create a phosphodiester linkage with an existing 5’ end. Thus, another enzyme, DNA ligase, is required to connect the 3’ hydroxyl of one DNA fragment with the 5’ phosphate of the adjacent fragment and seal the gap in the DNA backbone (Figure 20.13).

DNA ligase creates a phosphodiester linkage between DNA fragments to finish replication of this section of DNA.
Figure 20.13. Creation of the phosphodiester linkage between DNA strands. DNA ligase synthesizes the phosphodiester linkage between the 3′ end of one DNA fragment and the 5′ end of another existing fragment, creating a continuous strand of DNA. Dotted black lines indicate the phosphate-sugar backbone of the parent strands, while solid lines indicate the backbones of newly synthesized strands.

The sequence shown in Figure 20.6 and subsequent figures is very short for the purposes of demonstration, but real chromosomes are hundreds of thousands to millions of base pairs in length. The mechanics of DNA replication discussed here are similar across all forms of life, but the structure of chromosomes between bacteria and eukaryotes creates unique challenges. As you may recall, bacteria arrange their genomic DNA as one large, circular chromosome. As shown in Figure 20.14, DNA replication initiates at a single origin of replication, and the two replication complexes move away from the origin. These continue in opposite directions along the chromosome until they meet on the other side, and fully replicated chromosomes are separated into two daughter cells as part of binary fission.

The DNA of bacteria is replicated around the circular chromosome with no loss of sequence.
Figure 20.14. Replication of circular bacterial chromosomes from a single origin. Replication complexes move outward from the origin, working their way around the circle until all sequence is replicated with no loss of sequence. Created using BioRender.

In contrast to bacterial chromosomes, chromosomes in eukaryotes are linear and can be much larger. Additionally, the replication enzymes in eukaryotes work much more slowly. To ensure that S phase can be completed in a reasonable amount of time, each chromosome has many origins of replication. DNA replication complexes move outward in both directions from these origins, with replication bubbles merging as these complexes meet (Figure 20.15). As discussed in Chapter 19, Cyclin A-CDK complexes control when replication begins at each origin.

Replication of linear chromosomes requires multiple replication origins, which eventually merge together.
Figure 20.15. Replication of linear eukaryotic chromosomes with multiple origins. Each origin has a set of replication enzymes that move outward once replication is initiated by Cyclin A-CDK. Parental strands are shown in black, while new strands are shown in red. Replication bubbles grow in size, then merge with other replication bubbles, until all available sequence is replicated. Modification of an image by Dr. Todd Nickle and Isabelle Barrette-Ng, from Open Online Genetics, licensed under CC SA 3.0 licensing guidelines.

Section 20.5 The Telomere Problem

The difference in chromosome structure between bacteria and eukaryotes impacts not only the number of origins of replication, but also the ability of the replication process to fully replicate the DNA. The circular chromosomes of bacteria are fully replicated with each cell division. In contrast, the linear chromosomes of eukaryotes are difficult to fully replicate due to the way in which DNA replication occurs. A schematic showing key features of a eukaryotic chromosome is shown in Figure 20.16a.

Linear chromosomes lose sequences at telomeres with each cell division.
Figure 20.16. Telomeres at each end of chromosomes contain repetitive sequences that are not fully replicated. a, Schematic of a chromosome showing the location of telomeres and a short sequence found repeated at telomeres. b, Model of DNA replication at one telomere. RNA primers for lagging strand synthesis are placed near the end of the strand. c, Removal of the RNA primers leaves ~50-100 base pairs of sequence unreplicated, since there is no double-stranded 3′ end to build off of, leading to loss of sequence with each replication.

The region at each end of the chromosomes is called a telomere, and consist of hundreds to thousands of repeats of the same short sequence, such as 5’-TTAGGG-3’/3’-AATCCC-5’. This repetitive sequence doesn’t encode any proteins, but does serve some important functions for the cell. To understand one of these functions, we need to revisit the process of DNA replication and consider how this works at the end of a chromosome.

If we zoom into the end of a chromosome and imagine the replication machinery on the lagging strand (Figure 20.16b). Primase creates RNA primers for DNA polymerase to build off of, but rarely are these primers located at the very end of the sequence. Moreover, the primer on the very end of the chromosome cannot be replaced with DNA, since there is no double-stranded 3’ end for DNA polymerase. Hence, about ~50-100 bases on either end of the chromosome cannot be replicated (Figure 20.16c). Over time, these single-stranded sequences are lost, leading to progressive shortening of chromosomes with each generation. However, because the repetitive sequences at the telomere don’t code for genes, they can be lost without significant consequences for cell function. Once a cell has divided enough times to significantly shorten their telomeres (~50-70 times), further cell division is blocked.

Some cells can divide beyond this limit. These include the germ cells that produce gametes (egg and sperm) for the next generation, or stem cells within specific body tissues that help replace damaged or worn-out cells. These cells express the enzyme telomerase, which actively lengthens telomeres. As shown in Figure 20.17, telomerase has an RNA template sequence that base pairs with the end of the chromosome (Figure 20.17b). Because there is a double-stranded 3’ end and template strand for base pairing, telomerase can elongate the strand (Figure 20.17c, bolded purple bases). This binding and extension by telomerase are repeated many times to lengthen the lagging strand template (Figure 20.17d). Then, lagging strand synthesis occurs as normal to lengthen the other strand (Figure 20.17e).

Cells with telomerase can extend telomeres using an RNA template within the telomerase enzyme.
Figure 20.17. Telomerase lengthens the ends of chromosomes. a, Schematic of sequence at a telomere. b, Telomerase binds to telomere sequences using its RNA template for base pairing. c, Telomerase extends the DNA sequence, building off the 3′ end of the DNA strand and base pairing with its RNA template, then moves outward along the chromosome. d, Further extension and outward movement by telomerase lengthens the lagging strand template. e, Lagging strand synthesis lengthens the other DNA strand, similar to the process shown in Figures 20.10-20.13.

Importantly, while stem cells and germ cells normally express telomerase, other cells in the body do not, even though all cells have the gene for telomerase. Cancer cells, which divide far more than typical cells, have shortened telomeres. However, in many cases, these cancer cells also activate expression of telomerase, allowing continued cell division to occur.

License

Icon for the Creative Commons Attribution-NonCommercial 4.0 International License

Cells and Molecules Copyright © by Katherine Krueger is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License, except where otherwise noted.