The molecule that writes the instructions, and the rules that pass them on
Every cell in the human body — with a few specialised exceptions — carries a complete copy of the same molecule: deoxyribonucleic acid, or DNA. It is, in essence, an archive: roughly 3.2 billion chemical "letters" long, encoding somewhere around 20,000 protein-coding genes, plus a vast amount of regulatory material that controls when and where those genes are used. This guide covers what DNA is, how it's organised, how it's copied and passed on, and how errors and inheritance shape the traits of every living thing.
DNA is built from four chemical building blocks called nucleotide bases: adenine (A), thymine (T), guanine (G), and cytosine (C). Each base is attached to a sugar-phosphate backbone, and two such strands wind around each other to form the famous double helix — a structure worked out in 1953 by James Watson and Francis Crick, using critical X-ray diffraction data produced by Rosalind Franklin and Maurice Wilkins.
The two strands are held together by hydrogen bonds between bases on opposite strands, and the pairing is strict: adenine always pairs with thymine, and guanine always pairs with cytosine. This rule — complementary base pairing — is the single most important fact in molecular biology. It means each strand contains enough information to reconstruct its partner, which is precisely what makes DNA copyable, and it means the two strands run in opposite chemical directions (described as 5′ to 3′ and 3′ to 5′), a detail that turns out to matter enormously for how DNA is read and copied.
A single strand of human DNA, fully stretched out, would be about two metres long — yet it fits inside a cell nucleus roughly 6 micrometres across. This is achieved through an extraordinary degree of packaging: DNA winds tightly around clusters of proteins called histones, forming bead-like structures called nucleosomes, which coil further into the dense material called chromatin, which condenses further still into the X-shaped structures called chromosomes that are visible under a microscope during cell division.
Humans carry 23 pairs of chromosomes — 46 in total — one member of each pair inherited from each parent. Twenty-two pairs are autosomes, and the 23rd pair determines sex: two X chromosomes typically produce a female, one X and one Y typically produce a male. The full set of an organism's DNA is called its genome; the Human Genome Project, completed in 2003, produced the first essentially complete reference sequence of the human genome after 13 years of international collaboration.
| Term | What it means |
|---|---|
| Gene | A stretch of DNA that codes for a protein or functional RNA molecule |
| Allele | One of two or more alternative versions of a gene |
| Genotype | An organism's specific combination of alleles at a given gene or genes |
| Phenotype | The observable trait that results from a genotype, shaped also by environment |
| Chromatin | The combined material of DNA wound around histone proteins |
| Genome | An organism's complete set of DNA |
Before a cell divides, it must copy its entire genome so that each daughter cell receives a full set. This process, DNA replication, is semiconservative — each new double helix consists of one original ("parental") strand and one newly synthesised strand, a mechanism confirmed experimentally by Meselson and Stahl in 1958.
The enzyme helicase unwinds and separates the two original strands, exposing them so that DNA polymerase can build a new complementary strand against each one, following the same base-pairing rule that holds the helix together. Because the two original strands run in opposite directions, replication proceeds differently on each: the leading strand is synthesised continuously, while the lagging strand is built in short, discontinuous fragments called Okazaki fragments, later stitched together by the enzyme DNA ligase.
DNA polymerase is remarkably accurate — it makes roughly one error per billion bases copied, thanks to a built-in "proofreading" function that checks and corrects mismatched bases as it goes. Even so, on a genome of 3.2 billion bases, a handful of errors slip through every replication cycle, which is one of the ultimate sources of the genetic variation described further below.
DNA itself does very little directly — its instructions have to be read out and acted on. A gene is first transcribed into a working copy made of messenger RNA (mRNA), which is then translated by ribosomes into a chain of amino acids that folds into a functional protein. This two-step flow of information — DNA to RNA to protein — is known as the central dogma of molecular biology. The detailed machinery of transcription and translation is covered in full in our guide to how proteins are made; what matters here is that DNA functions as a stable, permanent reference copy, while RNA is the disposable working copy actually used to build things.
A mutation is any change to a DNA sequence — from a single swapped base to the loss or duplication of an entire chromosome. Mutations arise from replication errors, exposure to radiation or certain chemicals, or spontaneous chemical decay of DNA bases over time.
| Type | What happens | Example consequence |
|---|---|---|
| Point mutation (substitution) | One base is swapped for another | May change a single amino acid, or none at all, depending on the genetic code's redundancy |
| Insertion / Deletion | One or more bases are added or removed | Can shift the "reading frame" of the entire gene downstream, often severely disrupting the protein |
| Chromosomal duplication | A segment of DNA is copied an extra time | Can create raw material for new gene functions over evolutionary time |
| Aneuploidy | An abnormal number of whole chromosomes | An extra copy of chromosome 21 causes Down syndrome |
Most mutations are neutral — they occur in non-coding regions, or don't change the resulting protein, or change it in a way that has no meaningful effect. Some are harmful, disrupting a protein's function and potentially causing disease; a smaller number are beneficial, providing a slight advantage that natural selection can act on. Over enormous timescales, the accumulation of mutations across a population is the raw material of evolution — without mutation, there would be no genetic variation for natural selection to work with at all.
The foundational rules of inheritance were worked out by the Austrian monk Gregor Mendel in the 1860s, through careful breeding experiments with pea plants — decades before DNA's role was even suspected. Mendel showed that traits are passed on as discrete units (what we now call genes), inherited in pairs, one from each parent.
When the two inherited alleles differ, one is often dominant — its trait is expressed — while the other is recessive, its effect masked unless both inherited copies are the recessive version. A classic example is cystic fibrosis: a person needs two copies of the recessive CFTR mutation to develop the disease; a single copy makes them an unaffected "carrier" who can still pass the allele to their children.
Real inheritance is often more complex than simple dominant/recessive pairs. In incomplete dominance, heterozygous individuals show a blended intermediate trait. In codominance, both alleles are fully expressed simultaneously (as with the AB blood type). Most human traits of real interest — height, skin colour, susceptibility to common diseases — are polygenic, influenced by many genes at once, each with a small effect, layered on top of environmental factors. This is why so few real-world human traits follow a clean Mendelian pattern.
Not every heritable change to how a gene behaves involves altering its underlying DNA sequence at all. Epigenetics studies chemical modifications that switch genes on or off without changing the letters of the code itself — most importantly DNA methylation (the addition of small chemical tags directly to DNA, typically silencing the affected gene) and modifications to the histone proteins DNA is wound around, which can loosen or tighten how accessible a gene is to the cellular machinery that reads it.
Epigenetic patterns are how a skin cell and a neuron, despite carrying identical DNA, end up looking and behaving so differently — different genes are switched on and off in each. Some epigenetic marks can even be influenced by environment and, in certain documented cases, passed on to offspring, a finding that has reshaped how biologists think about the boundary between nature and nurture.
Understanding DNA's structure and behaviour has driven a series of technologies that now touch medicine, agriculture, and forensics directly:
PCR (Polymerase Chain Reaction), developed by Kary Mullis in 1983, exploits DNA polymerase to exponentially amplify a specific DNA sequence from a tiny starting sample — doubling it roughly 20-30 times over — making it detectable and usable. It underlies most modern genetic testing, including COVID-19 diagnostic tests.
DNA sequencing reads out the actual order of bases in a sample. Sanger sequencing, developed in 1977, made the original Human Genome Project possible; modern next-generation sequencing (NGS) platforms can now sequence an entire human genome in under a day for a few hundred dollars, down from roughly $3 billion and 13 years for the first one.
CRISPR-Cas9, adapted from a bacterial immune defence mechanism and developed into a gene-editing tool by Jennifer Doudna and Emmanuelle Charpentier in 2012 (Nobel Prize in Chemistry, 2020), allows researchers to cut DNA at a precisely targeted sequence and either disable a gene or insert a new one. It has dramatically accelerated genetic research and enabled early gene-therapy treatments for conditions such as sickle cell disease.
Perhaps the most striking fact about DNA is how universal it is: the same four-letter alphabet, the same triplet genetic code, and largely the same core replication machinery are shared by bacteria, plants, fungi, and animals alike, all tracing back to a common ancestor billions of years ago. A gene taken from a jellyfish can be inserted into a mouse and still be read correctly by the mouse's cellular machinery, because the underlying code is shared across nearly all of life. That universality is what makes genetics not just a description of heredity, but one of the deepest unifying threads in all of biology.
This document provides a general scientific overview of DNA and genetics for educational purposes.