Molecular Basis of Inheritance

The molecular story of how DNA stores, copies, and expresses genetic information through replication, transcription, and translation, plus operon-level regulation that decides when each gene speaks.

Part of Unit 21: Molecular Genetics & Gene Regulation in the NEET Biology syllabus.

Molecular Basis of Inheritance Introduction: From Heredity to Molecules Heredity at the molecular level is a story about one polymer doing four jobs at once. DNA stores the instructions, copies itself before a cell divides, hands a working transcript to the protein-making machinery, and quietly switches genes on or off as the cell's needs change. The chapter that begins with Watson and Crick's 1953 model ends with bacteria deciding whether to digest lactose, and every step in between is governed by base pairing, directionality, and a handful of dedicated enzymes. NEET problems repeatedly reward students who know the order of events, the enzymes by name, and the historical experiments that pinned each idea down. Three big experiments anchor the field. Frederick Griffith's 1928 transformation observation hinted that something in dead virulent pneumococci could re-program live avirulent ones. Avery, MacLeod, and McCarty showed in 1944 that the transforming agent was DNA , not protein. Meselson and Stahl confirmed in 1958 that replication is semi-conservative by tracking heavy and light nitrogen across bacterial generations. Modern molecular biology is essentially the steady mechanistic filling-in of the picture those three experiments outlined. Historical sequence to memorise: Griffith ( 1928 ) Avery--MacLeod--McCarty ( 1944 ) Hershey--Chase ( 1952 , 32 P vs 35 S ) Watson--Crick model ( 1953 , using Franklin's X-ray data) Meselson--Stahl ( 1958 , semi-conservative). neet-alert I. The Structure of DNA DNA is a double-stranded helix in which two polynucleotide strands run antiparallel . One strand reads 5' 3' while its partner reads 3' 5' . The backbone of each strand is a repeating sugar--phosphate chain in which deoxyribose is joined to the next residue by a phosphodiester bond. The nitrogenous bases project inward and pair across the helix, holding the structure together by hydrogen bonds and base-stacking interactions. The characteristic spiral structure of DNA , formed by two complementary, antiparallel polynucleotide strands wound around a common axis with a pitch of about 3.4 nm and roughly 10 base pairs per turn. Double Helix Antiparallel Strands Two DNA strands oriented in opposite chemical directions, so that the 5' end of one strand lies opposite the 3' end of the other. This geometry is required for correct base pairing and for the directionality of replication. Each base is either a purine ( adenine or guanine , double-ringed) or a pyrimidine ( cytosine or thymine , single-ringed). The pairing rule is strict: A pairs with T through two hydrogen bonds, and G pairs with C through three hydrogen bonds. The extra hydrogen bond in G--C pairs is why GC-rich regions of DNA are slightly harder to melt. Empirical observations by Erwin Chargaff that in any DNA sample A = T and G = C , and therefore total purines equal total pyrimidines ( A+G = T+C ). These ratios were a major clue pointing toward complementary base pairing. Chargaff's Rules Antiparallel sugar--phosphate backbones with complementary A--T and G--C base pairs. The major and minor grooves provide access points for proteins that read or modify DNA. DNA: Double, Deoxyribose, Thymine. RNA: Ribose, Uracil, single. DNA RNA Feature DNA versus RNA at a Glance Sugar Deoxyribose ( 2' -H) Ribose ( 2' -OH) Strandedness Double-stranded (usually) Single-stranded (usually) Pyrimidine bases Cytosine, Thymine Cytosine, Uracil Stability Chemically more stable; long-term store Less stable; transient working copy Main forms B-DNA (in vivo), A, Z mRNA, tRNA, rRNA, snRNA, miRNA remember B-form DNA is the cellular default: right-handed, 10 bp per turn, pitch 3.4 nm, distance between adjacent bases 0.34 nm. Z-DNA is left-handed and rare; A-DNA appears in dehydrated samples. Evidence that DNA Is the Genetic Material Before Watson and Crick's model, the chemical identity of the genetic material was contested. Avery, MacLeod, and McCarty isolated the active transforming principle from heat-killed virulent pneumococci and showed that protease and RNase left transforming ability intact, but DNase destroyed it. Hershey and Chase used radiolabeled bacteriophages, tagging protein with 35 S and DNA with 32 P , and demonstrated that only the phosphorus label entered infected bacteria. The verdict was the same in both lines of evidence: information is carried by DNA . The double helix was deduced largely from Rosalind Franklin and Maurice Wilkins' X-ray diffraction images (especially Photo 51) combined with Chargaff's base-ratio data. Watson, Crick, and Wilkins shared the 1962 Nobel; Franklin had died in 1958. The Watson and Crick model was built from scratch by the two of them alone. II. Packaging: From DNA to Chromatin A diploid human cell contains roughly two metres of DNA distributed across 46 chromosomes, yet it must fit inside a nucleus only a few microns across. The cell solves this by wrapping DNA around small, positively charged proteins called histones. The negative phosphate backbone is attracted to the basic histone surfaces (rich in lysine and arginine), and the complex they form is the fundamental packaging unit of the chromosome. The basic unit of DNA packaging; a segment of about 147 base pairs wrapped roughly 1.65 times around an octamer of histone proteins (two copies each of H2A, H2B, H3, and H4). Nucleosome Chromatin The complex structure formed by nuclear DNA and its associated proteins (histones and non-histone scaffold proteins). Euchromatin is loosely packed and transcriptionally active; heterochromatin is densely packed and largely silent. Beads-on-a-string nucleosomes coil into the 30 nm fibre and ultimately into the metaphase chromosome. Histone H1 sits on the linker DNA between nucleosomes and clamps them together. Levels of DNA Compaction Naked DNA double helix: the 2 nm fibre — the working width of the molecule itself. Nucleosome (beads on a string): 11 nm fibre; 147 bp of DNA per histone octamer, separated by linker DNA. Solenoid / chromatin fibre: 30 nm fibre stabilised by histone H1; six nucleosomes per turn. Looped domains: the 30 nm fibre is anchored to a non-histone scaffold, forming loops of 300 nm. Metaphase chromosome: maximal condensation, 1400 nm wide, visible under the light microscope. remember Histone tail modifications (acetylation, methylation, phosphorylation) form the so-called histone code. Acetylation generally loosens chromatin and favours transcription; methylation can do either, depending on the residue. III. DNA Replication: The Semi-Conservative Process Replication is the duplication of the genome before cell division. Three logically possible models existed at first: conservative (the original duplex stays intact and a brand new duplex is made), semi-conservative (each daughter has one parent and one new strand), and dispersive (each daughter strand is a patchwork). The Meselson--Stahl experiment in 1958 settled the question in favour of the semi-conservative model. Semi-Conservative Replication Mode of DNA replication in which each parent strand serves as a template for a newly synthesised complementary strand, so that every daughter molecule contains one parental strand and one new strand. The Meselson--Stahl Experiment Meselson and Stahl grew E. coli for many generations in medium containing the heavy nitrogen isotope 15 N , so that all DNA contained heavy nitrogen. They then transferred the cells to medium containing only light 14 N and sampled DNA at successive generations. The DNA was centrifuged in a caesium chloride density gradient, which separates molecules by buoyant density. Heavy parental DNA ( 15 N / 15 N ) gives a single dense band. After one generation in 14 N , all DNA is intermediate (hybrid). After two generations, half is hybrid and half is light — the signature of semi-conservative replication. Generation 0 (parent in 15 N ): one heavy band — all DNA is fully labelled. Generation I (one round in 14 N ): one intermediate (hybrid) band — disproves the conservative model. Generation II (two rounds in 14 N ): two bands, one intermediate and one light, in a 1 : 1 ratio — disproves the dispersive model. Generations III+: the light band intensifies and the hybrid band fades, confirming semi-conservative replication. Predicted Bands at Each Generation Machinery at the Replication Fork The Y-shaped junction at which the parental double helix is being unwound and the two new daughter strands are being synthesised. Bidirectional replication uses two forks moving in opposite directions from an origin. Replication Fork A short RNA segment, typically 10 -- 12 nucleotides long, synthesised by primase to provide a free 3' -OH group that DNA polymerase can extend. DNA polymerases cannot initiate a strand from scratch. Primer Okazaki Fragments Short stretches of DNA, 1000 -- 2000 nucleotides long in bacteria and 100 -- 200 in eukaryotes, synthesised discontinuously on the lagging strand. Each fragment requires its own primer and is later joined to its neighbour by DNA ligase. Helicase melts the duplex while SSBs hold strands apart. Primase lays down RNA primers and DNA polymerase extends them. The leading strand grows continuously toward the fork; the lagging strand grows in Okazaki fragments away from the fork. A complementary view of the fork showing the geometry of leading and lagging strand synthesis and the role of ligase in sealing Okazaki fragments. Helicase Unwinds the parental duplex by breaking hydrogen bonds ATP-dependent; creates the fork Single most asked unwinding enzyme Topoisomerase / DNA gyrase Relieves supercoiling ahead of the fork Cuts and re-ligates one or both strands Target of quinolone antibiotics Single-strand binding proteins (SSBs) Coat separated strands and prevent reannealing Non-catalytic, structural Protect from nucleases too Primase Synthesises short RNA primers Provides the first 3' -OH Required on both strands, repeatedly on lagging DNA Polymerase III Main replicative enzyme; extends in 5' 3' direction Has 3' 5' exonuclease (proof-reading) High processivity DNA Polymerase I Removes RNA primers and fills gaps with DNA Has both 5' 3' polymerase and 5' 3' exonuclease Key for primer removal DNA Ligase Seals nicks in the phosphodiester backbone ATP- or NAD-dependent Joins Okazaki fragments Helicase Hates double helix; Primase Pioneers RNA; Pol III Pumps DNA; Pol I Patches; Ligase Locks. Function Key Detail NEET Pointer Enzyme / Factor Enzymes of DNA Replication (Prokaryotic) DNA polymerases synthesise strictly 5' 3' . Because the two parental strands are antiparallel, one daughter strand ( leading ) is made continuously and the other ( lagging ) is made in pieces. This single geometric fact is the reason Okazaki fragments exist. remember Leading strand: synthesised continuously in the 5' 3' direction toward the moving fork; needs only one primer. Lagging strand: synthesised discontinuously, away from the fork, as a series of Okazaki fragments; each fragment needs its own primer. Common rule: both strands are extended by DNA polymerase III in the 5' 3' direction; the difference lies in geometry, not chemistry. Leading vs Lagging Strand DNA Polymerase can start synthesising a new strand from scratch. It cannot. DNA polymerase always requires a free 3' -OH group, which is provided by an RNA primer laid down by primase. The primer is later removed by Pol I in bacteria. Because synthesis is restricted to the 5' 3' direction, only the leading strand is continuous. The lagging strand is built in short Okazaki fragments, each with its own primer, later joined by ligase. Replication happens continuously on both strands. In eukaryotes, replication is bidirectional but slower ( 50 nucleotides per second) than in E. coli ( 2000 per second). Eukaryotic chromosomes use thousands of origins; the bacterial chromosome typically uses one ( oriC ). neet-alert IV. Transcription: DNA to RNA Transcription is the synthesis of an RNA molecule from a DNA template. Only one of the two DNA strands, the template (also called the antisense or 3' 5' strand), is read. The other strand, identical in sequence to the RNA except for T versus U, is called the coding or sense strand. RNA polymerase reads the template 3' 5' and synthesises RNA in the 5' 3' direction, using the base pairing rules A--U and G--C. Transcription Unit The stretch of DNA that is transcribed as a single RNA. It is bounded by a promoter upstream (where RNA polymerase binds) and a terminator downstream (where transcription ends). The promoter defines the template strand. Promoter at the 5' end of the coding strand, structural gene in the middle, terminator at the 3' end. The template strand reads 3' 5' and the nascent RNA grows 5' 3' . Bacteria use a single RNA polymerase made of a core enzyme and a sigma ( ) factor; sigma binds to the -35 and -10 boxes of the promoter and lets the core polymerase initiate. Eukaryotes use three nuclear RNA polymerases with sharply different jobs. RNA Pol I 28 S, 18 S, 5.8 S rRNAs Insensitive Nucleolus RNA Pol II All mRNAs, most snRNAs, miRNAs Highly sensitive Nucleoplasm RNA Pol III tRNAs, 5 S rRNA, some snRNAs Moderately sensitive Nucleoplasm Eukaryotic RNA Polymerases Enzyme RNAs Produced Sensitivity to -amanitin Where It Acts I = rRNA, II = mRNA, III = tRNA + 5S rRNA Post-Transcriptional Processing in Eukaryotes The primary transcript in eukaryotes is called hnRNA (heterogeneous nuclear RNA). It is not yet usable as a message. Three modifications convert it into mature mRNA: a 5' cap is added, a 3' poly-A tail is appended, and internal non-coding stretches are spliced out. Splicing Removal of non-coding introns from pre-mRNA and joining of coding exons. Catalysed by the spliceosome, a large ribonucleoprotein complex built from snRNAs (U1, U2, U4, U5, U6) and proteins. Addition of 200 -- 250 adenine residues to the 3' end of mRNA in a template-independent manner. The poly-A tail increases mRNA stability, helps export from the nucleus, and aids ribosome recruitment. Polyadenylation (Tailing) 5' Capping Addition of a 7 -methylguanosine cap, joined by an unusual 5' 5' triphosphate linkage, to the 5' end of mRNA . The cap protects against 5' exonucleases and is recognised by the small ribosomal subunit during translation initiation. Capping at 5' , polyadenylation at 3' , and removal of introns yield mature mRNA. Exons (coding) are stitched together; introns (non-coding) are excised as lariats. remember Exons exit the nucleus as part of mature mRNA. Introns are interrupting sequences that stay in the nucleus and are degraded. Prokaryotes have no introns and do no capping or tailing, so transcription and translation are coupled. The primary transcript (pre-mRNA) is the final message used for protein synthesis. Pre-mRNA must be capped, tailed, and spliced before it is exported as mature mRNA. Only then can it engage a ribosome for translation. V. The Genetic Code The genetic code is the mapping between three-nucleotide codons in mRNA and the twenty standard amino acids. Of the 4 3 = 64 possible codons, 61 specify amino acids and 3 are stop signals. Marshall Nirenberg, Har Gobind Khorana, and Severo Ochoa worked out the code in the 1960s, with Khorana making heteropolymers of defined sequence and Nirenberg using poly-U to translate phenylalanine. Codon A triplet of consecutive bases in mRNA that specifies one amino acid or a stop signal. The reading frame is set by the start codon AUG and read 5' 3' without overlap. Triplet: each codon is three bases; 61 sense codons plus 3 stop codons. Degenerate: most amino acids are specified by more than one codon (Leucine, Arginine, and Serine each have six). Unambiguous and specific: one codon codes for exactly one amino acid in a given context. Non-overlapping and commaless: bases are read in non-overlapping triplets with no punctuation between codons. Universal: the same code is used across nearly all organisms, with minor exceptions in mitochondria and some protozoa. Polarity: the code is read in a fixed 5' 3' direction on mRNA. Salient Features of the Genetic Code neet-alert Start codon: AUG (methionine in eukaryotes; formyl-methionine in bacteria). Stop codons: UAA (ochre), UAG (amber), UGA (opal). AUG is the only commonly used start codon and also codes for internal methionines. Amino Acid / Signal Why It Matters UAA, UAG, UGA — "U Are Annoying, U Are Going, U Go Away" Selected Codons Frequently Tested Codon AUG Methionine / Start Sets the reading frame UAA , UAG , UGA Stop signals Recognised by release factors, not tRNAs UUU / UUC Phenylalanine Decoded in Nirenberg's poly-U experiment GAG / GAA Glutamate Sickle-cell mutation changes GAG GUG (Glu Val) Stop codons: U A A , U A G , U G A — remember as U Are Away, U Are Gone, U Go Away.'' VI. Translation: mRNA to Polypeptide Translation is protein synthesis directed by mRNA. It happens on ribosomes in the cytoplasm and requires three classes of RNA, an army of factors, GTP, and ATP. The reading is done by transfer RNA molecules that carry an amino acid at one end and a complementary anticodon at the other. Ribosome The ribonucleoprotein machine of protein synthesis. The eukaryotic ribosome is 80 S (with 60 S and 40 S subunits), the prokaryotic ribosome is 70 S (with 50 S and 30 S subunits). Both contain three tRNA binding sites: A (aminoacyl), P (peptidyl), and E (exit). An adaptor RNA molecule, about 76 -- 90 nucleotides long, that carries a specific amino acid at its 3' end (the CCA acceptor end) and reads the mRNA codon through its anticodon loop. There is at least one tRNA species for each amino acid. Transfer RNA (tRNA) Anticodon A three-base sequence in the anticodon loop of a tRNA that base-pairs antiparallel with the mRNA codon. It ensures that the correct amino acid is delivered to the growing polypeptide. Clover-leaf secondary structure of tRNA showing the 5' end, the acceptor arm with its 3' -CCA where the amino acid attaches, the D loop, the anticodon loop with the three-base anticodon, the variable arm, and the T C loop. Stages of Translation Activation (amino acid charging): aminoacyl-tRNA synthetases attach the right amino acid to the right tRNA, using ATP. The product is an aminoacyl-tRNA, with the amino acid linked to the 3' -CCA end. Initiation: the small ribosomal subunit assembles on the 5' cap, scans to the first AUG , and is joined by the initiator tRNA (Met-tRNA in eukaryotes). The large subunit then docks. Elongation: an aminoacyl-tRNA enters the A site, a peptide bond is formed between the amino acid in A and the chain in P (catalysed by the 23 S rRNA acting as a ribozyme), and the ribosome translocates by one codon, moving the spent tRNA from P to E. Termination: a stop codon enters the A site, release factors bind, the polypeptide is released from the final tRNA, and the ribosomal subunits dissociate. Ribosome with mRNA threaded through, charged tRNAs in the A and P sites, peptide bond formation catalysed by the large-subunit rRNA, and translocation that moves tRNAs from A to P to E. The peptidyl transferase activity of the ribosome is a ribozyme function. The catalytic site lies in the 23 S rRNA of the large subunit, not in any protein. This was Tom Cech and Sidney Altman's broader lesson: RNA can catalyse reactions. remember Ribosome sites: A rrive, P eptide-bond, E xit. A site brings the new amino acid, P site holds the growing peptide, E site spits out the empty tRNA. Ribosomes are about two-thirds RNA by mass. The peptide bond is formed by the rRNA (a ribozyme), and ribosomal proteins largely play structural and stabilising roles. The ribosome itself is mostly protein, and proteins do the catalysis. VII. Gene Regulation: The Lac Operon Cells do not transcribe all their genes all the time. In bacteria, sets of related genes are organised into operons and switched on or off together by regulatory proteins. The classical example, worked out by Francois Jacob and Jacques Monod in 1961, is the lac operon of E. coli , which controls lactose metabolism. Operon A unit of prokaryotic gene regulation consisting of structural genes, a common promoter and operator, and one or more associated regulatory genes. All the structural genes in an operon are transcribed together as a single polycistronic mRNA. An inducible operon in E. coli that encodes the three enzymes needed for lactose use ( -galactosidase, permease, transacetylase). Structural genes ( z , y , a ) are transcribed only when lactose is present and glucose is absent. Lac Operon Top half: the central dogma — DNA mRNA protein. Bottom half: the lac operon in its repressed and induced states, showing how allolactose inactivates the repressor and lets RNA polymerase transcribe z , y , a . i gene Regulatory gene Encodes the lac repressor (constitutively expressed) Promoter ( P ) DNA binding site for RNA polymerase Determines transcription start and direction Operator ( O ) DNA site between promoter and z Repressor binds here to block transcription z gene Structural gene Encodes -galactosidase (cleaves lactose to glucose + galactose) y gene Structural gene Encodes permease (lactose uptake into the cell) a gene Structural gene Encodes transacetylase (modifies certain galactosides) i-P-O-Z-Y-A. "I Put On a Zippy Yellow Apron." Identity Role Components of the Lac Operon Component Lactose absent (OFF): the active repressor binds the operator and physically blocks RNA polymerase from transcribing z , y , a . The cell does not waste resources making lactose-handling enzymes. Lactose present (ON): a few molecules of lactose are converted by basal -galactosidase into allolactose, the true inducer. Allolactose binds the repressor, changes its shape, and the inactive repressor falls off the operator. Glucose absent (boosted ON): low glucose raises cAMP, which binds the catabolite activator protein (CAP). CAP--cAMP binds upstream of the promoter and recruits RNA polymerase strongly. Transcription rises sharply. Glucose also present (damped): cAMP is low, CAP cannot bind, and transcription of the operon is only weak even when lactose is present. This is catabolite repression — glucose is preferred over lactose. How the Lac Operon Switches The lac operon is negative inducible : control is negative because a repressor turns it OFF, and inducible because the substrate (via allolactose) flips it ON. Contrast with the trp operon, which is negative repressible (a co-repressor flips it OFF when tryptophan is plentiful). neet-alert The true inducer is allolactose, an isomer formed from lactose by basal -galactosidase. In the lab, the non-metabolisable analogue IPTG is often used because it induces without being broken down. Lactose itself is the inducer of the lac operon. VIII. Transcription versus Translation versus Replication Template Both DNA strands One DNA strand (template) mRNA Product Two daughter DNA duplexes RNA (mRNA / tRNA / rRNA) Polypeptide Main enzyme DNA polymerase RNA polymerase Ribosome (rRNA + proteins) Primer needed Yes (RNA primer) No Initiator tRNA, not a primer Direction of synthesis 5' 3' 5' 3' N -terminus C -terminus Location (eukaryotes) Nucleus Nucleus Cytoplasm / rough ER DNA copies DNA; DNA writes RNA; RNA reads protein. Replication Transcription Translation Replication, Transcription, and Translation Side by Side Feature Replication, transcription, and translation all share the 5' 3' rule for nucleotide chain growth. Protein chain growth is read N to C , but the mRNA driving it is still read 5' 3' . remember IX. The Human Genome Project and DNA Fingerprinting The Human Genome Project, launched in 1990 and completed in 2003, sequenced the roughly 3.2 10 9 base pairs of the human genome. It found about 20 , 000 -- 25 , 000 protein-coding genes — far fewer than the pre-project estimate. Less than 2 % of the genome codes for protein; the rest is regulatory DNA, repeats, and noncoding RNA genes. Chromosome 1 has the most genes; the Y chromosome has the fewest. An identification technique that compares variable regions of DNA, especially VNTRs (variable number tandem repeats) or short tandem repeats, to produce an individual-specific banding pattern. Introduced by Alec Jeffreys in 1984 and used in forensics, paternity testing, and population genetics. DNA Fingerprinting neet-alert HGP numbers to memorise: 3.2 billion base pairs, 20 , 000 -- 25 , 000 genes, less than 2 % coding. Average gene size 3000 bp. The longest gene ( dystrophin ) spans 2.4 million bp; chromosome 1 carries the most genes, chromosome Y the fewest. X. Things to Remember remember Historical milestones in order: Griffith's transformation Avery--MacLeod--McCarty Hershey--Chase Chargaff's ratios Watson--Crick model Meselson--Stahl semi-conservative proof. remember Enzyme roles to know cold: helicase (unwinds), topoisomerase / gyrase (relieves supercoiling), primase (makes RNA primer), Pol III ( 5' 3' synthesis with 3' 5' proof-reading), Pol I (removes primers, fills gaps), ligase (seals nicks). Directionality: replication and transcription proceed 5' 3' on the newly made strand. Synthesis on the lagging strand is discontinuous, in Okazaki fragments. remember Eukaryotic gene expression pathway: DNA Transcription pre-mRNA Capping, Tailing, Splicing mature mRNA Translation Polypeptide . remember The lac operon demonstrates negative inducible control. The repressor binds the operator in the absence of lactose; allolactose binds the repressor in the presence of lactose and switches the operon ON. remember remember Clinical connection: defects in DNA repair or in histone modification machinery can lead to cancer. Xeroderma pigmentosum (defective nucleotide excision repair) is a classical example. Base pair bonds: A -- T = 2 hydrogen bonds; G -- C = 3 hydrogen bonds. "AT pair, two; GC pair, three." Order of post-transcriptional processing: C ap, T ail, S plice — "Cats Tail Slips" — capping first ( 5' end), then tailing ( 3' end), then splicing of introns. XI. Closing Summary From the antiparallel double helix discovered in 1953 to the lac operon's elegant logic worked out in 1961, this chapter charts the molecular grammar of inheritance. DNA encodes the message; replication preserves it; transcription transcribes a working copy; splicing, capping, and tailing polish the message; translation builds the protein; and operons decide which messages are even worth writing in the first place. Mastery here means knowing not just what each player does, but when and why the cell deploys it — and being able to read a diagram or an experiment and tell at a glance which step in the dogma is being illustrated.