Saturday, 8 August 2026

SEQUENCING NOTES

Protein & DNA Sequencing

L6 – Methods in Biology | CSIR-NET / GATE Biotechnology Complete Notes

Protein Sequencing Edman Degradation Mass Spectrometry Tandem MS DNA Sequencing

1. Introduction to Protein and DNA Sequencing

  • Sequencing refers to determining the exact order of building blocks present in a biological molecule.
  • For proteins, sequencing means determining the order of amino acids in a polypeptide chain.
  • For DNA, sequencing means determining the order of nucleotides in a DNA molecule.
  • Protein sequencing is important for identifying proteins, studying structure-function relationships, detecting mutations, confirming recombinant proteins and characterizing post-translational modifications.
  • DNA sequencing is fundamental for genome analysis, gene identification, mutation detection, molecular diagnosis, evolutionary studies, recombinant DNA technology and genomics.
  • The sequence of a protein is normally represented from the N-terminus to the C-terminus.
  • The sequence of a DNA strand is normally written in the 5′ → 3′ direction.
  • Classical protein sequencing methods include Edman degradation and enzymatic or chemical fragmentation followed by analysis of peptides.
  • Modern protein identification and sequencing frequently use mass spectrometry, particularly tandem mass spectrometry or MS/MS.
  • Classical DNA sequencing is strongly associated with the Sanger chain-termination method, whereas modern sequencing includes next-generation sequencing and third-generation approaches.
CSIR-NET Focus: Remember that Edman degradation determines the N-terminal amino acid sequentially, whereas mass spectrometry determines peptide masses and fragmentation patterns that can be used to infer peptide sequences.
Basic Concept of Sequencing Protein Ala Gly Lys Val N → C DNA A T G C 5′ → 3′ Sequence information = ordered arrangement of molecular building blocks

2. Protein Sequencing – Basic Principles

A protein is composed of amino acids connected by peptide bonds. Determining the primary structure of a protein requires identifying the amino acid sequence and, when necessary, identifying disulfide bonds and post-translational modifications.

Major Steps in Classical Protein Sequencing

  • Purification of the protein: The protein should ideally be purified to homogeneity before sequencing.
  • Determination of molecular mass: Molecular mass provides useful information and can help detect whether the purified material corresponds to the expected protein.
  • Identification of terminal residues: The N-terminal and C-terminal residues can be investigated.
  • Cleavage into smaller peptides: Large proteins are generally fragmented into smaller peptides using specific proteases or chemical reagents.
  • Separation of peptides: Peptides may be separated using chromatography or electrophoretic techniques.
  • Sequencing of individual peptides: Edman degradation or mass spectrometry may be used.
  • Overlap analysis: Peptide sequences generated using different cleavage strategies are compared to reconstruct the complete protein sequence.
  • Sequence verification: The experimentally determined sequence may be compared with predicted sequences obtained from DNA or protein databases.
Important: Fragmentation at different sites is extremely useful because overlapping peptide sequences allow the original protein sequence to be reconstructed.

Why is Protein Cleavage Necessary?

  • Large proteins may be difficult to sequence directly.
  • Edman degradation works best when the N-terminal residue is accessible and the peptide is of manageable length.
  • Mass spectrometry provides highly informative spectra when peptides are within an appropriate mass range.
  • Multiple fragments provide overlapping information that helps establish the complete sequence.

3. Edman's Method / Edman Degradation

Edman degradation is a classical method for determining the amino acid sequence from the N-terminus of a peptide or protein. The method was developed by Pehr Edman and became one of the most important classical approaches to protein sequencing.

Principle

  • The free amino group of the N-terminal amino acid reacts with phenyl isothiocyanate (PITC).
  • The reaction forms a phenylthiocarbamoyl derivative.
  • Under acidic conditions, the N-terminal amino acid is released as a phenylthiohydantoin (PTH) derivative.
  • The remaining peptide chain remains intact.
  • The PTH-amino acid is identified chromatographically.
  • The procedure can then be repeated to identify the next amino acid.

Basic Steps of Edman Degradation

  1. Protein or peptide is exposed to PITC under mildly alkaline conditions.
  2. PITC reacts with the free N-terminal amino group.
  3. The peptide is treated under acidic conditions.
  4. The N-terminal amino acid is released as a PTH-amino acid.
  5. The PTH derivative is identified.
  6. The remaining peptide is subjected to another cycle.

Advantages

  • Provides direct information about the N-terminal sequence.
  • Historically important for determining protein primary structures.
  • Individual amino acid residues can be identified sequentially.
  • Useful for purified peptides and proteins with accessible N-termini.

Limitations

  • It is not efficient for very large proteins without prior fragmentation.
  • The N-terminus must be chemically accessible.
  • N-terminal blocking prevents conventional Edman sequencing.
  • Repeated cycles result in progressively lower yields.
  • Post-translationally modified residues can complicate interpretation.
Exam Trick: Edman degradation = N-terminal sequencing + PITC + PTH-amino acid.
Edman Degradation – Simplified Workflow Peptide N–A–B–C–D PITC reaction N-terminal labeling Acid treatment N-terminal release PTH-AA identified Next cycle: N–B–C–D → identify B Repeated cycles reconstruct the N-terminal sequence PITC → PTH-amino acid

4. Proteases and Peptide Fragments

Proteases are enzymes that hydrolyze peptide bonds. In protein sequencing, controlled proteolysis is used to convert a large protein into smaller peptides that can be individually analyzed.

Important Proteases

Protease Major Cleavage Specificity Importance
Trypsin After Lysine (K) or Arginine (R), generally unless followed by Proline One of the most widely used proteases in proteomics
Chymotrypsin Preferentially after aromatic residues such as Phe, Trp and Tyr Produces a complementary peptide set
Pepsin Broad specificity under acidic conditions Useful for producing peptide fragments
CNBr Chemically cleaves at methionine residues Important classical chemical cleavage reagent

Why Use Different Cleavage Reagents?

  • Different enzymes recognize different amino acid residues.
  • A single digestion may not provide enough information to reconstruct the complete sequence.
  • Using two independent cleavage strategies generates overlapping peptide fragments.
  • Overlapping fragments allow the order of peptides to be established.
Example: Suppose one digestion gives peptides ABC, DEF and GHI, while another gives BCD and EFG. The overlaps allow the larger sequence to be reconstructed as ABCDEFGHI.

5. Mass Spectrometry in Protein Sequencing

Mass spectrometry is a powerful analytical technique used to determine the mass-to-charge ratio (m/z) of ions. In proteomics, mass spectrometry is extensively used for protein identification, peptide characterization, sequencing and detection of post-translational modifications.

Basic Components

  • Sample preparation: Protein or peptide samples are prepared and often enzymatically digested.
  • Ionization: Molecules are converted into gas-phase ions.
  • Mass analyzer: Ions are separated according to their mass-to-charge ratio.
  • Detector: The ions are detected and converted into a measurable signal.
  • Mass spectrum: The result is represented as intensity versus m/z.

Common Ionization Methods

  • MALDI: Matrix-Assisted Laser Desorption/Ionization.
  • ESI: Electrospray Ionization.

MALDI

  • The analyte is mixed with a suitable matrix.
  • A laser pulse supplies energy.
  • The matrix assists desorption and ionization of the analyte.
  • MALDI is commonly coupled with TOF analyzers.

ESI

  • The sample is introduced as a solution.
  • A high electric potential produces charged droplets.
  • Solvent evaporation ultimately produces gas-phase ions.
  • ESI is particularly compatible with liquid chromatography.
  • ESI commonly generates multiply charged ions.
Remember: MALDI commonly produces predominantly singly charged ions, whereas ESI frequently produces multiply charged ions.
Mass Spectrometry Workflow Sample Protein/Peptide Ionization MALDI / ESI Mass Analyzer m/z separation Detector Spectrum Mass spectrum = signal intensity plotted against m/z m/z = mass-to-charge ratio

6. Tandem Mass Spectrometry (MS/MS)

Tandem mass spectrometry involves two or more stages of mass analysis. It is particularly useful for determining peptide sequence information.

Basic Concept

  • A peptide is first ionized.
  • The precursor ion is selected based on its m/z.
  • The precursor ion is fragmented into smaller product ions.
  • The product ions are analyzed according to their m/z values.
  • The differences between fragment masses provide sequence information.

Why Does Fragmentation Help Sequence Peptides?

Suppose a peptide has a sequence A-B-C-D. Fragmentation can generate ions corresponding to fragments such as A-B, A-B-C or B-C-D. The mass differences between neighboring fragments correspond approximately to the mass of individual amino acid residues, allowing the sequence to be reconstructed.

Important Ion Series

  • b-ions: Common N-terminal fragment ion series.
  • y-ions: Common C-terminal fragment ion series.
CSIR-NET Tip: In peptide MS/MS interpretation, b-ion series are generally associated with N-terminal fragments and y-ion series with C-terminal fragments.

Peptide Mass Fingerprinting

  • A protein is digested using a specific protease, commonly trypsin.
  • The masses of resulting peptides are measured.
  • The observed peptide mass pattern is compared with theoretical peptide masses from protein databases.
  • A protein can therefore be identified without determining every amino acid directly by Edman degradation.

7. Edman Degradation vs Mass Spectrometry

Feature Edman Degradation Mass Spectrometry
Primary information N-terminal amino acid sequence Peptide mass and fragment information
Key reagent/technology PITC Ionization + mass analysis
Terminal information N-terminal Can provide sequence information from peptide fragments
Speed Relatively slow Rapid and highly sensitive
Modern proteomics Limited use Extensively used
Blocked N-terminus Problematic Often still possible to analyze peptides

8. DNA Sequencing – Basic Principle

DNA sequencing determines the order of nucleotides in DNA. The four major bases in DNA are adenine (A), thymine (T), guanine (G) and cytosine (C).

  • A pairs with T through two hydrogen bonds.
  • G pairs with C through three hydrogen bonds.
  • DNA polymerases synthesize DNA in the 5′ → 3′ direction.
  • A sequencing reaction therefore uses a template DNA strand and a primer.
  • The primer provides the free 3′-OH required for DNA polymerization.

Major DNA Sequencing Approaches

  • Maxam-Gilbert sequencing: Chemical cleavage-based method; historically important but rarely used today.
  • Sanger sequencing: Chain termination method based on incorporation of dideoxynucleotides.
  • Next-generation sequencing: Massively parallel sequencing of large numbers of DNA molecules.
  • Third-generation sequencing: Includes single-molecule sequencing approaches, allowing direct analysis of individual DNA molecules.

9. Sanger Chain-Termination DNA Sequencing

The Sanger method is one of the most important DNA sequencing techniques for examinations. It uses dideoxynucleotides (ddNTPs) to terminate DNA synthesis.

Key Principle

  • DNA polymerase extends a primer using a DNA template.
  • The reaction contains normal deoxynucleotides (dNTPs).
  • It also contains chain-terminating dideoxynucleotides (ddNTPs).
  • A ddNTP lacks the 3′-OH group required for formation of the next phosphodiester bond.
  • Once a ddNTP is incorporated, DNA elongation terminates.
  • Fragments of different lengths are generated.
  • These fragments are separated according to size.
  • The sequence is read from the shortest fragment toward the longest fragment.
Very Important: The absence of a 3′-OH group in ddNTPs is the molecular reason for chain termination.

dNTP vs ddNTP

Feature dNTP ddNTP
2′-OH Present Absent
3′-OH Present Absent
DNA chain extension Can continue Terminates after incorporation
Use in Sanger sequencing Normal extension Chain termination
Sanger Sequencing – Chain Termination Template: 3′ — T A C G G A C T — 5′ New strand: 5′ — A T G C C T G A — 3′ dNTPs extension ddNTP incorporation No 3′-OH chain terminates DNA fragment Different termination positions → fragments of different lengths → sequence

10. Automated Sanger Sequencing

  • Modern Sanger sequencing uses fluorescently labeled ddNTPs.
  • Each ddNTP type may carry a different fluorescent dye.
  • The sequencing products are separated by capillary electrophoresis.
  • A detector records fluorescence as DNA fragments pass through the capillary.
  • The resulting data are represented as a chromatogram.
  • Different colored peaks correspond to different nucleotide bases.
  • The sequence can be read from the shortest fragments toward longer fragments.

Chromatogram Interpretation

  • Each peak corresponds to a detected sequencing product.
  • Peak position indicates fragment size and therefore sequence order.
  • Peak color identifies the nucleotide in a four-color sequencing system.
  • High-quality sequences generally have well-resolved, sharp peaks.
  • Overlapping or irregular peaks may indicate poor sequencing quality, mixed templates or technical problems.

11. General DNA Sequencing Workflow

  1. DNA preparation: Obtain purified DNA of sufficient quality.
  2. Template preparation: Prepare the DNA molecule to be sequenced.
  3. Primer binding: A sequencing primer anneals to the complementary template.
  4. DNA synthesis: DNA polymerase extends the primer.
  5. Termination or signal generation: Depending on the sequencing technology, nucleotide incorporation generates a detectable signal or terminates synthesis.
  6. Separation/detection: Sequencing products or signals are analyzed.
  7. Data processing: Raw sequencing data are converted into nucleotide sequence information.
  8. Quality assessment: Low-quality sequence regions may be removed or rechecked.

12. Brief Overview of Next-Generation Sequencing

Next-generation sequencing technologies allow thousands to millions of DNA fragments to be sequenced in parallel. Unlike classical Sanger sequencing, which generally analyzes one DNA fragment at a time, NGS provides massively parallel sequencing.

  • DNA is fragmented into many pieces.
  • Adapters are commonly attached to DNA fragments.
  • Fragments are amplified or prepared for single-molecule analysis depending on the platform.
  • Sequencing occurs simultaneously across many DNA molecules.
  • Millions of short sequence reads may be generated.
  • Bioinformatic tools are then used for quality control, alignment, assembly and variant analysis.

Important NGS Concepts

  • Read: A sequence generated from an individual DNA fragment.
  • Read length: Number of nucleotides contained in a sequencing read.
  • Coverage/depth: Number of times a genomic position is represented by sequencing reads.
  • Paired-end sequencing: Both ends of a DNA fragment are sequenced.
  • Adapter: Known DNA sequences attached to fragments to facilitate amplification and/or sequencing.
  • Library: Collection of DNA fragments prepared for sequencing.

13. High-Yield CSIR-NET Points

  • Protein sequencing determines the amino acid sequence of a protein.
  • Protein sequences are written from N-terminus to C-terminus.
  • Edman degradation identifies amino acids sequentially from the N-terminus.
  • PITC is the key reagent used in Edman degradation.
  • The released product is identified as a PTH-amino acid.
  • Trypsin cleaves peptide bonds after Lys and Arg under standard conditions, with the important Pro exception.
  • Chymotrypsin preferentially cleaves near aromatic amino acids.
  • CNBr cleaves at methionine residues.
  • Mass spectrometry measures ions according to mass-to-charge ratio.
  • MALDI means Matrix-Assisted Laser Desorption/Ionization.
  • ESI means Electrospray Ionization.
  • ESI commonly generates multiply charged ions.
  • Tandem MS provides fragmentation information useful for peptide sequencing.
  • b-ions are generally associated with N-terminal peptide fragments.
  • y-ions are generally associated with C-terminal peptide fragments.
  • DNA sequencing determines nucleotide order.
  • DNA polymerase synthesizes DNA in the 5′ → 3′ direction.
  • Sanger sequencing uses chain-terminating ddNTPs.
  • ddNTPs lack the 3′-OH group required for continued DNA synthesis.
  • Modern Sanger sequencing uses fluorescent labels and capillary electrophoresis.
  • A DNA chromatogram displays sequencing signals as peaks.
  • NGS performs massively parallel sequencing.
  • Paired-end sequencing obtains sequence information from both ends of a DNA fragment.

14. Common Exam Traps and Conceptual Confusions

Trap 1: Edman vs Sanger

  • Edman degradation → protein sequencing.
  • Sanger sequencing → DNA sequencing.

Trap 2: PITC vs ddNTP

  • PITC → Edman degradation.
  • ddNTP → Sanger DNA sequencing.

Trap 3: 3′-OH

  • dNTP has a 3′-OH.
  • ddNTP lacks the 3′-OH.
  • Therefore, ddNTP incorporation terminates DNA synthesis.

Trap 4: Direction

  • Protein sequence → N → C.
  • DNA sequence → conventionally 5′ → 3′.
  • DNA polymerase adds nucleotides to the 3′-OH of the growing strand.

Trap 5: Mass Spectrometry

  • Mass spectrometry does not simply measure the absolute mass of a neutral molecule.
  • The instrument measures ions according to their m/z.

15. One-Minute Revision Table

Topic Must Remember
Edman degradation N-terminal protein sequencing
PITC Edman reagent
PTH Product used to identify released amino acid
Trypsin Lys/Arg cleavage specificity
CNBr Methionine cleavage
MALDI Matrix-assisted laser ionization
ESI Electrospray; commonly multiply charged ions
MS/MS Precursor ion → fragment/product ions
b-ion N-terminal fragment series
y-ion C-terminal fragment series
Sanger DNA chain termination sequencing
ddNTP Lacks 3′-OH
Chromatogram Peaks representing sequencing signals
NGS Massively parallel sequencing
🧬 CSIR-NET / GATE Practice Quiz – 10 MCQs

Instructions: Select one option for each question and click Submit Quiz. Answers and explanations will appear only after submission.

Q1. Which reagent is specifically associated with Edman degradation?

Q2. Edman degradation primarily determines the sequence from which end of a protein?

Q3. Trypsin preferentially cleaves peptide bonds on the carboxyl side of:

Q4. Which feature of ddNTPs causes chain termination in Sanger sequencing?

Q5. Mass spectrometry primarily separates ions according to their:

Q6. Which ion series is generally associated with C-terminal fragments in peptide MS/MS?

Q7. In Sanger sequencing, which molecule provides the normal substrate for DNA chain extension?

Q8. Which statement about ESI is correct?

Q9. Which statement best describes next-generation sequencing?

Q10. Which statement correctly compares protein and DNA sequencing?

16. Final CSIR-NET Revision Checklist

  • Know the difference between protein sequencing and DNA sequencing.
  • Remember Edman = N-terminal protein sequencing.
  • Remember PITC → PTH-amino acid.
  • Know why proteins are fragmented before detailed sequencing.
  • Remember the specificity of trypsin, chymotrypsin and CNBr.
  • Understand the basic principle of mass spectrometry.
  • Remember that MS measures m/z.
  • Know the difference between MALDI and ESI.
  • Understand precursor and product ions in MS/MS.
  • Remember b-ion = N-terminal fragment.
  • Remember y-ion = C-terminal fragment.
  • Know that Sanger sequencing uses ddNTPs.
  • Remember that ddNTPs lack the 3′-OH.
  • Understand how chain termination produces fragments of different lengths.
  • Know that modern Sanger sequencing commonly uses fluorescent labels and capillary electrophoresis.
  • Understand the basic concept of NGS and massively parallel sequencing.
  • Know the meaning of read, read length, coverage, paired-end sequencing and sequencing library.
Final Memory Formula:

Protein → Edman → PITC → PTH → N-terminal sequence
Protein → Trypsin → Peptides → MS/MS → Fragment ions
DNA → Sanger → ddNTP → No 3′-OH → Chain termination
NGS → Massive parallel sequencing → Millions of reads

No comments:

Post a Comment

Mock Test 5

Mock Test 5: System Physiology CSIR NET Part C Level | Comprehensive Animal Physiology | 30 Questions ...