Saturday, 8 August 2026

Unit 11: Bioinformatics and Computational Biology

CSIR-NET LIFE SCIENCES

Unit 11: Bioinformatics and Computational Biology

Complete Topic-Wise Syllabus

Sequence Analysis • Structural Bioinformatics • Drug Design • Systems Biology

A. Major Bioinformatic Resources

๐Ÿงฌ 1. Sequence Databases

  • Nucleotide sequence databases
  • Protein sequence databases
  • Genome databases
  • Transcriptome databases
  • Reference sequence databases
  • Curated and non-curated databases
๐Ÿ“š Important Sequence Database Concepts:
  • Primary databases contain experimentally submitted sequence information.
  • Secondary databases contain derived or curated information.
  • Protein databases provide amino-acid sequence information and functional annotations.
  • Nucleotide databases contain DNA and RNA sequence information.

๐Ÿงช 2. Gene Expression Databases

  • Gene expression profiles
  • Transcriptomic data
  • RNA expression datasets
  • Microarray data
  • RNA-seq data
  • Expression patterns under different conditions
  • Tissue-specific expression

๐ŸงŠ 3. 3D Structure Databases

  • Experimental macromolecular structures
  • Protein structures
  • DNA structures
  • RNA structures
  • Protein-ligand complexes
  • Structural coordinates

๐Ÿ” 4. Pattern and Sequence Databases

  • Protein motifs
  • Functional domains
  • Conserved sequence patterns
  • Sequence signatures
  • Repeats
  • Profiles

๐Ÿ—‚️ 5. Database Classification

Database Type Main Information
Sequence Database DNA, RNA and protein sequences
Expression Database Gene expression and transcriptomic data
Structure Database 3D structures of biomolecules
Pattern Database Motifs, domains and sequence signatures

B. Basic Concepts of Sequence Analysis

๐Ÿ”Ž 1. Database Searches

  • Searching biological databases using nucleotide or protein sequences
  • Identification of homologous sequences
  • Functional annotation
  • Identification of conserved regions
  • Detection of evolutionary relationships

⚡ 2. BLAST

BLAST stands for Basic Local Alignment Search Tool. It identifies regions of local similarity between biological sequences.

  • DNA sequence searches
  • Protein sequence searches
  • Identification of homologues
  • Functional annotation
  • Similarity searching
๐Ÿ”ฌ Important BLAST Types:
  • BLASTN – nucleotide query against nucleotide database
  • BLASTP – protein query against protein database
  • BLASTX – translated nucleotide query against protein database
  • TBLASTN – protein query against translated nucleotide database
  • TBLASTX – translated nucleotide query against translated nucleotide database

⚡ 3. FASTA

  • Sequence similarity searching
  • Local and global sequence comparison approaches
  • Used for identifying similar sequences
  • Used in functional and evolutionary studies

๐Ÿงฌ 4. Sequence Identity

Sequence identity represents the percentage of positions at which two aligned sequences contain exactly the same residue or nucleotide.

Sequence Identity (%) = (Number of identical positions / Total aligned positions) × 100

๐Ÿ”— 5. Sequence Similarity

Sequence similarity includes identical residues as well as conservative substitutions with chemically or functionally similar properties.

๐Ÿงฌ 6. Homologues

Homologues are genes or proteins that share a common evolutionary origin.

๐ŸŒฟ 7. Orthologues

Orthologues are homologous genes separated by a speciation event.

๐Ÿงฌ 8. Paralougues

Paralogues are homologous genes produced through gene duplication.

Term Meaning
Homologue Common evolutionary origin
Orthologue Separated mainly by speciation
Paralogue Separated mainly by gene duplication

๐Ÿ” 9. Repeat Finding

  • Tandem repeats
  • Interspersed repeats
  • Short sequence repeats
  • Low-complexity regions
  • Transposable elements

๐Ÿ“Š 10. Scoring Matrices

  • Assign numerical scores to residue matches and mismatches
  • Used during sequence alignment
  • Reflect biological likelihood of substitutions
  • BLOSUM matrices
  • PAM matrices
๐Ÿ“Œ Important:

BLOSUM and PAM are commonly used substitution matrices for protein sequence alignment.

↔️ 11. Pairwise Sequence Alignment

  • Comparison of two sequences
  • Global alignment
  • Local alignment
  • Match
  • Mismatch
  • Gap
  • Gap penalties

๐ŸŒ 12. Multiple Sequence Alignment (MSA)

  • Alignment of three or more sequences
  • Identification of conserved regions
  • Identification of variable regions
  • Motif detection
  • Phylogenetic analysis
  • Protein family analysis
๐Ÿงฌ Applications of MSA:
  • Taxonomy
  • Phylogeny
  • Conserved motif identification
  • Functional prediction
  • Protein family classification
  • Comparative genomics

๐ŸŒณ 13. Phylogenetic Applications

  • Evolutionary relationships
  • Common ancestry
  • Sequence divergence
  • Species relationships
  • Gene family evolution

๐Ÿงฌ 14. Comparative Genomics

  • Comparison of genomes
  • Identification of conserved genes
  • Identification of species-specific genes
  • Genome rearrangements
  • Gene gain and loss
  • Evolutionary analysis

C. Gene Annotation

๐Ÿงฌ 1. Gene Annotation

Gene annotation is the process of identifying genes and assigning biological functions to them.

  • Identification of coding regions
  • Prediction of gene function
  • Functional classification
  • Identification of conserved domains
  • Identification of regulatory elements

๐Ÿ” 2. Homology-Based Function Prediction

  • Comparison with experimentally characterized genes
  • Identification of homologues
  • Functional transfer based on sequence similarity
  • Identification of conserved domains

๐Ÿงฉ 3. Context-Based Function Prediction

  • Gene neighborhood
  • Operon organization
  • Protein-protein interactions
  • Co-expression
  • Metabolic pathways

๐Ÿ—️ 4. Structure-Based Function Prediction

  • Protein structure comparison
  • Identification of structural domains
  • Active-site prediction
  • Ligand-binding site prediction

๐Ÿ•ธ️ 5. Network-Based Prediction

  • Protein interaction networks
  • Gene regulatory networks
  • Metabolic networks
  • Co-expression networks

๐Ÿงฌ 6. Genetic Variation

  • Single nucleotide polymorphisms (SNPs)
  • Insertions
  • Deletions
  • Copy-number variation
  • Structural variants

⚠️ 7. Deleterious Mutations

  • Mutations affecting protein function
  • Mutations affecting gene regulation
  • Loss-of-function mutations
  • Protein destabilizing mutations
  • Disease-associated variants

๐ŸŒณ 8. Computational Phylogenetics

  • Sequence alignment
  • Distance estimation
  • Phylogenetic tree construction
  • Tree interpretation
  • Evolutionary relationship analysis
๐Ÿงฌ Gene Annotation Workflow:

Genome Sequence → Gene Prediction → Sequence Similarity Search → Domain Identification → Structure Prediction → Functional Annotation → Pathway Assignment

D. Molecular Modelling and Dynamics

๐ŸงŠ 1. Three-Dimensional Structure Visualization

  • Protein structure visualization
  • DNA and RNA visualization
  • Ligand binding visualization
  • Protein-protein interaction visualization
  • Structural domain identification

๐Ÿ’ป 2. Molecular Modelling

Molecular modelling uses computational methods to predict, visualize and analyze molecular structures and interactions.

  • Structure prediction
  • Energy minimization
  • Conformational analysis
  • Molecular docking
  • Molecular dynamics

⚛️ 3. Molecular Mechanics

  • Atoms represented as particles
  • Bonds represented using force terms
  • Potential energy calculations
  • Bond stretching
  • Angle bending
  • Torsional rotation
  • Non-bonded interactions

๐Ÿงฎ 4. Force Fields

A force field is a mathematical representation used to calculate molecular potential energy.

  • Bond stretching terms
  • Angle bending terms
  • Dihedral terms
  • Van der Waals interactions
  • Electrostatic interactions
Total Potential Energy ≈ Bond Energy + Angle Energy + Dihedral Energy + Van der Waals Energy + Electrostatic Energy

๐ŸŒŠ 5. Molecular Dynamics

  • Simulation of atomic movement over time
  • Temperature control
  • Pressure control
  • Trajectory analysis
  • Conformational changes
  • Protein stability studies
  • Protein-ligand interaction studies
๐ŸงŠ Molecular Dynamics Workflow:

Structure → Parameterization → Energy Minimization → Equilibration → Production Simulation → Trajectory Analysis

E. Classification and Comparison of Protein 3D Structures

๐Ÿ—️ 1. Hierarchical Organization of Protein Structure

  • Primary structure
  • Secondary structure
  • Tertiary structure
  • Quaternary structure
๐Ÿ—️ Protein Structural Hierarchy:

Primary → Secondary → Tertiary → Quaternary

๐ŸŒ€ 2. Secondary Structure

  • ฮฑ-helix
  • ฮฒ-sheet
  • ฮฒ-turns
  • Loops

๐Ÿงฌ 3. Tertiary Structure

  • Three-dimensional folding of a polypeptide
  • Domain organization
  • Hydrophobic interactions
  • Hydrogen bonding
  • Electrostatic interactions
  • Disulfide bonds

๐Ÿ”ฎ 4. Secondary Structure Prediction

  • Prediction of ฮฑ-helices
  • Prediction of ฮฒ-sheets
  • Prediction of loops
  • Sequence-based prediction methods

๐Ÿงฌ 5. Homology / Comparative Modelling

Homology modelling predicts the structure of a target protein using a structurally characterized homologous protein as a template.

๐Ÿ”ฌ Basic Homology Modelling:

Target Sequence → Template Identification → Sequence Alignment → Model Building → Model Refinement → Model Validation

๐Ÿงฉ 6. Fold Recognition

  • Identification of structural folds
  • Comparison with known structural folds
  • Useful when sequence similarity is weak

๐Ÿงต 7. Threading

  • Target sequence fitted onto known structural folds
  • Structure-based sequence comparison
  • Useful for remote homology detection

๐Ÿ”ฌ 8. Ab Initio Structure Prediction

  • Prediction without requiring a closely related structural template
  • Uses physical and statistical principles
  • Searches conformational space

๐Ÿค– 9. AI-Based Structure Prediction

  • Machine-learning approaches
  • Deep-learning approaches
  • Prediction from amino-acid sequence
  • Protein structure prediction at large scale
  • Confidence estimation
๐Ÿค– AlphaFold:

AI-based protein structure prediction approach that uses deep-learning methods to predict protein three-dimensional structures from sequence information.

๐Ÿ“Š 10. Protein Structure Classification

  • Comparison of structural folds
  • Domain classification
  • Structural similarity
  • Evolutionary relationships
  • Functional relationships

F. Drug Design

๐Ÿ’Š 1. Chemical Databases

  • Small-molecule databases
  • Drug-like compound databases
  • Bioactive compound databases
  • Natural-product databases
  • Compound structure databases

๐Ÿงช 2. NCI and PubChem

  • Chemical structures
  • Compound identifiers
  • Molecular properties
  • Bioactivity information
  • Synonyms
  • Related chemical information

๐Ÿ”— 3. Receptor-Ligand Interactions

  • Ligand binding
  • Receptor binding sites
  • Hydrogen bonding
  • Hydrophobic interactions
  • Electrostatic interactions
  • Van der Waals interactions
  • Binding affinity

๐Ÿงฌ 4. Structure-Based Drug Design

Structure-based drug design uses the three-dimensional structure of a biological target to identify or optimize molecules capable of interacting with that target.

  • Target identification
  • Target structure determination
  • Binding-site identification
  • Virtual screening
  • Molecular docking
  • Lead optimization
๐Ÿ’Š Structure-Based Drug Design:

Target Structure → Binding Site → Virtual Screening → Docking → Hit Identification → Lead Optimization → Candidate Drug

๐Ÿงช 5. Ligand-Based Drug Design

Ligand-based drug design is used when information about active ligands is available but the detailed structure of the biological target may not be known.

  • Known active compounds
  • Structural similarity
  • Pharmacophore identification
  • QSAR analysis
  • Lead optimization

๐Ÿ“ˆ 6. Structure-Activity Relationship (SAR)

  • Relationship between chemical structure and biological activity
  • Identification of important functional groups
  • Lead optimization
  • Understanding activity changes caused by structural modifications

๐Ÿ“Š 7. QSAR

Quantitative Structure-Activity Relationship relates molecular descriptors to biological activity using mathematical or statistical models.

Molecular Structure → Molecular Descriptors → Mathematical Model → Predicted Biological Activity

๐Ÿงฉ 8. Pharmacophore

A pharmacophore represents the spatial arrangement of essential molecular features required for biological interaction with a target.

  • Hydrogen-bond donors
  • Hydrogen-bond acceptors
  • Hydrophobic regions
  • Aromatic regions
  • Charged groups

๐Ÿ’ป 9. In-Silico Drug Activity Prediction

  • Virtual screening
  • Molecular docking
  • QSAR prediction
  • Pharmacophore modelling
  • Machine-learning models
  • Binding-affinity prediction

๐Ÿงช 10. ADMET

Term Meaning
A – Absorption How a drug enters systemic circulation
D – Distribution Movement of drug throughout the body
M – Metabolism Biotransformation of the drug
E – Excretion Removal of drug or metabolites
T – Toxicity Potential harmful effects
๐Ÿ’Š ADMET is important for:
  • Drug safety
  • Pharmacokinetics
  • Drug efficacy
  • Lead optimization
  • Early elimination of unsuitable candidates

G. Systems Biology

๐Ÿงฌ 1. Concept of Systems Biology

Systems biology studies biological systems as interconnected networks rather than as isolated components.

  • Genes
  • Proteins
  • Metabolites
  • Cells
  • Pathways
  • Organ systems
๐Ÿงฌ Systems Biology Concept:

Genes ↔ Proteins ↔ Metabolites ↔ Pathways ↔ Cells ↔ Organisms

๐Ÿ“Š 2. Data Science in Biology

  • Large-scale biological data analysis
  • Machine learning
  • Statistical analysis
  • Data visualization
  • Pattern recognition
  • Predictive modelling

๐Ÿฅ 3. Data Science in Health

  • Clinical data analysis
  • Disease prediction
  • Biomarker discovery
  • Patient stratification
  • Medical image analysis
  • Risk prediction

๐Ÿ’Š 4. Data Science in Drug Discovery

  • Target identification
  • Virtual screening
  • Drug activity prediction
  • Drug repurposing
  • ADMET prediction
  • Lead optimization

๐Ÿ”ข 5. Mathematical Modelling of Metabolic Pathways

  • Metabolic networks
  • Reaction rates
  • Metabolite concentrations
  • Flux analysis
  • Pathway regulation
  • Dynamic models

๐Ÿฆ  6. Disease Modelling

  • Disease progression
  • Host-pathogen interactions
  • Gene regulatory networks
  • Signalling networks
  • Metabolic disorders
  • Therapeutic response prediction

๐Ÿ“ฑ 7. Digital Health

  • Electronic health records
  • Wearable devices
  • Remote monitoring
  • Digital biomarkers
  • Health applications
  • Artificial intelligence in healthcare

๐Ÿงฌ 8. Personalized Medicine

Personalized medicine uses individual biological and clinical information to guide prevention, diagnosis and treatment.

  • Genomic information
  • Individual genetic variation
  • Biomarkers
  • Pharmacogenomics
  • Patient-specific treatment
  • Precision medicine
๐ŸŽฏ Personalized Medicine:

Genomic Data + Clinical Data + Environmental Data + Lifestyle Data ↓ Individualized Diagnosis & Treatment

๐Ÿ“š Unit 11 – Bioinformatics & Computational Biology Quick Revision

A → Major Bioinformatic Resources
B → Sequence Analysis
C → Gene Annotation
D → Molecular Modelling & Dynamics
E → Protein 3D Structure
F → Drug Design
G → Systems Biology
๐ŸŽฏ High-Yield CSIR-NET Topics
  • Sequence databases
  • Gene expression databases
  • 3D structure databases
  • Pattern and motif databases
  • BLAST and FASTA
  • Sequence identity and similarity
  • Homologues, orthologues and paralogues
  • Repeat finding
  • BLOSUM and PAM matrices
  • Pairwise sequence alignment
  • Global and local alignment
  • Multiple sequence alignment
  • Phylogenetic applications
  • Comparative genomics
  • Gene annotation
  • Homology-based function prediction
  • Genetic variation and deleterious mutations
  • Phylogenetics
  • Molecular modelling
  • Molecular mechanics
  • Force fields
  • Molecular dynamics
  • Protein structural hierarchy
  • Secondary and tertiary structure prediction
  • Homology modelling
  • Fold recognition
  • Threading
  • Ab initio prediction
  • AI-based structure prediction
  • AlphaFold
  • Receptor-ligand interactions
  • Structure-based drug design
  • Ligand-based drug design
  • SAR
  • QSAR
  • Pharmacophores
  • Virtual screening
  • ADMET
  • Systems biology
  • Biological networks
  • Mathematical modelling
  • Data science in biology
  • Digital health
  • Personalized medicine
๐Ÿ”ฌ Sequence Analysis Workflow:

Sequence → Database Search → BLAST/FASTA → Similarity Analysis → Alignment → Conserved Motifs → Functional Prediction → Phylogenetic Analysis
๐Ÿ—️ Structural Bioinformatics Workflow:

Sequence → Structure Prediction → 3D Model → Structural Validation → Molecular Modelling → Molecular Dynamics → Functional Analysis
๐Ÿ’Š Computational Drug Discovery Workflow:

Target → Target Structure → Binding Site → Compound Library → Virtual Screening → Docking → SAR/QSAR → ADMET → Lead Optimization
๐Ÿงฌ Systems Biology Workflow:

Biological Data → Data Integration → Network Construction → Mathematical Modelling → Simulation → Prediction → Biological Interpretation
๐Ÿง  One-Line Unit 11 Revision:

Bioinformatics uses computational approaches to store, compare and interpret biological data, while computational biology applies modelling, structural analysis, data science and systems approaches to understand biological processes and develop applications in health and drug discovery.

๐Ÿงฌ CSIR-NET Life Sciences – Unit 11

Bioinformatics • Sequence Analysis • Structural Biology • Drug Design • Systems Biology

Study → Analyze → Practice → Revise → Master

No comments:

Post a Comment

Mock Test 5

Mock Test 5: System Physiology CSIR NET Part C Level | Comprehensive Animal Physiology | 30 Questions ...