Unit 11: Bioinformatics and Computational Biology
Complete Topic-Wise Syllabus
Sequence Analysis • Structural Bioinformatics • Drug Design • Systems Biology
A. Major Bioinformatic Resources
๐งฌ 1. Sequence Databases
- Nucleotide sequence databases
- Protein sequence databases
- Genome databases
- Transcriptome databases
- Reference sequence databases
- Curated and non-curated databases
- Primary databases contain experimentally submitted sequence information.
- Secondary databases contain derived or curated information.
- Protein databases provide amino-acid sequence information and functional annotations.
- Nucleotide databases contain DNA and RNA sequence information.
๐งช 2. Gene Expression Databases
- Gene expression profiles
- Transcriptomic data
- RNA expression datasets
- Microarray data
- RNA-seq data
- Expression patterns under different conditions
- Tissue-specific expression
๐ง 3. 3D Structure Databases
- Experimental macromolecular structures
- Protein structures
- DNA structures
- RNA structures
- Protein-ligand complexes
- Structural coordinates
๐ 4. Pattern and Sequence Databases
- Protein motifs
- Functional domains
- Conserved sequence patterns
- Sequence signatures
- Repeats
- Profiles
๐️ 5. Database Classification
| Database Type | Main Information |
|---|---|
| Sequence Database | DNA, RNA and protein sequences |
| Expression Database | Gene expression and transcriptomic data |
| Structure Database | 3D structures of biomolecules |
| Pattern Database | Motifs, domains and sequence signatures |
B. Basic Concepts of Sequence Analysis
๐ 1. Database Searches
- Searching biological databases using nucleotide or protein sequences
- Identification of homologous sequences
- Functional annotation
- Identification of conserved regions
- Detection of evolutionary relationships
⚡ 2. BLAST
BLAST stands for Basic Local Alignment Search Tool. It identifies regions of local similarity between biological sequences.
- DNA sequence searches
- Protein sequence searches
- Identification of homologues
- Functional annotation
- Similarity searching
- BLASTN – nucleotide query against nucleotide database
- BLASTP – protein query against protein database
- BLASTX – translated nucleotide query against protein database
- TBLASTN – protein query against translated nucleotide database
- TBLASTX – translated nucleotide query against translated nucleotide database
⚡ 3. FASTA
- Sequence similarity searching
- Local and global sequence comparison approaches
- Used for identifying similar sequences
- Used in functional and evolutionary studies
๐งฌ 4. Sequence Identity
Sequence identity represents the percentage of positions at which two aligned sequences contain exactly the same residue or nucleotide.
๐ 5. Sequence Similarity
Sequence similarity includes identical residues as well as conservative substitutions with chemically or functionally similar properties.
๐งฌ 6. Homologues
Homologues are genes or proteins that share a common evolutionary origin.
๐ฟ 7. Orthologues
Orthologues are homologous genes separated by a speciation event.
๐งฌ 8. Paralougues
Paralogues are homologous genes produced through gene duplication.
| Term | Meaning |
|---|---|
| Homologue | Common evolutionary origin |
| Orthologue | Separated mainly by speciation |
| Paralogue | Separated mainly by gene duplication |
๐ 9. Repeat Finding
- Tandem repeats
- Interspersed repeats
- Short sequence repeats
- Low-complexity regions
- Transposable elements
๐ 10. Scoring Matrices
- Assign numerical scores to residue matches and mismatches
- Used during sequence alignment
- Reflect biological likelihood of substitutions
- BLOSUM matrices
- PAM matrices
BLOSUM and PAM are commonly used substitution matrices for protein sequence alignment.
↔️ 11. Pairwise Sequence Alignment
- Comparison of two sequences
- Global alignment
- Local alignment
- Match
- Mismatch
- Gap
- Gap penalties
๐ 12. Multiple Sequence Alignment (MSA)
- Alignment of three or more sequences
- Identification of conserved regions
- Identification of variable regions
- Motif detection
- Phylogenetic analysis
- Protein family analysis
- Taxonomy
- Phylogeny
- Conserved motif identification
- Functional prediction
- Protein family classification
- Comparative genomics
๐ณ 13. Phylogenetic Applications
- Evolutionary relationships
- Common ancestry
- Sequence divergence
- Species relationships
- Gene family evolution
๐งฌ 14. Comparative Genomics
- Comparison of genomes
- Identification of conserved genes
- Identification of species-specific genes
- Genome rearrangements
- Gene gain and loss
- Evolutionary analysis
C. Gene Annotation
๐งฌ 1. Gene Annotation
Gene annotation is the process of identifying genes and assigning biological functions to them.
- Identification of coding regions
- Prediction of gene function
- Functional classification
- Identification of conserved domains
- Identification of regulatory elements
๐ 2. Homology-Based Function Prediction
- Comparison with experimentally characterized genes
- Identification of homologues
- Functional transfer based on sequence similarity
- Identification of conserved domains
๐งฉ 3. Context-Based Function Prediction
- Gene neighborhood
- Operon organization
- Protein-protein interactions
- Co-expression
- Metabolic pathways
๐️ 4. Structure-Based Function Prediction
- Protein structure comparison
- Identification of structural domains
- Active-site prediction
- Ligand-binding site prediction
๐ธ️ 5. Network-Based Prediction
- Protein interaction networks
- Gene regulatory networks
- Metabolic networks
- Co-expression networks
๐งฌ 6. Genetic Variation
- Single nucleotide polymorphisms (SNPs)
- Insertions
- Deletions
- Copy-number variation
- Structural variants
⚠️ 7. Deleterious Mutations
- Mutations affecting protein function
- Mutations affecting gene regulation
- Loss-of-function mutations
- Protein destabilizing mutations
- Disease-associated variants
๐ณ 8. Computational Phylogenetics
- Sequence alignment
- Distance estimation
- Phylogenetic tree construction
- Tree interpretation
- Evolutionary relationship analysis
Genome Sequence → Gene Prediction → Sequence Similarity Search → Domain Identification → Structure Prediction → Functional Annotation → Pathway Assignment
D. Molecular Modelling and Dynamics
๐ง 1. Three-Dimensional Structure Visualization
- Protein structure visualization
- DNA and RNA visualization
- Ligand binding visualization
- Protein-protein interaction visualization
- Structural domain identification
๐ป 2. Molecular Modelling
Molecular modelling uses computational methods to predict, visualize and analyze molecular structures and interactions.
- Structure prediction
- Energy minimization
- Conformational analysis
- Molecular docking
- Molecular dynamics
⚛️ 3. Molecular Mechanics
- Atoms represented as particles
- Bonds represented using force terms
- Potential energy calculations
- Bond stretching
- Angle bending
- Torsional rotation
- Non-bonded interactions
๐งฎ 4. Force Fields
A force field is a mathematical representation used to calculate molecular potential energy.
- Bond stretching terms
- Angle bending terms
- Dihedral terms
- Van der Waals interactions
- Electrostatic interactions
๐ 5. Molecular Dynamics
- Simulation of atomic movement over time
- Temperature control
- Pressure control
- Trajectory analysis
- Conformational changes
- Protein stability studies
- Protein-ligand interaction studies
Structure → Parameterization → Energy Minimization → Equilibration → Production Simulation → Trajectory Analysis
E. Classification and Comparison of Protein 3D Structures
๐️ 1. Hierarchical Organization of Protein Structure
- Primary structure
- Secondary structure
- Tertiary structure
- Quaternary structure
Primary → Secondary → Tertiary → Quaternary
๐ 2. Secondary Structure
- ฮฑ-helix
- ฮฒ-sheet
- ฮฒ-turns
- Loops
๐งฌ 3. Tertiary Structure
- Three-dimensional folding of a polypeptide
- Domain organization
- Hydrophobic interactions
- Hydrogen bonding
- Electrostatic interactions
- Disulfide bonds
๐ฎ 4. Secondary Structure Prediction
- Prediction of ฮฑ-helices
- Prediction of ฮฒ-sheets
- Prediction of loops
- Sequence-based prediction methods
๐งฌ 5. Homology / Comparative Modelling
Homology modelling predicts the structure of a target protein using a structurally characterized homologous protein as a template.
Target Sequence → Template Identification → Sequence Alignment → Model Building → Model Refinement → Model Validation
๐งฉ 6. Fold Recognition
- Identification of structural folds
- Comparison with known structural folds
- Useful when sequence similarity is weak
๐งต 7. Threading
- Target sequence fitted onto known structural folds
- Structure-based sequence comparison
- Useful for remote homology detection
๐ฌ 8. Ab Initio Structure Prediction
- Prediction without requiring a closely related structural template
- Uses physical and statistical principles
- Searches conformational space
๐ค 9. AI-Based Structure Prediction
- Machine-learning approaches
- Deep-learning approaches
- Prediction from amino-acid sequence
- Protein structure prediction at large scale
- Confidence estimation
AI-based protein structure prediction approach that uses deep-learning methods to predict protein three-dimensional structures from sequence information.
๐ 10. Protein Structure Classification
- Comparison of structural folds
- Domain classification
- Structural similarity
- Evolutionary relationships
- Functional relationships
F. Drug Design
๐ 1. Chemical Databases
- Small-molecule databases
- Drug-like compound databases
- Bioactive compound databases
- Natural-product databases
- Compound structure databases
๐งช 2. NCI and PubChem
- Chemical structures
- Compound identifiers
- Molecular properties
- Bioactivity information
- Synonyms
- Related chemical information
๐ 3. Receptor-Ligand Interactions
- Ligand binding
- Receptor binding sites
- Hydrogen bonding
- Hydrophobic interactions
- Electrostatic interactions
- Van der Waals interactions
- Binding affinity
๐งฌ 4. Structure-Based Drug Design
Structure-based drug design uses the three-dimensional structure of a biological target to identify or optimize molecules capable of interacting with that target.
- Target identification
- Target structure determination
- Binding-site identification
- Virtual screening
- Molecular docking
- Lead optimization
Target Structure → Binding Site → Virtual Screening → Docking → Hit Identification → Lead Optimization → Candidate Drug
๐งช 5. Ligand-Based Drug Design
Ligand-based drug design is used when information about active ligands is available but the detailed structure of the biological target may not be known.
- Known active compounds
- Structural similarity
- Pharmacophore identification
- QSAR analysis
- Lead optimization
๐ 6. Structure-Activity Relationship (SAR)
- Relationship between chemical structure and biological activity
- Identification of important functional groups
- Lead optimization
- Understanding activity changes caused by structural modifications
๐ 7. QSAR
Quantitative Structure-Activity Relationship relates molecular descriptors to biological activity using mathematical or statistical models.
๐งฉ 8. Pharmacophore
A pharmacophore represents the spatial arrangement of essential molecular features required for biological interaction with a target.
- Hydrogen-bond donors
- Hydrogen-bond acceptors
- Hydrophobic regions
- Aromatic regions
- Charged groups
๐ป 9. In-Silico Drug Activity Prediction
- Virtual screening
- Molecular docking
- QSAR prediction
- Pharmacophore modelling
- Machine-learning models
- Binding-affinity prediction
๐งช 10. ADMET
| Term | Meaning |
|---|---|
| A – Absorption | How a drug enters systemic circulation |
| D – Distribution | Movement of drug throughout the body |
| M – Metabolism | Biotransformation of the drug |
| E – Excretion | Removal of drug or metabolites |
| T – Toxicity | Potential harmful effects |
- Drug safety
- Pharmacokinetics
- Drug efficacy
- Lead optimization
- Early elimination of unsuitable candidates
G. Systems Biology
๐งฌ 1. Concept of Systems Biology
Systems biology studies biological systems as interconnected networks rather than as isolated components.
- Genes
- Proteins
- Metabolites
- Cells
- Pathways
- Organ systems
Genes ↔ Proteins ↔ Metabolites ↔ Pathways ↔ Cells ↔ Organisms
๐ 2. Data Science in Biology
- Large-scale biological data analysis
- Machine learning
- Statistical analysis
- Data visualization
- Pattern recognition
- Predictive modelling
๐ฅ 3. Data Science in Health
- Clinical data analysis
- Disease prediction
- Biomarker discovery
- Patient stratification
- Medical image analysis
- Risk prediction
๐ 4. Data Science in Drug Discovery
- Target identification
- Virtual screening
- Drug activity prediction
- Drug repurposing
- ADMET prediction
- Lead optimization
๐ข 5. Mathematical Modelling of Metabolic Pathways
- Metabolic networks
- Reaction rates
- Metabolite concentrations
- Flux analysis
- Pathway regulation
- Dynamic models
๐ฆ 6. Disease Modelling
- Disease progression
- Host-pathogen interactions
- Gene regulatory networks
- Signalling networks
- Metabolic disorders
- Therapeutic response prediction
๐ฑ 7. Digital Health
- Electronic health records
- Wearable devices
- Remote monitoring
- Digital biomarkers
- Health applications
- Artificial intelligence in healthcare
๐งฌ 8. Personalized Medicine
Personalized medicine uses individual biological and clinical information to guide prevention, diagnosis and treatment.
- Genomic information
- Individual genetic variation
- Biomarkers
- Pharmacogenomics
- Patient-specific treatment
- Precision medicine
Genomic Data + Clinical Data + Environmental Data + Lifestyle Data ↓ Individualized Diagnosis & Treatment
๐ Unit 11 – Bioinformatics & Computational Biology Quick Revision
- Sequence databases
- Gene expression databases
- 3D structure databases
- Pattern and motif databases
- BLAST and FASTA
- Sequence identity and similarity
- Homologues, orthologues and paralogues
- Repeat finding
- BLOSUM and PAM matrices
- Pairwise sequence alignment
- Global and local alignment
- Multiple sequence alignment
- Phylogenetic applications
- Comparative genomics
- Gene annotation
- Homology-based function prediction
- Genetic variation and deleterious mutations
- Phylogenetics
- Molecular modelling
- Molecular mechanics
- Force fields
- Molecular dynamics
- Protein structural hierarchy
- Secondary and tertiary structure prediction
- Homology modelling
- Fold recognition
- Threading
- Ab initio prediction
- AI-based structure prediction
- AlphaFold
- Receptor-ligand interactions
- Structure-based drug design
- Ligand-based drug design
- SAR
- QSAR
- Pharmacophores
- Virtual screening
- ADMET
- Systems biology
- Biological networks
- Mathematical modelling
- Data science in biology
- Digital health
- Personalized medicine
Sequence → Database Search → BLAST/FASTA → Similarity Analysis → Alignment → Conserved Motifs → Functional Prediction → Phylogenetic Analysis
Sequence → Structure Prediction → 3D Model → Structural Validation → Molecular Modelling → Molecular Dynamics → Functional Analysis
Target → Target Structure → Binding Site → Compound Library → Virtual Screening → Docking → SAR/QSAR → ADMET → Lead Optimization
Biological Data → Data Integration → Network Construction → Mathematical Modelling → Simulation → Prediction → Biological Interpretation
Bioinformatics uses computational approaches to store, compare and interpret biological data, while computational biology applies modelling, structural analysis, data science and systems approaches to understand biological processes and develop applications in health and drug discovery.
๐งฌ CSIR-NET Life Sciences – Unit 11
Bioinformatics • Sequence Analysis • Structural Biology • Drug Design • Systems Biology
Study → Analyze → Practice → Revise → Master
No comments:
Post a Comment