Skip to content

About

A consolidated workspace for structural bioinformatics and molecular modeling. Features AlphaFold2 3D kinase structural modeling (AKT1 for TNBC) and STN7/STN8 thylakoid kinase molecular docking (AutoDock Vina) pipelines.

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Latest commit

 

History

17,219 Commits

Folders and files

Repository files navigation

📚 Bioinformatics Education & Research Methodology Hub

Python R Nanopore Cambridge ML License: MIT

📌 Overview

This repository serves as a centralized knowledge hub for bioinformatics education, sequencing methodology, and advanced mathematical foundations for machine learning in biology. It integrates structured Biopython training, Oxford Nanopore sequencing protocols, and the official Cambridge University Mathematics of Machine Learning curriculum.


🗂️ Repository Structure

Designed for clean navigation and educational clarity:

Bioinformatics/
├── .gitignore                  # Git untracked pattern file
├── LICENSE
├── README.md
│
├── courses/                    # Structured learning curricula
│   ├── biopython/              # 30-module Biopython sequence and structural training
│   ├── cambridge_ml_math/      # Cambridge University Mathematics of Machine Learning course
│   └── wcu_comp_biomath/       # Computational biomathematics slides and labs
│
├── protocols/                  # Laboratory sequencing protocols
│   └── nanopore/               # Oxford Nanopore (ONT) metagenomics sequencing docs
│
├── literature/                 # Textbooks and reference publications
│   └── Bioinformatics for Beginners.pdf
│
├── starter_kit/                # Reusable bioinformatics templates
│   ├── genomic_analyzer.py     # DNA sequence quality & GC skew calculators
│   └── protein_3d_metrics.py   # AlphaFold coordinate & PDB backbone geometry parser
│
├── journal/                    # Academic logging
│   └── Bioinformatics_Research_Journal.md # Log of algorithms, models, and notes
│
└── career/                     # Academic and professional planning
    ├── Advice_for_Bioinformatics_AZE/ # Strategic career roadmaps
    └── IELTS_preparation_plan.md      # Language proficiency preparation goals

🔬 Core Learning Curricula

1. 🧬 Biopython Training (30 Modules)

A systematic, 30-lesson roadmap for mastering programmatic biology:

  • Modules 1–10: Sequence modeling (SeqIO, transcription, translation, FASTA/GenBank manipulation).
  • Modules 11–20: Querying biological databases (NCBI Entrez, AlignIO, Multiple Sequence Alignment).
  • Modules 21–30: Structural biology coordinate parsing (PDB and CIF), metagenomics, and advanced pipelines.

2. 🎓 Cambridge Mathematics of Machine Learning

Official lecture notes, problem sets, and textbooks from the University of Cambridge's Mathematics of Machine Learning course.

  • Covers the mathematical backbone (optimization, probability, learning theory) necessary for designing advanced genomic classifiers and models.

3. 🦠 Oxford Nanopore Metagenomic Protocols

Practical know-how and protocols for rapid pathogen surveillance (SQK-RPB114.24 kit) using Oxford Nanopore MinION/GridION platforms.


🔬 Mathematical & Algorithmic Foundations

To prepare for Cambridge University's mathematical machine learning standards, our biophysical and genomic analysis modules implement algorithms from their statistical definitions:

1. Genomic GC-Skew Windowing

To locate replication origins ($ori$) and termini in bacterial and viral genomes, we compute GC-skew over a sliding window: $$S_{GC} = \frac{C - G}{C + G}$$ where $C$ and $G$ represent the frequencies of cytosine and guanine bases within the window, respectively.


2. Kyte-Doolittle Hydrophobic Profile Windowing

For a polypeptide sequence of length $L$ and window size $w$ (typically 9 or 19 residues), the local hydrophobicity score at position $i$ is calculated as: $$H_i = \frac{1}{2w + 1} \sum_{j=-w}^{w} h_{i+j}$$ where $h_k$ is the hydropathy index of the residue at position $k$ according to the Kyte-Doolittle scale.


3. Miyazawa-Jernigan Statistical Contact Potentials

The non-covalent folding free energy ($E$) of a protein model is estimated using a simplified residue-contact energy grid: $$E = \sum_{i < j} e(R_i, R_j) \cdot \mathbb{I}(d(C_{\alpha,i}, C_{\alpha,j}) \leq d_{\text{cutoff}})$$ where:

  • $e(R_i, R_j)$ is the statistical contact energy between residue types $R_i$ and $R_j$ defined by the Miyazawa-Jernigan matrix.
  • $d(C_{\alpha,i}, C_{\alpha,j})$ is the 3D Euclidean distance between the Carbon-Alpha atoms of residues $i$ and $j$.
  • $\mathbb{I}(\cdot)$ is the indicator function flagging a physical contact (typically $d_{\text{cutoff}} = 6.5\text{ Å}$ or $8.0\text{ Å}$).

🔗 Related Portfolios

This hub links directly to other specialized codebases:

Scale Repository Focus
Transcriptomic & Clinical ML Bioinformatics-analysis RNA-seq · TNBC · Clinical genomics pipelines
Structural Biology (3D Protein AI) computational-structural-biology AKT1/STN7 kinase modeling & molecular docking
Phylogenomics & Evolution MEGA-Software-Genetics COL1A1 multi-species sequence alignment

Author: Suleiman Hajizadeh | Bioinformatician @ IMBB, Azerbaijan
📧 suleyman.hacizade1@gmail.com | 🔗 GitHub Portfolio

About

A consolidated workspace for structural bioinformatics and molecular modeling. Features AlphaFold2 3D kinase structural modeling (AKT1 for TNBC) and STN7/STN8 thylakoid kinase molecular docking (AutoDock Vina) pipelines.

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages