This repository serves as a centralized knowledge hub for bioinformatics education, sequencing methodology, and advanced mathematical foundations for machine learning in biology. It integrates structured Biopython training, Oxford Nanopore sequencing protocols, and the official Cambridge University Mathematics of Machine Learning curriculum.
Designed for clean navigation and educational clarity:
Bioinformatics/
├── .gitignore # Git untracked pattern file
├── LICENSE
├── README.md
│
├── courses/ # Structured learning curricula
│ ├── biopython/ # 30-module Biopython sequence and structural training
│ ├── cambridge_ml_math/ # Cambridge University Mathematics of Machine Learning course
│ └── wcu_comp_biomath/ # Computational biomathematics slides and labs
│
├── protocols/ # Laboratory sequencing protocols
│ └── nanopore/ # Oxford Nanopore (ONT) metagenomics sequencing docs
│
├── literature/ # Textbooks and reference publications
│ └── Bioinformatics for Beginners.pdf
│
├── starter_kit/ # Reusable bioinformatics templates
│ ├── genomic_analyzer.py # DNA sequence quality & GC skew calculators
│ └── protein_3d_metrics.py # AlphaFold coordinate & PDB backbone geometry parser
│
├── journal/ # Academic logging
│ └── Bioinformatics_Research_Journal.md # Log of algorithms, models, and notes
│
└── career/ # Academic and professional planning
├── Advice_for_Bioinformatics_AZE/ # Strategic career roadmaps
└── IELTS_preparation_plan.md # Language proficiency preparation goals
A systematic, 30-lesson roadmap for mastering programmatic biology:
- Modules 1–10: Sequence modeling (
SeqIO, transcription, translation, FASTA/GenBank manipulation). - Modules 11–20: Querying biological databases (
NCBI Entrez,AlignIO, Multiple Sequence Alignment). - Modules 21–30: Structural biology coordinate parsing (
PDBandCIF), metagenomics, and advanced pipelines.
Official lecture notes, problem sets, and textbooks from the University of Cambridge's Mathematics of Machine Learning course.
- Covers the mathematical backbone (optimization, probability, learning theory) necessary for designing advanced genomic classifiers and models.
Practical know-how and protocols for rapid pathogen surveillance (SQK-RPB114.24 kit) using Oxford Nanopore MinION/GridION platforms.
To prepare for Cambridge University's mathematical machine learning standards, our biophysical and genomic analysis modules implement algorithms from their statistical definitions:
To locate replication origins (
For a polypeptide sequence of length
The non-covalent folding free energy (
-
$e(R_i, R_j)$ is the statistical contact energy between residue types$R_i$ and$R_j$ defined by the Miyazawa-Jernigan matrix. -
$d(C_{\alpha,i}, C_{\alpha,j})$ is the 3D Euclidean distance between the Carbon-Alpha atoms of residues$i$ and$j$ . -
$\mathbb{I}(\cdot)$ is the indicator function flagging a physical contact (typically$d_{\text{cutoff}} = 6.5\text{ Å}$ or$8.0\text{ Å}$ ).
This hub links directly to other specialized codebases:
| Scale | Repository | Focus |
|---|---|---|
| Transcriptomic & Clinical ML | Bioinformatics-analysis | RNA-seq · TNBC · Clinical genomics pipelines |
| Structural Biology (3D Protein AI) | computational-structural-biology | AKT1/STN7 kinase modeling & molecular docking |
| Phylogenomics & Evolution | MEGA-Software-Genetics | COL1A1 multi-species sequence alignment |
Author: Suleiman Hajizadeh | Bioinformatician @ IMBB, Azerbaijan
📧 suleyman.hacizade1@gmail.com | 🔗 GitHub Portfolio