Phase 3 — Advanced & Specialised
Module 23

Capstone — Plant Genomics Pipeline

Tie everything from Phases 1–3 together. One polished end-to-end pipeline on public plant data — Snakemake, DESeq2, Docker, Jupyter notebook, full README with workflow diagram. Tag a v1.0 release on GitHub.

Weeks 29–30Timeline
~21 hrsStudy time
FREEAlways
What you'll learn

Topics covered in this module

Snakemake pipeline
DESeq2 report
Docker container
Jupyter notebook
v1.0 GitHub release
STAR · HISAT2 alignment
GATK variant calling
SnpEff annotation
Curriculum

10 lessons in this module

Repo: plant-genomics-pipeline · Weeks 29–30

1Capstone Overview & Project DesignPipeline goals, directory structure, tools review, data strategyLive
2Data Acquisition & Quality ControlDownloading Sorghum FASTQ data, FastQC, MultiQC reportSoon
3Read Trimming & Pre-alignment QCTrimmomatic/Fastp, adapter removal, quality filteringSoon
4Reference Genome SetupDownloading S. bicolor genome, indexing with STAR/HISAT2Soon
5Read Alignment & BAM ProcessingSTAR alignment, samtools sort/index, flagstatSoon
6Variant Calling with GATKHaplotypeCaller, GVCF mode, joint genotypingSoon
7Variant Filtering & AnnotationGATK VQSR/hard filters, SnpEff annotationSoon
8Differential Expression AnalysisfeatureCounts → DESeq2 full workflowSoon
9Integrating Results & VisualisationVolcano plots, Manhattan plots, combining DE + variant dataSoon
10Snakemake Automation & Final ReportFull Snakemake pipeline, HTML report, GitHub releaseSoon

📋 Not sure where this fits? Module 23 is part of the full bioinformatics curriculum — a structured 42-week learning path from Bash to single-cell RNA-seq.

See full curriculum →

Lesson 1 is live. The rest are on their way.

We're building this module carefully — every lesson comes with exercises, datasets, and a GitHub repository you push as proof of skill. Sign up to be notified the moment it goes live.

23
Module 23 of 31