General bacterial short-read mapping, variant calling & lineage/DR typing. Nextflow (DSL2) · containerized · self-contained interactive QC reports.
Docs site · Live QC report · Quick Start · Configuration · Citation
Paula Ruiz-Rodriguez1
and Mireia Coscolla1
1. I2SysBio, University of Valencia-CSIC, FISABIO Joint Research Unit Infection and Public Health, Valencia, Spain
Tip
New here? Start with the Quick Start - a run command, the samplesheet format, and where to find the report. Every parameter is listed in the configuration reference.
BAMpiro is a modular, containerized Nextflow (DSL2) pipeline that takes raw bacterial short reads all the way to annotated variants, consensus sequences, and a single interactive QC report. It is tuned by default for Mycobacterium tuberculosis but is organism-agnostic - point it at any reference genome + GFF.
Main features:
- Reference-agnostic mapping, variant calling (
FreeBayes), and consensus - Alignment-free MTBC lineage + WHO drug-resistance typing (Pathotypr)
- A self-contained, interactive HTML QC report with 21 panels
- Dual amino-acid numbering (used reference + H37Rv / Mycobrowser)
- Indels and codon-level (MNV) amino-acid changes, phased on the reads
- One digest-pinned container with every tool and marker panel built in
| Feature | Description |
|---|---|
| 🧬 Any bacterial genome | Reference-agnostic mapping + variant calling; TB-tuned defaults, works on any species |
| 🧹 Repeat & mappability masking | nucmer repeat exclusion plus a length-aware genmap read filter |
| 🧪 Variants & backbone | FreeBayes (ploidy 1/2) + "all-sites" VCFs for phylogenetic supermatrices |
| 🧩 Indels & codon-level changes | Every indel with the call rules a SNP has (PASS or LowSupport) + an indel matrix; SNPs of one codon read whole on the reads that carry them (get_MNV) |
| 🩺 Lineage & drug resistance | Alignment-free MTBC lineage + WHO DR typing (Pathotypr), reference-agnostic |
| 📊 Interactive QC report | Self-contained HTML dashboard, 21 panels + per-sample qc_flags.tsv |
| 🔤 Dual amino-acid numbering | Protein changes in both the used reference and H37Rv/Mycobrowser numbering |
| 📦 Containerized & reproducible | A single image, pinned by digest, with every tool + bundled marker panels (details) |
| ⚙️ Fully configurable | Every step exposed as a Nextflow parameter (reference) |
Requires Nextflow ≥ 24.04.2 and Docker or Singularity. The pipeline
pulls the paururo/bampiro image, pinned by digest, with every tool built in - nothing
else to install.
nextflow run main.nf --tsv samples.tsv --outdir results_bampiro -profile local,docker--tsv is required, and -profile says where the work runs: local on this machine,
slurm on any SLURM cluster, garnatxa on the I2SysBio one. With no -profile everything
runs on the current host, which on a cluster login node means the login node - so choose
one deliberately there. nextflow run main.nf --help lists every parameter with its
default.
Full requirements and the bundled software versions: Installation.
BAMpiro assumes M. tuberculosis settings by default (ploidy = 2 for mixed infections). Lineage/DR typing is off by default - enable it (and dual amino-acid numbering) with:
nextflow run main.nf \
--tsv samples.tsv --outdir results_bampiro -profile local,docker \
--run_pathotypr true --annotate_canonical trueThe samplesheet is a TSV (sampleId, r1, r2, refId, refFasta, refGff, …).
See the Quick Start guide for the full format and multi-run
merging, and Outputs for the result layout.
New to BAMpiro? Follow the hands-on
Jupyter tutorial - install → samplesheet →
run → explore the outputs with pandas / matplotlib on a bundled 17-sample example
cohort (runnable without running the pipeline first).
▶ Open the live interactive demo report - a full 17-sample demo cohort, every panel populated, right in your browser
Every run writes a single self-contained <samplesheet>_qc_report.html (no internet,
no CDN) that folds the whole cohort into one dashboard: live-adjustable QC
thresholds, a dark / light theme, a collapsible sidebar, and 21 linked
panels - general statistics, flagged samples, per-lineage summary (canonical
mycolorsTB palette), QC-space PCA, genome landscape, SNP dynamics, epistasis, a
full SNP matrix, and drug resistance - plus a machine-readable per-sample
qc_flags.tsv. It is organism-agnostic, mobile-responsive, and works offline on an
HPC login node.
Full panel list and interactive features: Interactive QC Report.
📖 Read the full documentation online → pathogenomics-lab.github.io/BAMpiro: a searchable site with a pipeline diagram, a hover glossary, a step-by-step tutorial series, a hands-on Jupyter notebook, and the live QC-report demo.
Tip
New here? Follow the Tutorials: a guided path from your first run to a finished phylogeny.
| Document | Description |
|---|---|
| Tutorials | A guided, hands-on path: first run → samplesheet → configuration → reading the report → typing → phylogeny |
| Introduction | What BAMpiro is, key features, and the workflow at a glance |
| Installation | Requirements, the container, and bundled software versions |
| Quick Start | Run commands, the samplesheet format, and multi-run merging |
| Configuration | The full parameter reference and feature toggles |
| Troubleshooting | Common first-run errors and how to fix them |
| Interactive QC Report | The self-contained HTML dashboard and its panels |
| Lineage & Drug-Resistance Typing | Pathotypr typing and dual amino-acid numbering |
| Tutorial | An end-to-end Jupyter walkthrough |
| Outputs | The result file tree and the repository layout |
| Changelog | Version history |
Prefer the source? Browse the Markdown under docs/, or build the site
locally with make docs-serve
(needs MkDocs Material:
pip install -r docs/requirements.txt).
tests/run_tests.shUnit tests for the Python under bin/, tests for the report front-end's hand-written statistics
(checked against SciPy), and a full -stub-run of the DAG over a 170 kB fixture cohort, none of
which needs a container, a reference genome or a network. One more leg runs the pipeline for real,
in its containers, on a cohort simulated at test time whose every SNP, codon change and indel is
known in advance (tests/run_tests.sh e2e, with Docker). tests/README.md has
the details; CI runs all of them on every pull request.
The name is a play on words combining bioinformatics and folklore:
- BAM - Binary Alignment Map, the standard format for reads aligned to a reference genome; the "heart" of this pipeline (mapping → variant calling).
- Piro - combined with "BAM" it sounds like Vampiro (Spanish/Portuguese for vampire).
Just as a vampire seeks blood, BAMpiro seeks BAM files (and FASTQ data) to extract vital information - variants, lineages, and stats. A creature that lives in your cluster and processes bacterial genomes.
If you use BAMpiro in your research, please cite:
Ruiz-Rodriguez P, Coscollá M. BAMpiro: a Nextflow pipeline for bacterial short-read mapping, variant calling and lineage/drug-resistance typing. https://github.com/PathoGenOmics-Lab/BAMpiro
@software{ruiz-rodriguez_bampiro,
title = {BAMpiro: bacterial short-read mapping, variant calling and lineage/drug-resistance typing},
author = {Ruiz-Rodriguez, Paula and Coscoll{\'a}, Mireia},
url = {https://github.com/PathoGenOmics-Lab/BAMpiro},
version = {1.1.0},
license = {GPL-3.0}
}GNU General Public License v3.0
|
Paula Ruiz-Rodriguez 💻 🔬 🤔 🔣 🎨 🔧 |
Mireia Coscolla 🔍 🤔 🧑🏫 🔬 📓 |
This project follows the all-contributors specification (emoji key).