Skip to content

BAMpiro logo

CI License: GPL v3 Nextflow Version Container PGO Docs Live QC report

General bacterial short-read mapping, variant calling & lineage/DR typing. Nextflow (DSL2) · containerized · self-contained interactive QC reports.

Docs site · Live QC report · Quick Start · Configuration · Citation

Paula Ruiz-Rodriguez1 and Mireia Coscolla1
1. I2SysBio, University of Valencia-CSIC, FISABIO Joint Research Unit Infection and Public Health, Valencia, Spain


What is BAMpiro?

Tip

New here? Start with the Quick Start - a run command, the samplesheet format, and where to find the report. Every parameter is listed in the configuration reference.

BAMpiro is a modular, containerized Nextflow (DSL2) pipeline that takes raw bacterial short reads all the way to annotated variants, consensus sequences, and a single interactive QC report. It is tuned by default for Mycobacterium tuberculosis but is organism-agnostic - point it at any reference genome + GFF.

Main features:

  • Reference-agnostic mapping, variant calling (FreeBayes), and consensus
  • Alignment-free MTBC lineage + WHO drug-resistance typing (Pathotypr)
  • A self-contained, interactive HTML QC report with 21 panels
  • Dual amino-acid numbering (used reference + H37Rv / Mycobrowser)
  • Indels and codon-level (MNV) amino-acid changes, phased on the reads
  • One digest-pinned container with every tool and marker panel built in

Features

Feature Description
🧬 Any bacterial genome Reference-agnostic mapping + variant calling; TB-tuned defaults, works on any species
🧹 Repeat & mappability masking nucmer repeat exclusion plus a length-aware genmap read filter
🧪 Variants & backbone FreeBayes (ploidy 1/2) + "all-sites" VCFs for phylogenetic supermatrices
🧩 Indels & codon-level changes Every indel with the call rules a SNP has (PASS or LowSupport) + an indel matrix; SNPs of one codon read whole on the reads that carry them (get_MNV)
🩺 Lineage & drug resistance Alignment-free MTBC lineage + WHO DR typing (Pathotypr), reference-agnostic
📊 Interactive QC report Self-contained HTML dashboard, 21 panels + per-sample qc_flags.tsv
🔤 Dual amino-acid numbering Protein changes in both the used reference and H37Rv/Mycobrowser numbering
📦 Containerized & reproducible A single image, pinned by digest, with every tool + bundled marker panels (details)
⚙️ Fully configurable Every step exposed as a Nextflow parameter (reference)

Installation

Requires Nextflow ≥ 24.04.2 and Docker or Singularity. The pipeline pulls the paururo/bampiro image, pinned by digest, with every tool built in - nothing else to install.

nextflow run main.nf --tsv samples.tsv --outdir results_bampiro -profile local,docker

--tsv is required, and -profile says where the work runs: local on this machine, slurm on any SLURM cluster, garnatxa on the I2SysBio one. With no -profile everything runs on the current host, which on a cluster login node means the login node - so choose one deliberately there. nextflow run main.nf --help lists every parameter with its default.

Full requirements and the bundled software versions: Installation.

Quick Start

BAMpiro assumes M. tuberculosis settings by default (ploidy = 2 for mixed infections). Lineage/DR typing is off by default - enable it (and dual amino-acid numbering) with:

nextflow run main.nf \
    --tsv samples.tsv --outdir results_bampiro -profile local,docker \
    --run_pathotypr true --annotate_canonical true

The samplesheet is a TSV (sampleId, r1, r2, refId, refFasta, refGff, …). See the Quick Start guide for the full format and multi-run merging, and Outputs for the result layout.

New to BAMpiro? Follow the hands-on Jupyter tutorial - install → samplesheet → run → explore the outputs with pandas / matplotlib on a bundled 17-sample example cohort (runnable without running the pipeline first).

Interactive QC Report

BAMpiro interactive QC report - click to open the live demo

▶ Open the live interactive demo report - a full 17-sample demo cohort, every panel populated, right in your browser

Every run writes a single self-contained <samplesheet>_qc_report.html (no internet, no CDN) that folds the whole cohort into one dashboard: live-adjustable QC thresholds, a dark / light theme, a collapsible sidebar, and 21 linked panels - general statistics, flagged samples, per-lineage summary (canonical mycolorsTB palette), QC-space PCA, genome landscape, SNP dynamics, epistasis, a full SNP matrix, and drug resistance - plus a machine-readable per-sample qc_flags.tsv. It is organism-agnostic, mobile-responsive, and works offline on an HPC login node.

Full panel list and interactive features: Interactive QC Report.

Documentation

📖 Read the full documentation online → pathogenomics-lab.github.io/BAMpiro: a searchable site with a pipeline diagram, a hover glossary, a step-by-step tutorial series, a hands-on Jupyter notebook, and the live QC-report demo.

Tip

New here? Follow the Tutorials: a guided path from your first run to a finished phylogeny.

Document Description
Tutorials A guided, hands-on path: first run → samplesheet → configuration → reading the report → typing → phylogeny
Introduction What BAMpiro is, key features, and the workflow at a glance
Installation Requirements, the container, and bundled software versions
Quick Start Run commands, the samplesheet format, and multi-run merging
Configuration The full parameter reference and feature toggles
Troubleshooting Common first-run errors and how to fix them
Interactive QC Report The self-contained HTML dashboard and its panels
Lineage & Drug-Resistance Typing Pathotypr typing and dual amino-acid numbering
Tutorial An end-to-end Jupyter walkthrough
Outputs The result file tree and the repository layout
Changelog Version history

Prefer the source? Browse the Markdown under docs/, or build the site locally with make docs-serve (needs MkDocs Material: pip install -r docs/requirements.txt).

Tests

tests/run_tests.sh

Unit tests for the Python under bin/, tests for the report front-end's hand-written statistics (checked against SciPy), and a full -stub-run of the DAG over a 170 kB fixture cohort, none of which needs a container, a reference genome or a network. One more leg runs the pipeline for real, in its containers, on a cohort simulated at test time whose every SNP, codon change and indel is known in advance (tests/run_tests.sh e2e, with Docker). tests/README.md has the details; CI runs all of them on every pull request.

Why "BAMpiro"?

The name is a play on words combining bioinformatics and folklore:

  • BAM - Binary Alignment Map, the standard format for reads aligned to a reference genome; the "heart" of this pipeline (mapping → variant calling).
  • Piro - combined with "BAM" it sounds like Vampiro (Spanish/Portuguese for vampire).

Just as a vampire seeks blood, BAMpiro seeks BAM files (and FASTQ data) to extract vital information - variants, lineages, and stats. A creature that lives in your cluster and processes bacterial genomes.

Citation

If you use BAMpiro in your research, please cite:

Ruiz-Rodriguez P, Coscollá M. BAMpiro: a Nextflow pipeline for bacterial short-read mapping, variant calling and lineage/drug-resistance typing. https://github.com/PathoGenOmics-Lab/BAMpiro

@software{ruiz-rodriguez_bampiro,
  title   = {BAMpiro: bacterial short-read mapping, variant calling and lineage/drug-resistance typing},
  author  = {Ruiz-Rodriguez, Paula and Coscoll{\'a}, Mireia},
  url      = {https://github.com/PathoGenOmics-Lab/BAMpiro},
  version = {1.1.0},
  license = {GPL-3.0}
}

License

GNU General Public License v3.0


BAMpiro is developed with ❤️ by:

Paula Ruiz-Rodriguez

💻 🔬 🤔 🔣 🎨 🔧

Mireia Coscolla

🔍 🤔 🧑‍🏫 🔬 📓

This project follows the all-contributors specification (emoji key).

About

General Bacterial Short Read Mapping, Variant Calling & Lineage/DR Typing Pipeline

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages