Skip to content
#gene editing Open access

aomlomics/tourmaline: v2.1.0

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Tourmaline v2.1.0 This release adds two new taxonomy features — the REVAMP classify method by Sean McAllister and optional Krona plots — and has a major overhaul of the documentation. Existing configuration files continue to work unchanged. No parameters were renamed or removed. Still requires QIIME 2 2024.10 in an environment named qiime2-amplicon-2024.10. New features REVAMP taxonomy assignment classify_method: revamp assigns taxonomy with REVAMP: BLASTn against a local NCBI nt database, then a lowest common ancestor over all best hits, bounded by per-rank percent identity cutoffs. This is the first classify method that does not use a curated reference database. Instead of refseqs_file / taxa_file, it searches nt directly with NCBI taxonomy as its backbone, which is useful for markers with no well-curated reference set, or as a cross-check on a curated one. classify_method: revamp revamp_dir: /path/to/REVAMP # a REVAMP clone revamp_blastdb: /path/to/blastdb # nt volumes plus a prepared taxdump/ revamp_blast_results: # optional: a BLASTn btab produced elsewhere revamp_blast_mode: mostEnvOUT # allIN | allEnvOUT | mostEnvOUT revamp_query_cov: 90 # percent of ASV length a hit must cover revamp_taxonomy_cutoffs: "97,95,90,80,70,60" # percent ID cutoffs, ordered S,G,F,O,C,P Notes for users: BLAST can run on another machine. nt is large and often lives elsewhere. Run BLASTn there against the ASVs Tourmaline exports, then point revamp_blast_results at the result; only the taxdump/ files are then needed locally. The exact command and required column order are in the taxonomy docs. Cutoffs are marker-dependent. 97,95,90,80,70,60 suits rRNA genes; 95,92,87,77,67,60 is a better starting point for protein-coding genes. The taxonomy TSV's third column is percent_id (0–100), not Confidence (0–1). REVAMP produces no confidence score. This column carries through to -asv_taxa_features.tsv, so downstream code that assumes a 0–1 confidence should be checked. Tourmaline never modifies the REVAMP clone, it is maintained separately. Requires a revamp conda environment (see install docs) and a taxdump/ prepared by REVAMP's ncbi_db_cleanup.sh. Optional Krona plots make_krona: True produces an interactive Krona chart at figures/{run_name}-krona.html, for any classify method. make_krona: False # opt-in krona_per_sample: True # one dataset per sample alongside the all-samples plot The plot always includes an all_samples dataset summed across the run. With krona_per_sample: True, each sample is added as its own selectable dataset — set it to False for runs with many samples. Rank prefixes are stripped, trailing empty/NA ranks dropped, and unassigned features grouped under Unassigned. The output is a self-contained HTML file, so open it directly in a browser rather than through qiime tools view. Requires a krona conda environment; no Krona taxonomy database download is needed. Improvements Taxonomy assignment rules extracted into rules/taxonomy_assignment.smk, shared by all five classify methods. This is an internal reorganization — behavior is unchanged — but anyone maintaining a fork or custom rules should expect the moved code. has_fa_suffix() now handles None and empty strings (#184, thanks @clementcoclet), so an optional config key left blank no longer raises an error. Empty strings are how the example configs express an unset key, making this the more common case of the two. Bug fixes BLCA scores were wrong when reference ranks were empty or NA. Scoring now stays in alias space and resolves names only for the taxonomy lookup. Previously a query could collapse into a reference record carrying the same name. Affects classify_method: bt2-blca; if you have BLCA results from an earlier version on a reference database with taxonomy strings containing NAs, they are worth regenerating. DADA2 rules now run with a clean R environment (R_LIBS_USER, R_LIBS, R_PROFILE_USER, R_ENVIRON_USER unset), so a user-level R library no longer interferes with denoising. Documentation The README and the docs/ site were rewritten. An intro for people new to amplicon analysis — a pipeline diagram, a vocabulary table (ASV, denoising, feature table, .qza), what you need before starting, and first-run advice. docs/configuration.md is now the complete parameter reference, and the README links to it rather than duplicating it. The previous duplication is what allowed the two to drift apart. New and expanded pages: per-step guides for QA/QC, repseqs and taxonomy; external data; running, including parameter sweeps and HPC; and a troubleshooting guide covering real failure modes. Upgrade notes Existing config files work unchanged. config_01_qaqc.yaml and config_02_repseqs.yaml have no key changes at all. config_03_taxonomy.yaml gains eight optional keys — make_krona, krona_per_sample, and six revamp_* keys — all of which default sensibly when absent. If you are not using REVAMP or Krona, you can reuse config files from the previous version. New optional conda environments, needed only for the corresponding feature: # REVAMP conda create -c conda-forge -c bioconda -n revamp "blast>=2.13" "taxonkit>=0.20" \ r-base r-dplyr bioconductor-biostrings perl perl-list-moreutils krona # Krona conda create -c conda-forge -c bioconda -n krona krona Verification performed This release was checked by workflow dry runs (DAG construction for the qaqc, repseqs and taxonomy steps), by confirming that a configuration file written for the previous release still parses against the new code, and by validating documentation links, mkdocs navigation targets and YAML syntax. Dry runs confirm that the workflow plans correctly; they do not execute the underlying tools. The REVAMP and Krona features were exercised during development but are new in this release — please report anything unexpected via GitHub issues. Note on AI assistance Parts of this release were prepared with AI assistance: Documentation. The README rewrite and the docs/ overhaul were AI-drafted, working from the Snakefiles, configuration templates and scripts. The factual corrections listed above were found by checking the documentation against the code. All of it was reviewed and edited by the maintainer before release. Pipeline and analysis code. The REVAMP integration, Krona plotting, BLCA fixes and R environment fixes were written by the project maintainers and contributors. AI was used in a supporting role on some commits, but all code was review by a human before commiting. Full changelog Full Changelog: https://github.com/aomlomics/tourmaline/compare/v2.0.1-beta...v2.1.0

View source

Similar papers

#computer vision Conference Aug 2008

A Preliminary Roadmap for Empirical Research on Agile Software Development

Some claim that especially in the field of agile software development the research lags years behind of the practice. In this paper, we characterize the status and main challenges for research on agile software development, and propose a preliminary roadmap, focusing on providing more empirical research, primarily on e...

Torgeir Dingsøyr, T. Dybå, P. Abrahamsson · 92 citations · ⚡7
#computer vision Book Open access Mar 2017

On the Unhappiness of Software Developers

The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 84 citations · ⚡6
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Conference Aug 2008

Scrum in a Multiproject Environment: An Ethnographically-Inspired Case Study on the Adoption Challenges

Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...

A. Marchenko, P. Abrahamsson · 59 citations · ⚡11

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.