Skip to content
#gene editing Open access

Bactomics HybAs v8.4-lite

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This release is a major workflow upgrade from the legacy 8.4-lite Snakefile to a batch-capable, provenance-aware, benchmarked, validation-enabled, report-ready HybAs implementation. The new workflow preserves the core hybrid assembly logic while substantially improving execution control, reproducibility, taxonomic decision support, metadata capture, validation depth, and operational robustness. In practical terms, HybAs now moves from a single-isolate pipeline into a more scalable, auditable, and engineering-grade workflow platform for repeated multi-isolate use. Highlights Batch-capable workflow architecture Native multi-isolate execution via samples.tsv Legacy single-isolate operation retained through config.yaml Cleaner workflow structure suitable for repeated runs and comparative isolate processing Explicit run-mode handling The workflow now resolves and records isolate processing mode explicitly: hybrid ont_only illumina_only This makes execution state clearer and improves downstream reporting and auditability. Configurable polishing strategy HybAs now formalizes polishing as an explicit workflow setting through polish_mode: none racon medaka tripolish tripolish is now defined as: Racon → Medaka → Polypolish This provides cleaner control over polishing depth and makes stage-resolved behavior easier to interpret. Taxon AWARE Engine The previous Kraken handling has been upgraded into the Taxon AWARE Engine, a structured taxonomic decision layer rather than a simple static filtering step. Supported modes include: off manual auto aware This engine improves decision transparency by recording: target selection awareness state confidence logic decision reason per-platform cleaning decisions for Illumina and ONT data Stronger preflight validation The workflow now performs stricter early validation of: configuration values run-mode compatibility required input reads Kraken database readiness BUSCO lineage availability This shifts failures earlier in execution and makes troubleshooting much more transparent. Optional module toggles Explicit enable/disable control is now available for: Kraken BUSCO Prokka MultiQC Polypolish This gives users more control over runtime, dependencies, and output scope. Polypolish hardening The Polypolish pathway has been made more robust by improving: mapping behavior SAM/BAM handling preservation of queryname-sorted alignment requirements safe handling of empty or invalid polishing-input conditions Provenance, metadata, and benchmarking Structured provenance capture This release adds both workflow-level and isolate-level provenance outputs, making HybAs substantially more reproducible and auditable. Captured provenance now includes: workflow metadata config and sample-sheet provenance runtime and system information conda environment manifests conda exports pip freeze outputs Python version capture database hashes input hashes output checksums Integrated per-isolate metadata outputs New isolate-level metadata products include structured summaries such as: run_summary.tsv run_summary.json tool_versions.tsv checksums.tsv skip_reasons.tsv qc_assessment.tsv kraken_decision.tsv These outputs make the workflow much easier to inspect, validate, and compare across isolates. Performance benchmarking Performance benchmarking is now a first-class workflow feature. Benchmark outputs are generated systematically and stored separately from analytical outputs, improving: runtime tracking rule-level performance review execution profiling workflow auditing across isolates Integrated validation subsystem This release introduces a formal downstream validation subsystem implemented as a dedicated validation workflow layered on top of HybAs outputs. This is not a loose script collection. It is a structured multi-phase validation module with configuration, orchestration, analytical cores, summary generation, and automated plot production. Dedicated validation workflow A dedicated validation subsnake is now included, driven through validation_config.yaml. It links validation execution to the HybAs directory structure and shared sample sheet, and coordinates: Phase 1 validation Phase 2 validation Phase 3 validation automated plotting Phase 1 validation extractor Phase 1 consumes existing HybAs outputs for one isolate and builds a structured validation layer without rerunning the main pipeline. It summarizes: stage FASTA progression structural metrics by stage N50, genome size, and contig count BUSCO by stage dnadiff stage-vs-final comparisons dnadiff prev-vs-next comparisons polishing modifications versus previous stage ONT and Illumina retention ONT coverage Kraken report summaries Illumina mapping QC CDS counts circularity evidence bundled validation exports Phase 2 analytical validation layer Phase 2 adds a Python-driven analytical layer over Phase 1 outputs. It extends validation into: BUSCO locus tracking across stages BUSCO degradation, recovery, gain, and loss analysis hotspot window analysis stage-wise CDS-space profiling protein-space summary behavior across polishing stages This pushes validation beyond surface assembly metrics into gene-space consequences of polishing decisions. Phase 3 spatial and mechanistic validation layer Phase 3 builds a third validation layer using Phase 1 and Phase 2 products. It includes: edit extraction from dnadiff .snps files genome-wide edit coordinate tabulation fixed-window edit-density mapping hotspot ranking spatial clustering statistics negative-binomial descriptor fitting sequence-complexity profiling hotspot-versus-background analysis bundled validation exports This makes the polishing trajectory interpretable as a genome-wide edit topology rather than only a final endpoint. Phase 3 extension layer An additional Phase 3 extension module expands biological interpretation through: hotspot feature overlap independent rRNA adjudication with Barrnap Barrnap profiling across polishing stages hotspot versus rRNA overlap hotspot core-sequence export motif and repeat microprofiling motif enrichment against genome background hotspot classification read-level inspection support Compact user-facing validation summary A final summary layer converts the detailed validation outputs into a compact isolate-level interpretation, including: convergence status dispersion behavior dominant hotspot mode rRNA-associated instability signals read-support summaries Automated validation plotting package A dedicated plotting layer now reads completed Phase 1–3 validation tables and generates a user-facing figure package under: {isolate}/validation/plots/ This gives HybAs a clearer interpretive output layer for validation, review, and presentation. Output and reporting improvements Report-friendly structure Outputs are now organized in a cleaner, more structured, report-compatible layout. This improves downstream use for: validation review benchmarking manuscript support reproducibility reporting MultiQC aggregation Improved directory organization Internal products are now separated into clearer directories such as: merged inputs cleaned inputs filtered inputs assembly polishing mapping final metadata benchmarks validation This reduces clutter and improves workflow traceability. Centralized database management Databases are now expected under a shared db root, with explicit update-style rules for major resources such as Kraken and BUSCO. Normal workflow execution does not auto-download databases, improving reproducibility and administrative control. Changed behavior Compared with the older 8.4-lite Snakefile, this release changes HybAs in several important ways: HybAs is no longer just a flat single-isolate Snakefile; it now behaves as a more formal workflow platform taxonomic handling is now governed by the Taxon AWARE Engine provenance and benchmarking are now formal workflow outputs output layout is substantially more structured validation is now an integrated subsystem rather than an external afterthought final assembly handling is more traceable and reproducibility-oriented downstream scripts relying on older flat output paths may require updating Why this release matters This version marks the transition from a functional assembly pipeline to a more reproducible, validation-aware, and engineering-grade workflow platform. The main gains are not only in assembly execution, but in: auditability traceability taxonomic decision transparency provenance capture performance benchmarking structured post-run validation batch-scale usability report-ready output design This release is especially important for users who need: repeated multi-isolate processing explicit taxonomic decision logic stage-resolved polishing interpretation provenance suitable for validation and reproducibility rule-level performance benchmarking integrated validation outputs beyond standard assembly metrics cleaner workflow-state tracking for reporting, review, or publication support Upgrade notes Users upgrading from the legacy workflow should review: config.yaml samples.tsv output-path assumptions database paths Taxon AWARE Engine settings module toggles run-mode controls validation workflow configuration

View source

Similar papers

#computer vision Conference Aug 2008

A Preliminary Roadmap for Empirical Research on Agile Software Development

Some claim that especially in the field of agile software development the research lags years behind of the practice. In this paper, we characterize the status and main challenges for research on agile software development, and propose a preliminary roadmap, focusing on providing more empirical research, primarily on e...

Torgeir Dingsøyr, T. Dybå, P. Abrahamsson · 92 citations · ⚡7
#computer vision Book Open access Mar 2017

On the Unhappiness of Software Developers

The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 84 citations · ⚡6
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Conference Aug 2008

Scrum in a Multiproject Environment: An Ethnographically-Inspired Case Study on the Adoption Challenges

Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...

A. Marchenko, P. Abrahamsson · 59 citations · ⚡11

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.