Reference databases shape the taxonomic resolution, uncertainty, and reproducibility of metabarcoding analyses. For ITS barcodes, public references are distributed across repositories with different taxonomic conventions, geographic coverage, and annotation practices, creating conflicts, missing ranks, and misannotations when databases are merged or compared. We introduce CurateMake, a reproducible Snakemake workflow for ITS reference database construction, harmonisation, and validation. It integrates four public sources (UNITE, BOLD, PLANiTS, and CALeDNA) and user-supplied databases, combines Catalogue of Life name harmonisation with ITSx-based region standardisation, MSA/HMM-based alignment grouping, and SATIVA phylogenetic validation. Raw, CoL-harmonised, and SATIVA-validated annotation layers are retained throughout to compare curation effects while preserving flagged records for review. We evaluated CurateMake on 3.58 million ingested sequences and controlled error-injection simulations. ITSx expanded the final harmonised database to 5.19 million barcode-resolved entries by recovering ITS1 and ITS2 sub-regions from full-length ITS records. Across the full dataset, normalised intra-cluster entropy decreased from Raw to CoL-harmonised to SATIVA-validated annotations, consistent with improved taxonomic coherence. In simulations, CurateMake achieved the highest correction rate across 1%–50% corruption and, at 15% corruption, corrected 42% ± 1% of introduced errors, compared with 28% ± 1% for CoL alone and 0% for SATIVA without the workflow’s alignment infrastructure. These results show that nomenclatural harmonisation and phylogeny-informed validation address complementary error classes, with phylogenetic validation contributing measurably only within taxon-coherent alignments in this benchmark. CurateMake therefore provides a reproducible, provenance-tracked framework for auditable ITS reference database curation in metabarcoding workflows.
Auguste Gardette, Eugeni Belda, Edi Prifti et al.· bioRxiv· 0 citations
Digitized herbarium collections, now comprising over 100 million freely accessible specimen images, have become a critical resource for addressing fundamental questions in ecology and evolutionary biology. Yet the rich metadata encoded in herbarium labels (collector identities, geographic localities, collection dates, and ecological observations) remains largely inaccessible at scale, constraining both biodiversity informatics and the construction of specimen-specific image-text corpora for multimodal AI. We present HERBIOME, a modular end-to-end pipeline for automated herbarium label digitization, integrating YOLOv8-based component detection, CRAFT Hezar word-level text localization, fine-tuned TrOCR for recognition of mixed handwritten and printed text, and GPT-4o Mini for semantic metadata structuring into standardized fields. TrOCR was trained on a multi-source dataset combining general transcription corpora (CREMMA-AN, PictoCatalogs) with herbarium-specific data (R\'eColNat), achieving a Character Error Rate of 4.05-4.10%. End-to-end evaluation on 450 French herbarium specimens, using a dual-metric framework of Maximum Window Similarity (MWS: 0.614-0.618) and Semantic Metadata Accuracy (SMA: 0.440-0.445), reveals that hybrid training strategies improve semantic fidelity while random sampling maximizes surface similarity, with taxonomic fields remaining the principal bottleneck. By automating the extraction of structured metadata from complex, heterogeneous labels, HERBIOME reduces transcription burden, enables the construction of paired image-text datasets that faithfully capture specimen individuality, which is a prerequisite for next-generation multimodal biodiversity AI systems.
Hiba Abbad, Hanane Ariouat, Eva Perez Pimparé et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.