Jul 2026· Journal of Natural History· Vol 60, pp. 1535 - 1551· 0 citations· 24 references
TL;DR
The software has been applied across diverse biological collections and labelling scenarios, demonstrating its utility as a practical and accessible solution for researchers, curators, collection managers and private collectors seeking to standardise and streamline specimen label production.
Abstract
ABSTRACT Labels are essential components of biological collections, preserving specimen-associated metadata required for identification, curation, reuse and long-term scientific use. Despite their importance, label production often relies on general-purpose software or specialised workflows that can be difficult to standardise, time-consuming, or inaccessible to users without programming expertise. As digitisation efforts expand, there is a growing need for flexible and reproducible label-production tools. Here we present EntomoLabels, a free Windows desktop application developed to support the design and batch production of labels for biological collections. Originally created for entomological collections, where labels are typically small and highly standardised, the software can be adapted to a broad range of specimen types and collection workflows. Users can define label dimensions and layout, design labels using text fields, tags, lines, frames, images and machine-readable codes, and generate batches of labels from structured comma-separated values (CSV) files. Additional features include adjustable label and page dimensions, Unicode support, reusable templates, automatic serial numbering and generation of both one-dimensional and two-dimensional barcodes. Representative applications demonstrated here include compact entomological pin labels, barcode identifier labels, herbarium labels, labels for fluid-preserved specimens, tissue or DNA tubes, double-sided study-skin tags, storage-box labels and slide-mounted specimens. By separating label design from specimen, EntomoLabels promotes repeatable workflows, reduces formatting inconsistencies and facilitates integration between physical specimens and digital collection records. Unlike many existing label-production approaches that depend on scripts, databases or barcode-centred workflows, EntomoLabels emphasises visual label composition through an interactive graphical interface, allowing users to directly control layout and printing without requiring programming skills. The software has been applied across diverse biological collections and labelling scenarios, demonstrating its utility as a practical and accessible solution for researchers, curators, collection managers and private collectors seeking to standardise and streamline specimen label production.
New approaches are proposed to simplify and formalize multi-panel figure composition, emphasizing rapid layout construction and experimentation while maintaining explicit control over publication parameters.
Abstract Premise Accessioning herbarium specimens is labor intensive, yet remains vital for research in ecology, evolution, and conservation. As institutional support for herbaria declines, efficient tools are needed to streamline this process. The R package BarnebyLives was developed to assist collectors by supplementing collection notes, verifying taxonomic data, conducting quality checks, generating labels, and submitting digital records. Methods and Results BarnebyLives integrates geospatial data from U.S. government sources to provide jurisdictional and site information and checks taxonomic names using in‐house spell checkers, International Plant Names Index (IPNI) author standards, and Kew's Plants of the World Online. Optional features include generating Google Maps driving directions. The tool outputs data in tabular and spatial formats for review before producing LaTeX‐based labels and shipping manifests. Conclusions BarnebyLives improves data accuracy, ensures up‐to‐date taxonomy, and significantly reduces the time and effort required to accession herbarium specimens in the United States.
Reed Clark Benkendorf, J. Fant· Applications in Plant Scienc...· 0 citations
Digitized herbarium collections, now comprising over 100 million freely accessible specimen images, have become a critical resource for addressing fundamental questions in ecology and evolutionary biology. Yet the rich metadata encoded in herbarium labels (collector identities, geographic localities, collection dates, and ecological observations) remains largely inaccessible at scale, constraining both biodiversity informatics and the construction of specimen-specific image-text corpora for multimodal AI. We present HERBIOME, a modular end-to-end pipeline for automated herbarium label digitization, integrating YOLOv8-based component detection, CRAFT Hezar word-level text localization, fine-tuned TrOCR for recognition of mixed handwritten and printed text, and GPT-4o Mini for semantic metadata structuring into standardized fields. TrOCR was trained on a multi-source dataset combining general transcription corpora (CREMMA-AN, PictoCatalogs) with herbarium-specific data (R\'eColNat), achieving a Character Error Rate of 4.05-4.10%. End-to-end evaluation on 450 French herbarium specimens, using a dual-metric framework of Maximum Window Similarity (MWS: 0.614-0.618) and Semantic Metadata Accuracy (SMA: 0.440-0.445), reveals that hybrid training strategies improve semantic fidelity while random sampling maximizes surface similarity, with taxonomic fields remaining the principal bottleneck. By automating the extraction of structured metadata from complex, heterogeneous labels, HERBIOME reduces transcription burden, enables the construction of paired image-text datasets that faithfully capture specimen individuality, which is a prerequisite for next-generation multimodal biodiversity AI systems.
Hiba Abbad, Hanane Ariouat, Eva Perez Pimparé et al.· 0 citations
Herbaria serve as invaluable spatio-temporal repositories of plant diversity information. Digitization of herbarium collections enhances the accessibility, discoverability, and long-term preservation of this important plant information, yet financial and infrastructural constraints often prevent herbaria in resource-constrained regions from digitizing their collections. Consequently, critical plant diversity data gaps remain due to underrepresentation of these collections in global biodiversity databases. Here, we describe an AI-assisted modular digitization toolkit specifically designed for herbaria operating under limited funding, developed and refined through our experience digitizing the crop wild relative (CWR) collection of the National Herbarium of Zimbabwe. The toolkit comprises three core components: (1) a portable, cost-effective photostation assembled from commodity parts, (2) a streamlined cascade workflow for systematic digital imaging, and (3) an AI-assisted data management pipeline for image quality control, label transcription, data analysis, and presentation. Compared to manual transcription and legacy optical character recognition approaches, AI-based transcription achieves lower time cost while maintaining high accuracy, and AI-driven data management delivers accessibility and reduced expenditure relative to conventional database infrastructure. The toolkit is designed to allow herbarium staff full autonomy over the digitization procedure, ensuring institutional ownership and the capacity for independent continuation beyond initial project support. By prioritizing affordability, modularity, and simplicity, this toolkit provides a replicable framework that may enable resource-constrained herbaria to locally generate high-quality scientific data for conservation and the sustainable utilization of plant genetic resources.
Langalenkosi Gatula, I. Bezrukov, Joseph Atemia et al.· bioRxiv· 0 citations
Natural history collections are structured knowledge systems in which taxonomic and historical information is embedded within specimens, labels and their physical arrangement. However, in large entomological collections, much of this information remains inaccessible because it is not captured in standardised digital formats. Whole-drawer imaging offers a scalable approach to digitisation, but existing workflows primarily focus on specimen-level extraction and often neglect the curatorial structure encoded in drawer organisation. Here, we present an open-source workflow that transforms whole-drawer images into structured, machine-readable inventory data, while preserving their spatial arrangement and taxonomic context. The pipeline integrates low-cost imaging, barcode linkage, deep learning-based object detection (YOLOv.11) and optical character recognition (OCR) to reconstruct taxonomic series and estimate specimen counts per taxon from drawer images. The approach was developed for the Coleoptera collection of the Senckenberg Research Institute Frankfurt, encompassing approximately 1.9 million specimens which are stored in more than 5,000 drawers. Model training on Carabidae collection drawers overall achieved high performance (precision 99.2%, recall 97.9%, mAP0.5 98.5%), with reasonable transferability to other beetle families and Hymenoptera drawers tested. By treating the drawer as the primary unit of digitisation, this workflow provides a scalable intermediate layer between physical collections and specimen-level databases. It enables rapid assessment of taxonomic composition, supports collection management and digitisation planning and contributes to the mobilisation of biodiversity data from large historical collections.
Sneha Bhansali, Nikolai Ignatev, M. Simões· Natural History Collections...· 0 citations
Bioinformatics workflows rely heavily on visual representations. Quality-control plots, cell embeddings, heatmaps, genome-browser tracks, and interactive dashboards are not merely illustrations, but instruments for making analytical decisions. For blind and low-vision researchers who use screen readers, braille displays, or audio-based interfaces, these create a barrier: the evidence used to justify an analysis is often encoded in visual form, while the underlying decision remains undocumented. We argue that non-visual accessibility and computational reproducibility are closely aligned, as they both require analyses to be transparent and to record why decisions were made. We present ten simple rules for non-visual bioinformatics, covering plots as decision records, cautious use of AI-generated figure descriptions, accessible computing environments, text-first literate programming, structured data and metadata, compact object summaries, accessible publication formats, collaboration practices, shared community infrastructure, and accessibility as part of FAIR research. The intended audience is computational biologists and developers. Using single-cell RNA-seq as a running example, we show that the accessible equivalent of a plot is a structured decision record. That is, a plot companion that goes beyond storing the underlying data by also stating the purpose of the analysis and the resulting quantitative evidence and uncertainty. We argue that treating accessibility in this way makes bioinformatics more inclusive and also more transparent and auditable.
Jacqueline G. Kientsch, S. Neuhauss, I. Mallona· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.