Qualitative and quantitative analyses underline that current LTR optimisations cannot fully overcome representational bottlenecks, resulting in catastrophic predictive breakdown for rare `Tail'classes under severe domain shift.
Abstract
As with most ``in the wild''collections of the natural world, the North America Camera Trap Images (NACTI) dataset exhibits long-tailed class imbalance, with the largest class covering over 50% of its 3.7M images. Building on the PyTorch Wildlife model, we systematically evaluate Long-Tail Recognition (LTR) methodologies to benchmark species recognition performance, including specialised loss functions and LTR-sensitive regularisation. Our optimised configuration achieves state-of-the-art 99.40% Top-1 accuracy on the NACTI test split, significantly outperforming standard baselines and previously reported top performances. To assess robustness under domain shifts (e.g., night-time captures, occlusion, motion-blur), we extend our evaluation across three independent reduced-bias test sets (including ENA-Detection, Caltech Camera Traps and Missouri Camera Traps). Across these out-of-distribution (OOD) evaluations, our LTR-enhanced model consistently demonstrates substantially stronger generalisation capabilities compared to standard cross-entropy approaches. However, qualitative and quantitative analyses underline that current LTR optimisations cannot fully overcome representational bottlenecks, resulting in catastrophic predictive breakdown for rare `Tail'classes under severe domain shift. For maximum reproducibility, all dataset splits, key code, and network weights are published with this paper at https://github.com/ZehuaLiuY/Species-Classification.
It is tested whether models in this deployment-relevant 2--8B range carry genuine taxonomic knowledge, evaluating four such VLMs against the domain-specific specialist BioCLIP on a 96-species task, and comparing clean iNaturalist photographs against camera-trap imagery from 6 LILA.
William Zhou, Mayukha Siripuram, Xiao-Wei Yan et al.· 0 citations
Passive acoustic monitoring produces far more bat recordings than experts can label. We show that simple model-generated pseudo-labels turn this surplus into effective supervision. We compare pseudo-labeling with other semi-supervised learning methods on an 18-species European corpus using only 10% of its training labe...
Zero-shot vision-language models (VLMs) are increasingly used as training-free species recognizers, but reported accuracy can reflect more than visual species knowledge. We audit CLIP, BioCLIP, BioCLIP2, and a multilingual Jina CLIP v2 control on seven freshwater-fish categories from two Bangladeshi sources (10,321 ima...
Crop and forest damage from sika deer, wild boar, and Japanese macaque is a serious economic problem in Japanese satoyama, where farmland and forest intermingle. Camera traps enable scalable monitoring, yet existing benchmarks evaluate recognition in-distribution, rarely prioritize night infrared imagery, and do not jo...
Keito Inoshita, Kohei Hisayama, Haruto Sugeno et al.· 0 citations
A novel end-to-end framework integrating a self-attention mechanism to address limitations in effectively detecting small animals in low-contrast trap images and small animals while also demonstrating zero-shot detection capability leveraging the MLLM.
Nowshin Amin, Nafisa Tabassum Oyshi, Tahmid Abrar Zidan et al.· 0 citations
Alignment-free lineage assignment from k-mer frequency profiles is widely used for SARS-CoV-2 surveillance, and the methods that do it are ranked against each other by margins of one or two percentage points. Those rankings rest on an unchecked protocol. Public repositories hold many near-duplicate genomes, and stratif...
Mohammad Jamhuri, A. Irawan· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.