Skip to content
Preprint

Shortcut Learning in a Public Grape Disease Dataset: Annotation Granularity as a Modulator, Not a Cause

Aug 2026 · 0 citations
Computer Science

TL;DR

Annotation granularity is a modulator of this shortcut, not its cause: it can amplify or attenuate a sink that already exists, but cannot create one, and what fixes the destination remains open.

Abstract

Public datasets for agricultural disease detection are usually judged fit for use from reported metrics, which say nothing about whether the annotation scheme is internally consistent. On one public grape disease dataset (3288 images, 11995 boxes, 6 classes), varying model capacity, input resolution and detection paradigm yields a test-set mAP50 range comparable to seed-to-seed noise, with the bottleneck at small objects across all five architectures. The finding lies on the data side: one class is annotated at whole-leaf level (median box area 43.16% of the image) while the other five are annotated at lesion level. On 5156 cross-species images containing no grape, 65.7% of the false-positive boxes fall into that one class, an over-representation of 13.41x relative to its share of the training annotations. Counterfactual retraining establishes a causal effect of granularity on the magnitude of the shortcut: shrinking only that class's boxes cuts its cross-species false positives by 66%, and a placebo control confirms the effect is specific to the manipulated class. A manipulation in the opposite direction, with criteria registered in advance, returns a negative result: coarsening the finest class to whole-leaf level (0.57% to 40.37%), matched in box count and share of annotations and with higher in-distribution AP, still leaves its cross-species false positives at zero boxes, while the unmanipulated original class holds 50.0% of them. Annotation granularity is therefore a modulator of this shortcut, not its cause: it can amplify or attenuate a sink that already exists, but cannot create one, and what fixes the destination remains open. We also give a granularity screening statistic requiring neither images nor training, and show airborne lesion-level detection to be optically out of reach. The failure mode is invisible to in-distribution evaluation.

View source

Similar papers

Conference Aug 2026

Systematic Evaluation of Cross-Dataset Domain Shift and Shortcut Learning in Lightweight CNNs for Tomato Disease Classification

The reliable deployment of deep learning architectures for automated plant pathology remains a significant challenge due to the pronounced performance gap between laboratory-curated datasets and real-world field conditions. This study presents a rigorous investigation into cross-dataset domain shift and shortcut learni...

D. Sreekanth, Dilshad Shaik, Pothuraju Siva Kumar et al. · 0 citations
Open access Aug 2026

Annotation Refinement and Minority-Class Augmentation for Coffee Leaf Disease Detection Using YOLOv8

Coffee leaf diseases negatively affect crop productivity, making early detection an essential task in precision agriculture. The widely used BRACOL dataset presents critical object detection challenges related to spatial localization consistency and extreme class imbalance. Rather than proposing a new detection archite...

Naila Jinan Gaisani, E. Y. Puspaningrum, B. M. Mulyo · 0 citations
Preprint Aug 2026

A Dataset-Centric Benchmark of Deep Learning Methods for Grape Leaf Disease Classification and Detection

Grape leaf disease recognition is important for precision agriculture, enabling early diagnosis, timely intervention, and improved vineyard management. Although deep learning has achieved strong results, many studies rely on few datasets, often acquired under controlled conditions, and may not reflect real vineyard cha...

Petar Canoski, Vlatko Spasev, Ivica Dimitrovski et al. · 0 citations
Open access Sep 2026

Saturated In-Domain, Separable Out-of-Domain: A Hierarchical Multi-Scale Lesion-Attention Network and a Zero-Shot Cross-Corpus Protocol for Multi-Crop Leaf Disease Classification

Leaf disease classifiers are ranked by held-out accuracy on their training corpus, where that number now exceeds 98% for many modern backbones. We ask what such a measurement can still resolve. On a four-crop, 21-class corpus of 7179 de-duplicated images, we train 26 architectures under one protocol with three seeds an...

Songul Karakus, Mehmet Burukanli, Davut Ari · 0 citations
Open access Aug 2026

Data-Centric Evaluation of Protein Function Prediction Pipelines

Findings show that performance estimates in protein function prediction should be interpreted as outcomes of complete data-centric workflows rather than isolated properties of predictive models.

Nicole Soto-García, Norma Murillo-Acevedo, Julián García-Vinuesa et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.