Deep Learning Approaches for Apple Disease Classification and Detection: A PRISMA 2020 Systematic Literature Review
Apple diseases cause great losses in fruit yield and quality, while manual diagnosis is time-consuming, subjective, and difficult to scale across orchards. Deep learning has thus become a staple in image-based disease recognition, but evidence remains scattered across classification and object-detection paradigms, heterogeneous datasets, and inconsistent evaluation protocols. This systematic literature review integrates 30 peer-reviewed studies published between 2020 and 2026, selected from 180 database records using a PRISMA 2020-aligned protocol. The final corpus includes 22 image-classification studies and eight object-detection studies. One non-peer-reviewed preprint was kept as contextual evidence only and was excluded from all corpus counts and comparative analyses. Architectures, datasets, preprocessing and augmentation strategies, training configurations, evaluation metrics, evidence of generalization, and deployment characteristics are discussed separately for the two task paradigms. Reported classification accuracies vary from 91.0% to 99.99%, whereas detection studies report mAP@0.5 values from 82.1% to 99.99%; these ranges are descriptive and are not pooled estimates. These values are heavily dependent on dataset composition and evaluation design: controlled-background datasets often yield near-perfect scores, while field transfer can lead to a drop of almost 30 percentage points. CNNs still dominate the field, but hybrid transformer-based attention mechanisms, multi-scale feature fusion, class-imbalance-aware training, and lightweight YOLO variants are gaining traction. Long-standing limitations are the lack of reporting of hyperparameters, an over-reliance on accuracy, the lack of external validation, heterogeneity in annotation, the lack of uncertainty analysis, and limited reporting of latency, memory, and energy consumption. Accordingly, the research agenda emphasizes standardized real-orchard benchmarks, cross-dataset validation, calibrated and explainable predictions, reproducible experimental protocols, and deployment-aware model design.