Reliability Beyond Accuracy in Crop Classification Benchmarks (Supplementary Materials)
Abstract
Reliability Beyond Accuracy in Crop Classification Benchmarks (Supplementary Materials)Introduction: Near-perfect crop-label accuracy can conceal uncertainty, perturbation sensitivity, and weak explanations. This study evaluates these reliability dimensions without treating benchmark classification as agronomic recommendation. Materials and Methods: Ten classifiers were evaluated using five matched stratified fitting, validation, and test partitions on a 2,200-record crop-label benchmark and an independent 3,810-record rice-grain benchmark. A Calibrated Selective Forest combined Random Forest, validation-fitted temperature scaling, and confidence-based abstention. Analyses included feature and component ablations, probability calibration, Gaussian input perturbations, computational cost, and explanation fidelity. Results: Mean crop-label accuracy was 99.59% for Random Forest, 99.23% for XGBoost, and 98.32% for FT-Transformer. Random Forest calibration reduced mean expected calibration error from 0.0449 to 0.0038. At a 0.95 threshold, calibrated selection retained 98.59% of predictions with 0.091% mean selective error. Under perturbations of 0.20 training standard deviations, its accuracy fell to 86.96% and selective error rose to 4.29%. On rice grains, Random Forest accuracy was 91.71%, and calibration did not improve every probability metric. Local explanation rankings were often stable despite a median neighborhood weighted R² of 0.293 under the specified local interpretable model-agnostic explanations (LIME) configuration. Conclusions: Accuracy, calibration, selective reliability, and explanation fidelity are distinct properties. The composite pipeline provides an auditable benchmark procedure, not a new learning algorithm or a validated farm recommendation system. Geographic and temporal agronomic validation remain necessary.