Predictive Model for Last-Mile Delivery Failure Detection: A Systematic Literature Review and Comparative Experiment on a National Logistics Operator’s Data
Abstract
Failed last-mile deliveries impose cascading operational costs on postal networks, yet most predictive approaches rely on synthetic datasets or single-method evaluations that limit real-world applicability. This paper addresses that gap through a PRISMA-compliant Systematic Literature Review (SLR) of 261 Scopus records (2020–2025), retaining 22 peer-reviewed studies, and a rigorous comparative experiment on 100,000 real shipment records from the Operator, a national last-mile logistics enterprise in Indonesia. Three machine learning classifiers are benchmarked under a severe class imbalance of 16.7:1 (94.4% successful vs. 5.6% failed deliveries): XGBoost, Random Forest, and a Multilayer Perceptron (MLP). Random Forest with class weighting achieves the highest discriminative performance (AUC-ROC = 0.842, AUC-PR = 0.986 for success detection / 0.499 for failure detection ≈9× above the 5.6% failure-class baseline, Balanced Accuracy = 0.756) and is recommended for operational deployment. Beyond outcome prediction, this study introduces a courier-centric operational friction analysis that translates XG-Boost feature importance scores into four actionable burden categories: transactional & wait-time friction (cash payment: 0.196), administrative & regulatory friction (stamp service/regulated category: 0.124–0.086), physical & mobility friction (package weight: 0.041), and temporal friction (weekend patterns: 0.048). These findings offer the Operator concrete SOP improvements, including pre-notifying cash recipients, streamlining government-program compliance, enforcing weight-based courier assignment, and shifting at-risk deliveries to weekday slots.