Carbon-Aware Offline Reinforcement Learning for Joint Inventory Replenishment and Transportation Scheduling in Multi-Echelon Supply Chains
Joint replenishment and transportation decisions couple service, inventory, vehicle utilization, and carbon emissions over time. Offline reinforcement learning (RL) is attractive when operational logs exist but online exploration is unsafe; however, distribution shift and cumulative carbon constraints make direct offline Q-learning unreliable. This paper proposes budget-conditioned carbon-aware conservative Q-learning (CA-CQL), a discrete-action offline method that combines a remaining-budget state, carbon-shaped conservative value learning, a learned behavior-support prior, a state-dependent carbon shadow price, and a pathwise carbon-feasibility screen. A reproducible supplier–distribution-center–three-retailer benchmark is constructed because public sales datasets lack synchronized replenishment actions, vehicle schedules, loads, and emissions. Transport emissions are calibrated with the official UK Government 2026 vehicle-kilometer conversion factors. Across five independently generated offline datasets and 250 evaluation episodes per policy, CA-CQL achieves an operating cost of 16,865 ± 404, emissions of 14,119 ± 57 kg CO₂e, a 92.18 ± 0.88% fill rate, and zero carbon-cap violations. Relative to a support-regularized CQL baseline, it reduces emissions by 110.8 kg and the violation rate by 13.6 percentage points, with a 1.36% cost increase and a 0.63-point fill-rate decrease. Ablations show that the behavior prior prevents severe under-replenishment and that the carbon screen is necessary for hard pathwise compliance. The results establish a transparent compliance–service trade-off rather than claiming universal dominance.