ProcessLight: Process Supervision for Large Language Model Based Traffic Signal Control
This work proposes an LLM-based framework ProcessLight to decompose signal decisions into verifiable semantic steps and develops Step-wise Traffic Process Policy Optimization (STeP-PO), a novel reinforcement learning framework that optimizes structured reasoning processes through step-level credit assignment.