Credit assignment in multi-turn agent reinforcement learning operates at two levels: assigning trajectory-level credit to actions and distributing each action's credit across its tokens. In this paper, we introduce FACTOR, which separates these decisions. FACTOR uses checkpoint-calibrated TD residuals to assign per-action credits that telescope to the trajectory advantage, and feedback-conditioned teacher-student likelihood gaps to allocate each credit across the realized action tokens. Per-action normalization preserves the action-average coefficient and prevents token-level sign flips. We pair this construction with an action-mean reduction, removing the implicit dependence of an action's scalar surrogate weight on its token length. At the behavior policy and before clipping, each action's inner action-mean surrogate equals its TD credit. FACTOR consistently improves over competitive baselines across ALFWorld, WebShop, and ScienceWorld, with every environment-seed comparison favoring FACTOR and the largest gains emerging on the longest-horizon environment. The same hyperparameters transfer without retuning to a larger backbone and to a different model family. Ablations identify TD action credit as the dominant driver of the improvement, with hindsight token allocation contributing complementary gains.
Li-Chao Ma, Yang Sun, Shuai Zhao et al.· 0 citations
Imitation learning has emerged as a promising paradigm for autonomous vehicle control. However, existing policies often suffer from severe covariate shift and lack explicit stability guarantees in closed-loop trajectory tracking. This paper proposes an adaptive Lyapunov-constrained imitation learning framework for vehicle trajectory tracking. To address the covariate shift problem, we explicitly embed a discrete-time kinematic bicycle model into the policy training process. This differentiable physical prior enables propagation of the nominal tracking error and allows a global Lyapunov decay constraint to be directly employed on the policy optimization via an Actor-Critic architecture. To prevent the overly aggressive penalization near the reference trajectory, we introduce a state-dependent adaptive decay rate that relaxes the constraint for small errors while strengthening it under large deviations. As a result, the proposed method achieves both high tracking precision and strong recovery from model uncertainties and disturbances. Extensive experiments in interactive traffic environments demonstrate that the proposed method significantly outperforms state-of-the-art baselines in collision avoidance and disturbance rejection.
Yuchen Wei, Yu-Hsiang Su, F. Arvin et al.· 2026 IEEE/ASME International...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.