Skip to content
Preprint

On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models

Aug 2026 · 0 citations · 25 references
Computer Science

TL;DR

The irreducible action-specific prediction error of future models that do not condition on the candidate action is characterized, conditions under which a world-action joint can recover an interventional forward model are identified, and an environment family is constructed in which every observational learner has positive worst-case regret.

Abstract

World-action models predict a future outcome and then infer an associated action. Although this factorization can improve representation learning and data efficiency, it is unclear whether it provides stronger control capability than direct behavior cloning when both are trained from the same observational demonstrations. We compare a direct behavior-cloning policy, an imitation-trained world-action policy, and a policy optimized with an action-conditioned world model. At the controller-class level, every world-action policy can be flattened into a direct stochastic policy with the same closed-loop trajectory distribution. At the population level, under realizability, exact optimization, common deployment information, and distribution-preserving deployment, direct behavior cloning and world-action imitation both recover the observational behavior policy. Thus, future prediction changes the learning factorization but not the unrestricted external policy class or ideal imitation target. Action-conditioned world-model learning differs by predicting outcomes under specified actions and comparing them through a control objective. We characterize the irreducible action-specific prediction error of future models that do not condition on the candidate action, identify conditions under which a world-action joint can recover an interventional forward model, and show that observational demonstrations do not identify action effects in general. Finally, we construct an environment family in which every observational learner has positive worst-case regret, whereas one informative intervention permits zero regret. The key distinction is therefore between predicting futures associated with observed behavior and predicting consequences of specified actions for policy optimization.

View source

Similar papers

Preprint Aug 2026

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

This work proposes a compatibility prediction Latent World Model for robot navigation that predicts action-conditioned latent feature compatibility rather than reconstructing observations and demonstrates how the learned world model can supervise policy learning from unlabeled video data and improve policies through reinforcement learning entirely within the world model.

Zengmao Wang, Wei Gao, Shuhan Shen · 0 citations
Preprint Aug 2026

IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning

Model-based reinforcement learning (MBRL), which learns environment dynamics to generate synthetic experience, is a promising approach to sample-efficient decision making. Numerous methods have been developed to improve dynamics prediction and policy optimization for MBRL through uncertainty estimation, model regularization, and conservative value learning. However, these methods typically treat the transition model and critic as monolithic predictors, overlooking the policy-induced data bias. Consequently, action can become entangled with environmental evolution, while uneven action coverage may distort the counterfactual value estimates used for policy improvement. To address this, we propose IADD-TR, a unified framework combining Intervention-Aware Dynamics Decoupling (IADD) and Targeted Regularization (TR). IADD factorizes transitions into an action-intervention stage and an action-free natural evolution stage, using a zero-action anchor to resolve the non-uniqueness of this two-stage factorization for robust generalization. Its latent and state-aligned components are identifiable up to an invertible within-block transformation and pointwise, respectively. For policy learning, we derive TR from the efficient influence function of a replay-state policy-gradient functional. TR augments the critic with an action-density-scaled residual correction and optimizes a targeted loss, yielding doubly robust policy-gradient estimation when either the critic or the replay action density is consistently specified. Extensive experiments on five MuJoCo tasks show that IADD-TR achieves competitive returns with improved sample efficiency.

Ze-Feng Liang, Jie Qiao, Ruichu Cai et al. · 0 citations
#diffusion models Preprint Sep 2026

JEPA Policy: Diffusion-Free Imitation Learning via Paired Action and Future Representation Prediction

JEPA Policy is introduced, a diffusion-free framework that uses the action chunk and its observed future representation as paired training targets as paired training targets and supports paired future-representation supervision supervision as a practical approach to low-latency visuomotor imitation without iterative generative sampling.

Unknown authors · 0 citations
Preprint Aug 2026

Counterfactual Quotient Models: Learning What Actions Change, Not What the World Does

This work introduces the Counterfactual Quotient Model, which treats action-conditioned futures as equivalent when they differ only by a component shared across actions, and establishes the decision sufficiency, identifiability, common-mode invariance, approximation behavior, and regret properties of the resulting representation.

Junlin Chen, Ruijie Wang, Jianxin Li · 0 citations
#artificial intelligence Preprint Aug 2026

Q-Learning With World Models

This work proposes QWM, a framework that leverages world models to perform test-time search over imagined trajectories on top of Q-learning to select high-value actions during both online rollouts and evaluation, and significantly outperforms strong prior state-of-the-art methods on both sample efficiency and performance.

Perry Dong, Yueru Jia, Chelsea Finn et al. · 0 citations
Jul 2026

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

This work introduces OVI, an interactive on-policy IL algorithm that is statistically efficient whenever the learner can represent the expert's value function and computationally efficient given access to a linear maximization oracle, and introduces a negative result showing that interaction is necessary.

Luca Viano, Antoine Moulin, Audrey Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.