Skip to content
Preprint

Dueling World Models: Advantage-Style Action Channels for Common-Mode Distractor Rejection

Aug 2026 · 0 citations · 46 references
Computer Science

TL;DR

In latent dynamics, subtracting a prediction's mean effect over actions cancels whatever the actions share--the action-independent variation where distractors live--leaving a clean, controllable channel, with no reward, no reconstruction, and no distractor-specific auxiliary loss.

Abstract

Latent world models plan by predicting future states from an action, but when a scene contains motion the agent does not control, they quietly go action-blind: predictions for different actions become indistinguishable even as the training loss keeps improving. Existing remedies suppress this distraction with reconstruction, task reward, or auxiliary objectives, each adding machinery or assumptions. We show that a minimal alternative suffices, borrowed from the dueling decomposition of value into a state baseline and an action advantage: in latent dynamics, subtracting a prediction's mean effect over actions cancels whatever the actions share--the action-independent variation where distractors live--leaving a clean, controllable channel, with no reward, no reconstruction, and no distractor-specific auxiliary loss. Because this is only a subtraction at readout time, it applies unchanged to any action-conditioned world model, including frozen pretrained ones. Across a gridworld, synthetic generators with known factors, distracting continuous control, and natural-pixel Atari, the isolated channel recovers the agent's own effect where entangled predictors fail, with nuisance leak indistinguishable from zero; applied post hoc it surfaces an action channel in off-the-shelf models that their raw readouts miss, and it converts into goal-reaching control in the gridworld. We prove the cancellation is exact in finite samples for both discrete and sampled action sets, and we state its measured boundary--distractors whose motion tracks the action--together with the remaining limitations in the appendix.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

CAER: Causal Action Effect Reweighting for World Model Training

World models are becoming core infrastructure for embodied intelligence, with action-conditioned video generation providing controllable predictions of how scenes evolve after agent interventions. Yet existing models are commonly trained with space-time-uniform mean squared error, allowing abundant background tokens to dominate the gradient while sparse interaction dynamics remain under-optimized; such uniform fitting rewards reconstructing appearance rather than learning how actions change the world. We introduce Causal Action Effect Reweighting (CAER), a general training paradigm that redistributes supervision toward the tokens whose predicted future is causally affected by the action. CAER contrasts the model's own predictions with and without action conditioning to localize these tokens online, then normalizes the resulting effect map into a weight that preserves the total coefficient mass and changes only where it is spent. This online signal requires no external annotations or offline preprocessing, avoids additional data-processing time, and scales naturally with model and dataset size. Experiments across heterogeneous action-conditioned world-model tasks show that CAER converges to better solutions than uniform MSE training, with consistent improvements in the physical consistency, controllability, and visual quality of generated videos.

Jianjie Fang, Xvyuan Liu, Zi-You Wang et al. · 0 citations
Preprint Aug 2026

GlanceWAM: Sparse Test-Time Imagination for World-Action Models

GlanceWAM is introduced, which decouples imagination from control within a single video DiT: an asynchronous proposer glances ahead on a slow clock to imagine a single lookahead frame seconds into the future in the background, while an action head decodes action chunks at control rate purely in latent space without blocking.

Linhan Wang, Zijian An, Mingyuan Zhang et al. · 0 citations
Preprint Aug 2026

Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency

Action-Conditioned Predictive Consistency (ACPC), a diagnostic that measures how far a clean history and a visually perturbed view of it diverge after being rolled forward under the same action sequence, is introduced and it is proved that this divergence bounds the perturbation-induced change in multi-step prediction error and planner cost.

Guo An, Zijing Wu, Hongzhuang Dong et al. · 1 citation
Preprint Aug 2026

Counterfactual Quotient Models: Learning What Actions Change, Not What the World Does

This work introduces the Counterfactual Quotient Model, which treats action-conditioned futures as equivalent when they differ only by a component shared across actions, and establishes the decision sufficiency, identifiability, common-mode invariance, approximation behavior, and regret properties of the resulting representation.

Junlin Chen, Ruijie Wang, Jianxin Li · 0 citations
#machine learning Preprint Aug 2026

Three Steps at a Time: Learning Representations from Action Sequences in Contrastive RL

While self-supervised approaches to reinforcement learning have achieved strong results by learning representations of states and actions, a key open question is the time scale over which actions should be modeled. Departing from the standard formulation relying on single-step actions, we extend contrastive reinforcement learning (CRL), a prototypical self-supervised method, to operate over action chunks, and find that this results in large, pervasive gains across established offline and online benchmarks: +31.7% and +93.1% across 18 and 11 environments respectively. While action-chunking-driven gains are generally explained through the ability to model non-Markovian, temporally extended policies, and to propagate unbiased multi-step returns, interestingly, we find that these arguments only partially apply to CRL. Our empirical studies suggest that, in the context of CRL, an action chunk carries more information about the goal than a single action, measurably improving the critic's representations, and rendering the algorithm significantly more effective.

Michal Korniak, Kamil Dybek, Benjamin Eysenbach et al. · 0 citations
#artificial intelligence Preprint Aug 2026

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

Contrastive inverse dynamics thus provides a distribution-free anti-collapse signal that requires no target network, stop-gradient, pretrained encoder, or reconstruction objective, and it is argued that the anti-collapse pressure can instead come from the transition data itself.

Jack Boylan, Chris Hokamp · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.