Skip to content

Conservation Buys Stability and Factoring Buys Counterfactuals in Physical World Models

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

A concrete design principle for physical world models: long-horizon stability and changed-law generalization arise from distinct structural commitments, and each can be imposed deliberately without requiring the other.

Abstract

A learned simulator can reproduce its training conditions accurately yet fail in two distinct ways once those conditions change. Over long rollouts, small errors accumulate until the trajectory drifts away from physically plausible behavior; under an intervention on a physical parameter, the model may continue to follow the law seen during training rather than the intervened one. We show that these two failures require different structural remedies. Evolving a learned energy with a symplectic integrator preserves the geometry of the conservative dynamics and keeps rollouts bounded and physically meaningful for up to $100\times$ the training horizon, while equal-capacity predictors, an energy-regularized predictor, and a tuned neural ODE diverge. By contrast, encoding the physical coupling through an explicit linear factorization enables the model to follow a never-seen sign of that coupling, whereas an unrestricted parameterization remains locked to the training law. Crucially, the two mechanisms are separable: removing the structure responsible for long-horizon stability leaves counterfactual transfer intact, while removing the factorized coupling destroys counterfactual transfer without eliminating stability. This double dissociation, established with matched controls that remove or replace one structural component at a time, persists beyond the headline three-body system and remains visible when the physical state must be inferred from pixels rather than provided directly. The result is a concrete design principle for physical world models: long-horizon stability and changed-law generalization arise from distinct structural commitments, and each can be imposed deliberately without requiring the other.

View source

Similar papers

#machine learning Preprint Sep 2026

Stable and Counterfactually Robust Physical World Models from Imposed Structure and Learned Physics

A world model learns to forecast how a physical system evolves from recorded trajectories, yet the systems it imitates obey physical laws that are neither fully supplied nor reliably respected. The model may create energy, drift or diverge over long rollouts, and answer a changed law query using the law observed during...

Yu-Feng Wang, Parivesh Priye, Lu Wei et al. · 0 citations
#artificial intelligence Preprint Sep 2026

World Models with Predictable Long-Horizon Marginals

Three properties of a world model are distinguished: the distribution it approaches, the rate of approach, and the conditional dynamics it learns, which derive an absolute convergence bound from finite initialization banks and control departure from the reference through conditional action-space divergence.

Yu-Hao Du, Shu-Nian Chen · 0 citations
Preprint Sep 2026

Compact but Moving: Intervention-Relevant Geometry in Recurrent World Models

Results show that compact intervention structure can persist as a moving, state-dependent local geometry embedded in high-dimensional recurrent dynamics, without implying a fixed or dynamically closed low-dimensional state.

Yu-Ming Chen, Yang Liu · 0 citations
Preprint Aug 2026

Counterfactual Quotient Models: Learning What Actions Change, Not What the World Does

This work introduces the Counterfactual Quotient Model, which treats action-conditioned futures as equivalent when they differ only by a component shared across actions, and establishes the decision sufficiency, identifiability, common-mode invariance, approximation behavior, and regret properties of the resulting repr...

Jun-Lin Chen, Rui-Jie Wang, Jian-Xin Li · 1 citation
Preprint Aug 2026

CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift

CausalNav, a controller built around a signed, action-conditioned transition graph over identified state coordinates, is studied with CausalNav, a controller built around a signed, action-conditioned transition graph over identified state coordinates, which attains the best average rank.

Yi-Yao Zhang, Diksha Goel, Hussain Ahmad et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.