Skip to content
Preprint

PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models

Aug 2026 · 2 citations · 31 references
Computer Science

TL;DR

PhyLatent, a dynamics-relevant training objective for JointEmbedding Predictive Architecture (JEPA) world models, shows that global non-collapse alone is insufficient for learning a reliable JEPA worldmodel state space.

Abstract

We propose PhyLatent, a dynamics-relevant training objective for JointEmbedding Predictive Architecture (JEPA) world models. Our key observation is that preventing global latent collapse does not ensure that a representation preserves physical states and action consequences. We identify three failure modes in JEPA world models: physical invariance collapse, physical identifiability collapse, and counterfactual dynamics collapse. PhyLatent addresses them through three training pathways: physical invariance, physical identifiability, and counterfactual dynamics, implemented with physical state grounding, future representation alignment, static visual invariance, counterfactual branch separation, and latent denoising. On OGBench-Cube, PhyLatent reduces the three failure rates from 15.60%, 6.71%, and 8.41% to 7.53%, 0.95%, and 4.62%, respectively, and improves model predictive control (MPC) success from 70.0% to 78.1%. With the same architecture and planner, it further improves success from 81.0% to 98.0% on TwoRooms and remains competitive on Reacher and PushT. These results show that global non-collapse alone is insufficient for learning a reliable JEPA worldmodel state space.

View source

Similar papers

Preprint Aug 2026

Is Forward Prediction Enough? Physical State Grounding for JEPA World Models

PSG-JEPA is proposed, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes.

Haodong Yan, Jiaguang Zhu, Ming-Ming Jia et al. · 2 citations
Preprint Jul 2026

Branch-JEPA: Finite-Support Predictive Distributions for JEPA World Models

Branch-JEPA is introduced, which replaces this point-valued transition with a context-weighted finite set of latent successors, and preserves more distinct futures, while full-set scoring improves the quality of the resulting predictive distribution.

Zhi Song, Ximing Xing, Zhenchao Tang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Flow-JEPA: Flow Matching for Robust Latent Dynamics in JEPA World Models

This work proposes Flow-JEPA (F-JEPA), a conditional flow matching dynamics model that jointly generates a sequence of future latent states conditioned on the current observation and actions, suggesting that conditional flow matching provides a promising alternative to deterministic autoregressive dynamics in JEPA world models.

Yan-Chen Huo, Zi-Ying Song, Yadan Luo · 0 citations
#artificial intelligence Preprint Aug 2026

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

Contrastive inverse dynamics thus provides a distribution-free anti-collapse signal that requires no target network, stop-gradient, pretrained encoder, or reconstruction objective, and it is argued that the anti-collapse pressure can instead come from the transition data itself.

Jack Boylan, Chris Hokamp · 0 citations
#machine learning Preprint Aug 2026

JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics

This work introduces the cross-predictive JEPA (JEPA-x), which grounds latent dynamics in privileged physical trajectories, and shows that direct physical-state regression improves decodability without improving forecastability or control, indicating that the benefit comes from shaping latent dynamics rather than merely encoding physical variables.

Kehan Wen, Ziming Li, Siyuan Luo et al. · 0 citations
Preprint Aug 2026

Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency

Action-Conditioned Predictive Consistency (ACPC), a diagnostic that measures how far a clean history and a visually perturbed view of it diverge after being rolled forward under the same action sequence, is introduced and it is proved that this divergence bounds the perturbation-induced change in multi-step prediction error and planner cost.

Guo An, Zijing Wu, Hongzhuang Dong et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.