Skip to content
Book Open access

Causal Abstraction Learning for Multi-Modal Grounded Planning

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 2790-2801 · 0 citations · 26 references

Abstract

Recent advances in multimodal embodied agents have enabled long-horizon planning in visually rich environments via natural language. Yet, their generalization remains brittle when task instructions deviate from familiar examples, exposing a reliance on surface imitation rather than structural understanding. We propose Causal Abstraction Learning for Multi-Modal Grounded Planning (CALM), a framework that enhances planning agents with the ability to discover and exploit causal regularities across tasks. CALM incrementally develops a causal library by abstracting precondition–effect structure from successful executions, yielding compact representations that emphasize stable dependencies beyond incidental context. When execution diverges from expectation, these abstractions are refined through contrastive causal reasoning, enabling targeted adjustments that resolve underlying mechanism mismatch. The resulting structure serves as a transferable prior for planning in novel settings, integrating perceptual cues with mechanism-informed knowledge. Without retraining or task-specific heuristics, CALM generalizes robustly and efficiently to linguistic and perceptual variation. Experiments on ALFRED and VirtualHome demonstrate consistent gains, highlighting causal abstraction as a scalable inductive bias for grounded planning.

Read PDF

Similar papers

Conference Jul 2026

SPaRL: Spatially-aware Reinforcement Learning with Language

Large Language Models (LLMs) have demonstrated strong generalization and reasoning capabilities across a wide range of domains, including embodied decision making and robotics. Despite this progress, existing reinforcement-based embodied agents often struggle with novel tasks requiring complex spatial reasoning and multi-object manipulation when relying solely on the robot’s egocentric view. In this paper, we propose SPaRL (Spatially-aware Reinforcement Learning with Language), a novel framework that explicitly integrates a structured 3D scene graph into an LLM-based reinforcement learning policy. By representing objects and their spatial relationships in a compact, semantically meaningful graph and conditioning it on task instructions, SPaRL provides the policy with explicit relational context beyond raw visual inputs. The scene graph is dynamically pruned to retain instruction-relevant objects and relations, serialized into natural language, and jointly processed with visual observations and task descriptions by a frozen LLM backbone. We evaluate SPaRL on language-conditioned rearrangement tasks in Habitat. Results show that incorporating an instruction-conditioned scene graph consistently improves performance over a vision-only LLM policy. Across curriculum training, SPaRL achieved improved performance on the Language Rearrangement benchmark. These results suggest that explicit spatial representations provide useful inductive bias for object reasoning, particularly as task difficulty increases.

Jeyoung Lee, Jaewon Lee, J. Oh et al. · 0 citations
Aug 2026

Towards a Causally-inspired Evolving World Model for Vision-and-Language Navigation in Continuous Environments.

Vision-and-Language Navigation in Continuous Environments (VLN-CE) has emerged as a pivotal challenge in Embodied AI, requiring an agent to navigate 3D spaces guided by natural language instructions. Drawing inspiration from human cognition, world models provide a powerful paradigm by predicting environment dynamics and enabling reasoning beyond immediate observations. However, existing world model-based VLN methods remain static once trained - their representations rely on fixed correlationbased priors rather than adaptive causal structures, making them unable to accommodate evolving confounders and changing observation-action dependencies across environments. This rigidity leads to overfitting to training-specific patterns and degraded performance under distribution shifts. To address this limitation, we propose a causally-inspired evolving world modeling framework that formulates VLN-CE as a sequence of causal partially observable Markov decision processes. Our model learns unified latent states that integrate vision, language, and action, while addressing spurious correlations through a dual-level intervention mechanism: at the observation level, frequency-domain perturbations simulate superficial appearance variations to enhance perceptual robustness; at the representation level, cross-episode confounder buffers perform counterfactual substitution to approximate the influence of latent confounding factors. Beyond static world modeling, our framework continuously evolves, refining these proxy representations across episodes, enabling efficient adaptation to previously unseen environments. Building on this evolving causally-inspired foundation, our world model supports counterfactual reasoning and strengthens generalization across diverse navigation contexts. Extensive evaluations on established VLN-CE benchmarks demonstrate that our method outperforms existing approaches, delivering superior navigation performance across diverse scenarios. Real-world robot evaluations further validate the practicality of our approach. Code is available in the Supplementary Material.

Xuan Yao, Junyu Gao, Changsheng Xu · 0 citations
Preprint Aug 2026

ParallelWorld: Test-Time Scaling for Embodied Reasoning

This work proposes ParallelWorld, a multi-horizon test-time scaling framework for embodied reasoning that empowers agents to simulate and evaluate multi-step future trajectories in parallel before committing to an action.

Min Chen, Shengjun Zhang, Yuxin Li et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Neurosymbolic Embodied Agents

A neurosymbolic agent that factors long-horizon household tasks into task-directed visual exploration and constrained symbolic planning and evaluates executable continuations using a domain-independent planning heuristic is presented.

Mohammad Albinhassan, Yuming Feng, Alessandra Russo et al. · 0 citations
Dec 2025

Embodied Tree of Thoughts: Deliberate Manipulation Planning With Embodied World Model

Embodied Tree of Thoughts (EToT), a novel Real2Sim2Real planning framework that leverages a physics-based interactive digital twin as an embodied world model, is validated on a suite of short- and long-horizon manipulation tasks, where it consistently outperforms baselines by effectively predicting physical dynamics and adapting to potential failures.

Wenjiang Xu, Mingkan Zhang, Cindy Wang et al. · 2 citations
Preprint Aug 2026

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

XEWorld is introduced, a controlled cross-embodiment testbed for world models that isolates embodiments by evaluating held-out robots within physically identical scenes, highlighting that achieving true cross-embodiment generalization requires architectural innovations that decouple visual appearance from underlying physical dynamics.

Yixiang Chen, Jiabing Yang, Yuan Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.