Causal Abstraction Learning for Multi-Modal Grounded Planning (CALM) is proposed, a framework that enhances planning agents with the ability to discover and exploit causal regularities across tasks.
Xin-Shu Li, Shiyi Yang, Ziqi Xu et al.· Proceedings of the 32nd ACM...· 0 citations
Offline reinforcement learning (RL) is a useful approach for recommender systems because it can optimize long-term user feedback from logged interaction data without online exploration. A key challenge is the multi-modal nature of user preferences: a user may like several unrelated item types, so a unimodal policy (for example, a Gaussian) tends to average across modes and generate actions that do not match any interest. Recent diffusion-based policies can model complex preference distributions, but they often require many denoising steps. We propose PerfRec (Preference-aware Flow for Recommendation), a flow-matching offline RL framework that learns an expressive behavioral policy and distills it into an efficient one-step policy. PerfRec (i) trains a conditional flow model to clone the logged action distribution, (ii) trains twin Q-networks using next actions sampled from the learned flow policy, and (iii) trains an advantage-conditioned one-step policy with Q-guidance for improvement and a distillation loss that keeps the policy close to the flow policy. We use binary advantage conditioning to separate high-advantage and low-advantage regions of the flow-induced action distribution, so that at inference we can sample from the high-advantage mode with a single forward pass. Experiments on five benchmark datasets and one online simulation platform show that PerfRec improves recommendation performance over strong offline RL baselines.
Self-supervised Causal Effects Estimation is proposed, a novel framework that integrates causal priors with self-supervised learning to construct balanced and predictive representations for causal effects estimation that consistently outperforms state-of-the-art methods.
Xin-Shu Li, Shiyi Yang, Venus Haghighi et al.· ACM Transactions on Intellig...· 0 citations
Recent advances in multimodal embodied agents have enabled long-horizon planning in visually rich environments via natural language. Yet, their generalization remains brittle when task instructions deviate from familiar examples, exposing a reliance on surface imitation rather than structural understanding. We propose Causal Abstraction Learning for Multi-Modal Grounded Planning (CALM), a framework that enhances planning agents with the ability to discover and exploit causal regularities across tasks. CALM incrementally develops a causal library by abstracting precondition–effect structure from successful executions, yielding compact representations that emphasize stable dependencies beyond incidental context. When execution diverges from expectation, these abstractions are refined through contrastive causal reasoning, enabling targeted adjustments that resolve underlying mechanism mismatch. The resulting structure serves as a transferable prior for planning in novel settings, integrating perceptual cues with mechanism-informed knowledge. Without retraining or task-specific heuristics, CALM generalizes robustly and efficiently to linguistic and perceptual variation. Experiments on ALFRED and VirtualHome demonstrate consistent gains, highlighting causal abstraction as a scalable inductive bias for grounded planning.
Xinshu Li, Shiyi Yang, Ziqi Xu et al.· Proceedings of the 32nd ACM...· 0 citations
This work presents RegionSLM, a region-aware SLM designed to explicitly connect the question to its supporting regions, and curates ReDoc, a region-supervised corpus with 105k documents and 350k question-answer pairs obtained via a question-guided two-step filtering procedure.
Chao Wang, Hehe Fan, Huichen Yang et al.· Annual International ACM SIG...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.