Skip to content

A Simple Baseline for Learning Approximate State Abstractions in Factored State Spaces

· 0 citations · 32 references

TL;DR

This work introduces a surprisingly simple neural network architecture change: a learnable, state-independent attention mask applied to the inputs of the policy and value networks and trained end-to-end using only the RL objective.

View source

Similar papers

From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning

LA-MAML (Language Adapted MAML), which modifies the inner loop by adapting the global policy parameters in a single step through a learned embedding of the task instruction, replacing the inner loop trajectory collection and gradient-based updates.

Garvit Singla, U. M. Natarajan, Raghuram Bharadwaj Diddigi · 0 citations
#artificial intelligence Preprint Aug 2026

Q-Learning With World Models

This work proposes QWM, a framework that leverages world models to perform test-time search over imagined trajectories on top of Q-learning to select high-value actions during both online rollouts and evaluation, and significantly outperforms strong prior state-of-the-art methods on both sample efficiency and performance.

Perry Dong, Yueru Jia, Chelsea Finn et al. · 0 citations
Jul 2026

Constrained Reinforcement Learning Using Successor Representations

The Safe Deep Successor Representation is proposed, a novel method that allows quick retraining of policies towards new cost structures and is competitive on a simple navigation task while being considerably more flexible.

Michaela Girstl, Alexander Mattick, Christopher Mutschler · 0 citations
#machine learning Preprint Aug 2026

PAC: Progress-Augmented Advantage Curriculum for Multi-Task Reinforcement Learning of LLMs

Reinforcement learning (RL) is used to improve the reasoning abilities of LLMs, while training data span heterogeneous tasks. However, most RL post-training pipelines rely on fixed or manually designed task mixtures, even though task usefulness changes as training progresses. Online curriculum methods often define learnability by update magnitude, ignoring whether the update translates into reward gains, which can misallocate rollout budget toward tasks with large but ineffective updates. We propose PAC, a Progress-Augmented Advantage Curriculum for multi-task RL of LLMs that combines two task-level signals: advantage-derived learnability, which measures the magnitude of the policy update a task can induce, and recent reward gains, which show whether those updates have improved task performance. A Bayesian Thompson Sampling controller uses these signals to allocate rollouts across tasks during GRPO training. We evaluate PAC under two settings: a multi-level reasoning setting and a multi-domain reasoning setting. PAC improves sample efficiency and final performance: it reaches comparable validation scores with fewer rollout steps and achieves higher final averages than random sampling and advantage-based curriculum baselines in both settings. These results show that jointly tracking advantage signals and actual reward gains yields an effective online curriculum for LLM post-training.

Yuan-Qiang Yu, Yan-Zhao Zheng, Zhen-Tao Zhang et al. · 0 citations
Preprint Aug 2026

State2State: Environment-Derived Mid-Training for LLM Agents

State2State is proposed, an environment-derived mid-training method that converts explored environment states into training objectives, challenging agents to reach a specified target state by deriving tasks from environment exploration and verifying success through rule-based state matching.

Xuanyu Lei, Yiqi Zhu, Chenliang Li et al. · 1 citation
Preprint Jul 2026

Cross-Benchmark Generalization in Long-Horizon Agents

Desc descriptive evidence is provided that long-horizon multi-tool post-training can change ways of working that transfer beyond its training domain, and both software-engineering benchmarks improve despite the training collection containing no software-engineering tasks.

Sushant Mehta, Logan Ritchie, Liudas Panavas et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.