Skip to content

Compositional Reasoning in Language Models under Reinforcement Learning Post-Training

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

A dependency-graph framework is proposed to formalize compositional reasoning, yielding three levels of compositionality with increasing complexity, and preliminary evidence that the decomposed-to-composed asymmetry can extend to practical settings is presented.

Abstract

Compositional reasoning is critical for real-world problem solving: since training data is necessarily limited, models must generalize by composing learned skills in new ways. While post-training methods such as reinforcement learning (RL) have substantially improved the reasoning abilities of language models (LMs), their effects on compositional reasoning remain less well understood. We propose a dependency-graph framework to formalize compositional reasoning, yielding three levels of compositionality with increasing complexity. Empirically, we instantiate this framework with data-structure tasks, which provide deterministic reward computation and clear compositional structure. We find a consistent decomposed-to-composed asymmetry: decomposed-skill training does not reliably transfer to composed tasks, whereas composed-task training transfers more readily back to decomposed tasks. We provide theoretical explanation for this asymmetry, and further evaluate compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. Finally, we present a pilot study on real-world tool-calling benchmarks, showing preliminary evidence that the decomposed-to-composed asymmetry can extend to practical settings.

View source

Similar papers

Review Sep 2026

Reinforcement Learning Post-Training for Reasoning Large Language Models: Methods, Systems, and Evaluation

Reinforcement learning (RL) has become a central post-training approach for reasoning and agentic large language models (LLMs), particularly when task outcomes can be verified automatically. Comparisons across this literature remain difficult because a reported gain may combine changes to the learning signal, policy co...

Liu Yang, Han Zhu, Zheng-Yang Zhong et al. · 1 citation
#natural language process... Preprint Aug 2026

SPEAR: Distilling Domain-Adaptive Reasoning Skeletons via Sequential Symbolic Alignment in Reinforcement Learning

This work introduces SPEAR (Symbolic Process Evaluation and Alignment Reward), a training-free and plug-and-play process reward method for sequence-level on-policy distillation that effectively bridges the reasoning gap between student and teacher models via sequence-level distillation with efficient dense process rewa...

Zhuo-Chun Li, Yuelyu Ji, Yiming Zeng et al. · 0 citations
#machine learning Preprint Aug 2026

Learning Generalizable Behaviors for Terminal Agents

River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization is proposed, which achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks.

Yi-Fan Yao, Bo Pang, Xuan-Phi Nguyen et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Extremely Sparse Supervision Incentivizes Reasoning Ability

Overall, the results challenge the assumption that effective post-training must be token-intensive and point to a new direction for understanding and designing more efficient post-training algorithms.

Zhi-Shuai Liu, Xingzi Xu, Mehmet Saygin Seyfioglu et al. · 2 citations
#machine learning Preprint Sep 2026

Learning to Optimize through Solver-Grounded Self-Play

Results indicate that with zero curated data, OPT-Zero matches state-of-the-art data-dependent methods while exhibiting substantially stronger generalizability, establishing self-play training as a highly scalable paradigm for advancing LLM reasoning in modeling and solving optimization problems.

Xia Jiang, Yao-Xin Wu, Chen-Yu Zhou et al. · 0 citations

Delayed Convergence and Emergent CoT Reliance in RL-Tuned Language Models

This analysis reveals that intermediate log-probability is an unreliable indicator for reasoning capability; instead, reasoning performance results from a shift of internal confidence allocation where RL fine-tuning delays internal convergence, exhibiting prolonged mid-layer exploration before converging sharply at the...

Pablo Pérez-Lázaro, Rocío Aznar-Gimeno, F. J. Lacueva-Pérez et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.