RLVR has substantially improved the reasoning capabilities of LLMs. However, existing methods typically parameterize temporal progression in the Markov Decision Process by token-by-token generation, despite the highly non-uniform information flow along autoregressive trajectories. In this paper, we propose InfoPPO, whi...
Yong-Cheng Zeng, Xin-Yu Cui, Yan Song et al.· 0 citations
Biological swimmers and flyers exploit unsteady vortices for propulsion, whereas engineered vehicles usually suppress them as disturbances. Learning such flow exploitation in machines is difficult because real-fluid interaction data are scarce and unstructured exploration is unstable in high-dimensional, history-depend...
Fei Han, Xin-Yu Cui, Zhi-Peng Wang et al.· 0 citations
This work proposes R$^2$-MAD (Remember and Reweight for Multi-Agent Debate), a framework that equips agents with an experience memory accumulated from past debates that achieves consistent improvements over existing single-agent and MAD baselines.
Xuan-Fa Jin, Zhijian Ma, Yong-Cheng Zeng et al.· 0 citations
Inspired by the human brain, which balances plasticity and stability through complementary episodic storage and gradual consolidation, UniMem is proposed, a self-routing framework for autonomous memory management that consistently outperforms baselines while maintaining execution fidelity.