Skill-based LLM agents select reusable procedures from an external library to solve complex tasks, yet their routing decisions rely entirely on text-level signals such as task descriptions, verbal reflections, and experience-derived rules, while the model's own internal representational state remains unobserved. Recent...
Bo Lin, He-Jia Geng, Xinyi Xie et al.· 2 citations
Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback...
Fang Wu, Dan-Lei Xing, Yan-Jie Huang et al.· 0 citations
LatticeMind is presented, a conflict-aware structured memory that handles contradiction at write time, which maintains explicit item status, applies cheap symbolic conflict checks, and invokes LLM reconciliation only for unresolved semantic cases.
Heng Zhou, Lian Zhang, Yutao Fan et al.· 0 citations
This work argues RSI must next operate across real, diverse scientific, engineering, and meta-scientific domains, not where formal evaluation is merely tractable, and presents MetaRSI-v1, where improvement is the scheduled composition of three typed operators over one unified paradigm.
Zi-Hang Tan, Lei-Xin Sun, Zi-Tong Shi et al.· 1 citation
Experiments across diverse models, benchmarks, and agent harnesses show that supervised fine-tuning on SKT-generated trajectories consistently improves skill-use performance, establishing verified data synthesis as an effective and scalable approach for skill-use training.
Zelin Tan, Yi-Qun Zhang, Hao Li et al.· 2 citations
We propose Process-Aware Policy Optimization (PAPO), a method that integrates process-level evaluation into Group Relative Policy Optimization (GRPO) through decoupled advantage normalization, to address two limitations of existing reward designs. Outcome reward models (ORM) evaluate only final-answer correctness, trea...
Zelin Tan, Zhouliang Yu, Bo-Cheng Lin et al.· arXiv.org· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.