By shifting RL from token-level exploration to experience-level reasoning, CodeSkill improves optimization efficiency and long-horizon behavioral coherence, highlighting the effectiveness of explicit behavioral abstraction for scalable agentic code generation.
Song-Li Wu, Jing-Yi Wang, Zhao-Cheng Du et al.· 0 citations
MCPGen is introduced, an executable benchmark for Model Context Protocol (MCP) workflow development that evaluates three diagnostic tasks: workflow reconstruction, tool creation, and backward-compatible workflow extension and evaluates 11 representative LLMs in a single-turn foundation-model setting.
Yingxuan Yang, Jia-Qi Liu, Li-Rui Guan et al.· 0 citations
The proposed REALM framework models long-term memory as a continual lifecycle by autonomously organizing memories into a heterogeneous cognitive graph, retrieving evidence via adaptively composed graph-search atoms, and continually reconsolidating memories based on retrieval feedback.
Yuan-Yi Song, Yukai Wang, Xin-Bei Ma et al.· 2 citations
The key observation is that although expert trajectories are scarce, high-quality final artifacts such as literature reviews, analyst reports and legal judgments, are abundant in pre-training data and can be viewed as compressed traces of the evidence-seeking processes that produced them.
Junjie Huang, Jiarui Qin, Di Yin et al.· 0 citations
This work presents SAGE (Skill with Attribution-Guided Evolution), a deployed framework that learns, attributes, evolves, and routes directing knowledge from expert demonstrations, and releases PROSE, the first public dataset pairing screenplays with storyboards by professional directors across 68 episodes.
Mao-Lin Ran, Xiaoyan Lu, Jia-Qi Liu et al.· 0 citations
This work introduces Harness-R1, the first method, to the authors' knowledge, that makes failure-conditioned, lifecycle-wide editing of an existing executable runtime a learned capability, and post-trains a dedicated harness engineer with online reinforcement learning so that its edits are optimized for the realized ta...
Shuai Shao, Kangning Zhang, Qingyao Li et al.· 10 citations· ⚡2
A training-free, model-agnostic memory graph that repairs continuity entirely at the prompt-text layer by extracting a multi-dimensional state for every shot, retrieving only causally prior context over the resulting graph, filtering it selectively, and injecting the surviving constraints by natural-language prompt rew...
Jiaqi Liu, Maolin Ran, Xiaoyan Lu et al.· 0 citations
SkillGate lifts a 9B policy from 40.8% to 53.2% trial success, well ahead of the identical budget spent on outcome reward alone, while cutting exposure to misleading candidates by two thirds and reading fewer skills.
Qingyao Li, Wenxiang Jiao, Shuai Shao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.