GraphForge, an evidence-graph based framework that grounds both the task and its verification in real files, and rejects fine-tuning on the SFT model's own rollouts yields further improvements on all three benchmarks, suggesting that the rubrics provide a useful selection signal.
Qi-Sheng Su, Han-Chen Wang, Guang Zhu et al.· 0 citations
This work proposes AsyncTool, a benchmark for assessing LLM-based agents in interactive multi-task tool-use environments with delayed tool feedback, and constructs a diverse asynchronous multitasking dataset that covers multiple scenarios and tool-use patterns.
Video-DR is introduced, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval, enabling autonomous exploration that breaks the imitation-learning ceiling.
VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought, demonstrates that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.
Haodong Li, Tianfei Ren, Xiaoxiao Ma et al.· arXiv.org· 8 citations
These results establish source-intent preservation and shared executable-state grounding as key principles for scalable terminal-task synthesis as key principles for scalable terminal-task synthesis.
Kou Shi, Zun Wang, Qi-Sheng Su et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.