The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore...
Tingyu Qu, Wei-Gao Sun, Yuecheng Liu et al.· 0 citations
Omni-Decision is presented, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict.
Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The standard recipe assigns a single outcome reward at the end and compares trajectories sampled for the same task. As a result, a group with no successful trajectory yields...
A statistical framework for quantitatively assessing LLM reliability under multiple evaluation criteria and shows that the proposed framework can identify reliability differences between different LLMs is reasonably robust to variations in indicator weights.
Yi Zhu· Advances in Engineering Inno...· 0 citations
Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world requests are often open-ended, casually specified, and context-dependent, requiring models not only to follow explicit instructions but als...
Yong Peng, Qing-Shui Gu, Li-Ya Zhu et al.· 0 citations
This work proposes AgentSpec, a speculative decoding algorithm that addresses the limitations of existing methods for LLM agents and incorporates structure-isolated drafting that constrains speculation to semantically coherent segments of the agent workflow, reducing the drafts of irrelevant semantic paths and achievin...
Xin Wang, Zi-Ming Miao, Yi Zhu et al.· 0 citations
ADSD is introduced, which is the first prompt-suffix attack that collapses verifier acceptance by pushing draft probability mass toward tokens the target is unlikely to accept, and successfully generates highly effective adversarial suffixes.
Run-Min Wang, Chaoyi Zhou, Xi Liu et al.· arXiv.org· 0 citations
By pairing an interactive function-calling sandbox with evidence-based verification, MobilePA-Bench serves as both a practical diagnostic benchmark and an interactive foundation for agentic reinforcement learning---accelerating the development of dependable mobile agents.
Yi Zhu, Xiong-Wei Wu, Qiyi Wang et al.· 1 citation
This work systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains and establishes StartupBench as an empirical measure of progress toward E2E completions o...
Li-Ya Zhu, Xin Ma, Tao Liu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.