LLM agents that invoke external tools face critical safety vulnerabilities when malicious manipulations exploit their implicit trust in tool outputs and metadata. However, identifying these vulnerabilities through testing is challenging due to the need to bypass safety guardrails with semantically legitimate inputs, th...
Yu-Chen Shao, Zi-Qun Bao, Yu-Heng Huang et al.· 0 citations
Recently, Diffusion Large Language Models (dLLMs) have demonstrated unique efficiency advantages, enabled by their inherently parallel decoding mechanism and flexible generation paradigm. Meanwhile, despite the rapid advancement of Search Agents, their practical deployment is constrained by a fundamental limitation, te...
Jiahao Zhao, Shaoxuan Xu, ZhongXiang Sun et al.· Annual International ACM SIG...· 0 citations
This work introduces early outcome prediction, a complementary axis of efficiency that instead cuts cost within each task within each task, and instantiates EarlyEval, a lightweight framework that trains a pair of LightGBM success and failure classifiers over behavioral, textual, and reference-solution features and hal...
This work identifies a phenomenon termed token-level memorization asymmetry through theoretical analysis of diffusion training dynamics and proposes Q-Skew, a quantile-weighted skewness-based indicator for membership inference on finetuned DLMs.
Shengfang Zhai, Leo Marchyok, Yu-Ling Shi et al.· 0 citations
DeepRepoQA is proposed, a novel question answering (QA) framework for repository-level code understanding that builds on an agentic framework where LLM agents find answers through a systematic tree search over the repository structure.
Wei Peng, Yu-Ling Shi, Yingwei Ma et al.· 0 citations
LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone to factual errors caused by insufficient repository understanding. Recent methods attempt to mitigate this limitation through pre-repair repository exploration; however, their fix-driven strategies expl...
ParaTempo is a training-free asynchronous parallel reasoning framework driven by temporal confidence, a branch-local measure of answer-space convergence that reduces average latency, and exhibits stronger temporal stability and predictive power for future branch convergence than token-level and instantaneous signals.
Xuteng Zhang, Wenhao Zeng, Xiao-Dong Gu et al.· 1 citation
SWE-Pruner Pro is proposed, which prunes tool outputs directly inside the agent, with a small head turns the agent's own internal representations into a keep-or-prune label for each line, with a length-aware embedding keyed to each tool output's line count.
Yuhang Wang, Yuling Shi, Shaoqiu Zhang et al.· 2 citations
SWE-Bench ProMax is introduced, an expert-curated, multilingual code refactoring benchmark of 170 instances drawn from real commits across seven programming languages, which presents a meaningful and unsaturated challenge for current AI coding agents.
Yu-Ling Shi, Jing-Heng Xu, Kelin Fu et al.· 6 citations
Self-Reflective Policy Optimization (SRPO) enables LLMs to analyze their own completed trajectories, synthesize errors into concise"reflection patches," and use reflection-conditioned teacher scores on student on-policy rollouts as dense token-level training signals.
Jialong Liu, Yu-Ling Shi, Ning Yang et al.· 1 citation
Repo0 is presented, a continuous structural evolution framework for zero-to-all code generation that maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation.
Si-Lin Chen, Haoyi Teng, Xiao-Dong Gu et al.· 2 citations
SkillForge is proposed, a self-distillation framework that proactively acquires project-specific knowledge from the repository itself that substantially improves downstream software issue resolution.
Si-Lin Chen, Han Li, Xiao-Dong Gu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.