This work examines repair hallucination in final patches and understanding hallucination in intermediate artifacts through three tasks, namely triggering testcase identification, line coverage prediction, and additional testcase generation, showing that both repair and understanding hallucinations remain prevalent.
Xue-Meng Cai, Jia-Kun Liu, Lin-Han Yang et al.· 0 citations
Prototype pollution is a critical class of taint-style vulnerabilities in JavaScript programs, enabling attackers to tamper with object prototypes and thereby alter program behavior in unexpected and often dangerous ways. Despite its severity, existing detection techniques struggle with excessive false positives and po...
De-Zhen Kong, Pei-Sen Yao, Jia-Kun Liu et al.· ACM Transactions on Software...· 0 citations
Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents'implementation capability to produce correct code edits from detailed specifications. However, p...
Yun Peng, Zi-Han Wu, Ze-Yang Zhuang et al.· 0 citations
VulAgentRL is proposed, an agentic RL framework for interprocedural vulnerability detection built on a Code Property Graph (CPG), which outperforms state-of-the-art baselines, including frontier models, on the strict pair-wise-correct metric while issuing fewer tool calls, and its advantage persists on an out-of-distri...
Yi-Kun Li, Ting Zhang, Jia-Kun Liu et al.· arXiv.org· 2 citations
DyRetriever is an efficient context retrieval method via partial dependency graphs that uses an LLM to first select a set of entry-point functions and then perform multi-hop reasoning along the code dependency graph, eliminating manually designed rules and enabling flexibility across scenarios.
Zhongxin Liu, Zhonghao Jiang, Zhi-Fan Ye et al.· 1 citation
Optimo is proposed, a multi-level LLM-based code optimization approach built on a novel Mixture-of-Prompts (MoP) architecture that achieves up to 57.48% opt%, and consistently outperforms the best baseline by up to 96.51% in terms of opt%.
Yun Peng, Jun Wan, Jiakun Liu et al.· arXiv.org· 0 citations
A high-quality benchmark of 1,000 code refinement instances from 328 Python, Java, and JavaScript repositories that focused on one of the most challenging code refinement scenarios that strictly requires repository-level knowledge reasoning, and a straightforward method, RepoRefiner, which retrieves repository-level co...
Ke Wang, Peng Lan, Jia-Kun Liu et al.· ACM Transactions on Software...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.