The goal is to understand the code generation errors of foundation LLMs and explore the solution to resolve directly fixable errors, and to design and evaluate the LlmFix fixing method and constructed the LlmErrorEval dataset.
This paper presents a formal mathematical model for categorizing the outcome of generated-tests into four classes, a couple of basic metrics: Bug-Revealing Rate (BRR) and Bug-Validating Rate (BVR); and two basic statistical tests to ensure that the results are rigorous.
Zeyad Farooq Lutfi· Al-Noor Journal of Engineeri...· 0 citations
The results demonstrate that EXCODER provides an effective, though imperfect, solution to this problem in automated code generation, offering developers the first way to implement ERC following test-driven development.
Ling-Tao Zhong, Jiyang Zhang, Jayanth Srinivasa et al.· 0 citations
Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in problem solving or bug fixing independently, but do not explore the relationship between thes...
Alexandru Stefan Stoica, Traian Rebedea, M. Mihăescu· 0 citations
Hunk-Constrained Direct Preference Optimization is introduced, a training framework that unifies security hardening and functional correction in large language models and demonstrates that HPO achieves substantial security improvements—up to 28 percentage points—while preserving or enhancing functional correctness.
Qian-Shuo Huang, Xin Yin, Xin-Rui Li et al.· ACM Transactions on Software...· 0 citations
DuoSteer is proposed, a double-steering approach that simultaneously applies safety and code-correctness steering to attention heads and outperforms not only other steering variants but also prompting and supervised fine-tuning baselines for inference-time vulnerability reduction.
Goanna is introduced, a novel type checker for Haskell that focuses on improving error diagnostics, and shows performance constraints when diagnosing large programs containing complex errors, but remains responsive enough to provide real-time debugging assistance for small to medium-sized programs.
Shuai Fu, Tim Dwyer, Peter James Stuckey et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.