Skip to content

Analyzing and fixing code generation errors of foundation large language models

Sep 2026 · Empirical Software Engineering · Vol 32 · 0 citations · 56 references

TL;DR

The goal is to understand the code generation errors of foundation LLMs and explore the solution to resolve directly fixable errors, and to design and evaluate the LlmFix fixing method and constructed the LlmErrorEval dataset.

View source

Similar papers

Open access Aug 2026

Improving Bug Detection in LLM-Generated Unit Tests: Revisiting Test-Oracle Reliability Across Modern Large Language Models

This paper presents a formal mathematical model for categorizing the outcome of generated-tests into four classes, a couple of basic metrics: Bug-Revealing Rate (BRR) and Bug-Validating Rate (BVR); and two basic statistical tests to ensure that the results are rigorous.

Zeyad Farooq Lutfi · 0 citations
#natural language process... Preprint Sep 2026

Retrofitting Code Using LLMs to Support Exceptional Behavior

The results demonstrate that EXCODER provides an effective, though imperfect, solution to this problem in automated code generation, offering developers the first way to implement ERC following test-driven development.

Ling-Tao Zhong, Jiyang Zhang, Jayanth Srinivasa et al. · 0 citations
#natural language process... Preprint Sep 2026

Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?

Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in problem solving or bug fixing independently, but do not explore the relationship between thes...

Alexandru Stefan Stoica, Traian Rebedea, M. Mihăescu · 0 citations

Hunk-Constrained DPO: Segment-Level Optimization for Secure and Correct LLM Code Generation

Hunk-Constrained Direct Preference Optimization is introduced, a training framework that unifies security hardening and functional correction in large language models and demonstrates that HPO achieves substantial security improvements—up to 28 percentage points—while preserving or enhancing functional correctness.

Qian-Shuo Huang, Xin Yin, Xin-Rui Li et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Interpreting and Steering for Safe and Correct Code Generation

DuoSteer is proposed, a double-steering approach that simultaneously applies safety and code-correctness steering to attention heads and outperforms not only other steering variants but also prompting and supervised fine-tuning baselines for inference-time vulnerability reduction.

Hao Yan, Zi-Yu Yao · 0 citations
Open access Sep 2026

Goanna: a novel approach for automated type error debugging

Goanna is introduced, a novel type checker for Haskell that focuses on improving error diagnostics, and shows performance constraints when diagnosing large programs containing complex errors, but remains responsive enough to provide real-time debugging assistance for small to medium-sized programs.

Shuai Fu, Tim Dwyer, Peter James Stuckey et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.