Skip to content

Bencmarking Generative AI for Debugging Support in Program Education

Sep 2026 · Mendeley Data

Abstract

Abstract—Programming activities, such as automatic debugging and program repair, are increasingly being supported by Generative Artificial Intelligence (GenAI). However, a rigorous empirical evaluation of the efficacy of various large language models (LLMs) across several debugging phases is necessary. Using a benchmark of 35 programming issue instances across seven categories—syntax, logic, runtime, boundary, type, performance, and security—this study experimentally assesses the debugging capabilities of ChatGPT and Gemini. Bug detection, localization, diagnosis, repair, compilation, execution, explanation, root-cause identification, and test-case validation were all used to assess each model. The findings demonstrate that ChatGPT and Gemini both attained a 100% success rate in every assessed dimension. No additional introduced defects were found, and all generated corrections passed compilation, execution, and test-case validation. Additionally, no performance difference was found between the two models in any of the seven bug areas that were examined. These results show that, within the parameters of the assessed benchmark, both models showed consistent debugging performance. However, the findings' generalizability is restricted by the small number of test cases and the lack of failure cases. Therefore, larger and more varied benchmarks with more complicated programming bugs and other programming languages should be evaluated in future studies. Keywords—Generative AI, Large Language Models, Program Debugging, Automated Program Repair, ChatGPT, Gemini, Software Engineering.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#computer vision Review Mar 2008

Agile methods in European embedded software development organisations: a survey on the actual use and usefulness of Extreme Programming and Scrum

The results show that the embedded industry has been able to apply agile methods in its development processes and that the appreciation of the agile methods and their individual practices appears to increase once adopted and applied in practice.

O. Salo, P. Abrahamsson · 238 citations · ⚡9
#computer vision Open access Jul 2017

What happens when software developers are (un)happy

Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 236 citations · ⚡13
#computer vision Open access Oct 2004

Mobile-D: an agile approach for mobile application development

The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.

P. Abrahamsson, Antti Hanhineva, H. Hulkko et al. · 225 citations · ⚡18

Related blog posts

MIT News · Artificial Intelligence Sep 14, 2026

New method enables AI for safety-critical situations

The “HardFlow” algorithm could help generative AI models produce high-quality outputs that obey strict requirements when “pretty close” doesn’t cut it.

GPT-Lab Sep 10, 2026

Responsible AI Must Consider Its Afterlife

AI may appear weightless, but every model depends on physical infrastructure. To understand responsible AI, we need to look beyond algorithms and consider the entire lifecycle of the hardware behind them. The post Responsible AI Must Consider Its Afterlife appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.