Abstract—Programming activities, such as automatic debugging and program repair, are increasingly being supported by Generative Artificial Intelligence (GenAI). However, a rigorous empirical evaluation of the efficacy of various large language models (LLMs) across several debugging phases is necessary. Using a benchmark of 35 programming issue instances across seven categories—syntax, logic, runtime, boundary, type, performance, and security—this study experimentally assesses the debugging capabilities of ChatGPT and Gemini. Bug detection, localization, diagnosis, repair, compilation, execution, explanation, root-cause identification, and test-case validation were all used to assess each model. The findings demonstrate that ChatGPT and Gemini both attained a 100% success rate in every assessed dimension. No additional introduced defects were found, and all generated corrections passed compilation, execution, and test-case validation. Additionally, no performance difference was found between the two models in any of the seven bug areas that were examined. These results show that, within the parameters of the assessed benchmark, both models showed consistent debugging performance. However, the findings' generalizability is restricted by the small number of test cases and the lack of failure cases. Therefore, larger and more varied benchmarks with more complicated programming bugs and other programming languages should be evaluated in future studies. Keywords—Generative AI, Large Language Models, Program Debugging, Automated Program Repair, ChatGPT, Gemini, Software Engineering.
Muhammad adrevi zaki rukmana, Sasmoko Sasmoko, Novtryananda M.S Ghunu et al.· Mendeley Data· 0 citations
Generative AI (GenAI) tools have been adopted across all levels of computer science education, from K-12 through graduate study, offering learners opportunities to improve learning efficiency and support creative problem solving. However, this rapid adoption has outpaced critical evaluation: students demonstrate high acceptance of AI-generated outputs but significant difficulty correcting AI-generated code compared to instructor-designed tasks, an asymmetry that reflects a deeper structural gap where students are taught to use GenAI tools, but not to critically evaluate or correct what these tools produce. Existing learning theories such as constructivism, sociocultural theory, and connectivism presuppose human-centered epistemic agency and do not account for GenAI's role in simulating reasoning or co-constructing meaning with learners. At the institutional level, GenAI governance has been active but limited in scope: policy analysis shows that guidelines remain largely prescriptive and output-focused, emphasizing academic integrity, privacy, and security while offering little direction on how students should reason through, reflect on, or take responsibility for AI-assisted decisions. This rapid review synthesizes findings across GenAI adoption, human-AI collaboration, and institutional governance in computer science education to clarify the conditions under which human oversight should occur. We propose a framework, organized around task classification, output verification, confidence checking, correction and revision, and learning reflection that operationalizes a human-in-the-loop approach to GenAI use, positioning learners and educators as the final authority over AI-generated outputs rather than passive recipients of them.
Wilson Feraldo, Sasmoko Sasmoko, Arao Ternorio Alves dos Santos et al.· Mendeley Data· 0 citations
Abstract—Programming activities, such as automatic debugging and program repair, are increasingly being supported by Generative Artificial Intelligence (GenAI). However, a rigorous empirical evaluation of the efficacy of various large language models (LLMs) across several debugging phases is necessary. Using a benchmark of 35 programming issue instances across seven categories—syntax, logic, runtime, boundary, type, performance, and security—this study experimentally assesses the debugging capabilities of ChatGPT and Gemini. Bug detection, localization, diagnosis, repair, compilation, execution, explanation, root-cause identification, and test-case validation were all used to assess each model. The findings demonstrate that ChatGPT and Gemini both attained a 100% success rate in every assessed dimension. No additional introduced defects were found, and all generated corrections passed compilation, execution, and test-case validation. Additionally, no performance difference was found between the two models in any of the seven bug areas that were examined. These results show that, within the parameters of the assessed benchmark, both models showed consistent debugging performance. However, the findings' generalizability is restricted by the small number of test cases and the lack of failure cases. Therefore, larger and more varied benchmarks with more complicated programming bugs and other programming languages should be evaluated in future studies. Keywords—Generative AI, Large Language Models, Program Debugging, Automated Program Repair, ChatGPT, Gemini, Software Engineering.
Muhammad adrevi zaki rukmana, Sasmoko Sasmoko, Novtryananda M.S Ghunu et al.· Mendeley Data· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.