Skip to content

Author

Muhammad adrevi zaki rukmana

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Dataset Open access Sep 2026

Bencmarking Generative AI for Debugging Support in Program Education

Abstract—Programming activities, such as automatic debugging and program repair, are increasingly being supported by Generative Artificial Intelligence (GenAI). However, a rigorous empirical evaluation of the efficacy of various large language models (LLMs) across several debugging phases is necessary. Using a benchmark of 35 programming issue instances across seven categories—syntax, logic, runtime, boundary, type, performance, and security—this study experimentally assesses the debugging capabilities of ChatGPT and Gemini. Bug detection, localization, diagnosis, repair, compilation, execution, explanation, root-cause identification, and test-case validation were all used to assess each model. The findings demonstrate that ChatGPT and Gemini both attained a 100% success rate in every assessed dimension. No additional introduced defects were found, and all generated corrections passed compilation, execution, and test-case validation. Additionally, no performance difference was found between the two models in any of the seven bug areas that were examined. These results show that, within the parameters of the assessed benchmark, both models showed consistent debugging performance. However, the findings' generalizability is restricted by the small number of test cases and the lack of failure cases. Therefore, larger and more varied benchmarks with more complicated programming bugs and other programming languages should be evaluated in future studies. Keywords—Generative AI, Large Language Models, Program Debugging, Automated Program Repair, ChatGPT, Gemini, Software Engineering.

Muhammad adrevi zaki rukmana, Sasmoko Sasmoko, Novtryananda M.S Ghunu et al. · 0 citations
#large language models Dataset Open access Sep 2026

Bencmarking Generative AI for Debugging Support in Program Education

Abstract—Programming activities, such as automatic debugging and program repair, are increasingly being supported by Generative Artificial Intelligence (GenAI). However, a rigorous empirical evaluation of the efficacy of various large language models (LLMs) across several debugging phases is necessary. Using a benchmark of 35 programming issue instances across seven categories—syntax, logic, runtime, boundary, type, performance, and security—this study experimentally assesses the debugging capabilities of ChatGPT and Gemini. Bug detection, localization, diagnosis, repair, compilation, execution, explanation, root-cause identification, and test-case validation were all used to assess each model. The findings demonstrate that ChatGPT and Gemini both attained a 100% success rate in every assessed dimension. No additional introduced defects were found, and all generated corrections passed compilation, execution, and test-case validation. Additionally, no performance difference was found between the two models in any of the seven bug areas that were examined. These results show that, within the parameters of the assessed benchmark, both models showed consistent debugging performance. However, the findings' generalizability is restricted by the small number of test cases and the lack of failure cases. Therefore, larger and more varied benchmarks with more complicated programming bugs and other programming languages should be evaluated in future studies. Keywords—Generative AI, Large Language Models, Program Debugging, Automated Program Repair, ChatGPT, Gemini, Software Engineering.

Muhammad adrevi zaki rukmana, Sasmoko Sasmoko, Novtryananda M.S Ghunu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.