Abstract—Programming activities, such as automatic debugging and program repair, are increasingly being supported by Generative Artificial Intelligence (GenAI). However, a rigorous empirical evaluation of the efficacy of various large language models (LLMs) across several debugging phases is necessary. Using a benchmark of 35 programming issue instances across seven categories—syntax, logic, runtime, boundary, type, performance, and security—this study experimentally assesses the debugging capabilities of ChatGPT and Gemini. Bug detection, localization, diagnosis, repair, compilation, execution, explanation, root-cause identification, and test-case validation were all used to assess each model. The findings demonstrate that ChatGPT and Gemini both attained a 100% success rate in every assessed dimension. No additional introduced defects were found, and all generated corrections passed compilation, execution, and test-case validation. Additionally, no performance difference was found between the two models in any of the seven bug areas that were examined. These results show that, within the parameters of the assessed benchmark, both models showed consistent debugging performance. However, the findings' generalizability is restricted by the small number of test cases and the lack of failure cases. Therefore, larger and more varied benchmarks with more complicated programming bugs and other programming languages should be evaluated in future studies. Keywords—Generative AI, Large Language Models, Program Debugging, Automated Program Repair, ChatGPT, Gemini, Software Engineering.
Muhammad adrevi zaki rukmana, Sasmoko Sasmoko, Novtryananda M.S Ghunu et al.· Mendeley Data· 0 citations
Abstract—Programming activities, such as automatic debugging and program repair, are increasingly being supported by Generative Artificial Intelligence (GenAI). However, a rigorous empirical evaluation of the efficacy of various large language models (LLMs) across several debugging phases is necessary. Using a benchmark of 35 programming issue instances across seven categories—syntax, logic, runtime, boundary, type, performance, and security—this study experimentally assesses the debugging capabilities of ChatGPT and Gemini. Bug detection, localization, diagnosis, repair, compilation, execution, explanation, root-cause identification, and test-case validation were all used to assess each model. The findings demonstrate that ChatGPT and Gemini both attained a 100% success rate in every assessed dimension. No additional introduced defects were found, and all generated corrections passed compilation, execution, and test-case validation. Additionally, no performance difference was found between the two models in any of the seven bug areas that were examined. These results show that, within the parameters of the assessed benchmark, both models showed consistent debugging performance. However, the findings' generalizability is restricted by the small number of test cases and the lack of failure cases. Therefore, larger and more varied benchmarks with more complicated programming bugs and other programming languages should be evaluated in future studies. Keywords—Generative AI, Large Language Models, Program Debugging, Automated Program Repair, ChatGPT, Gemini, Software Engineering.
Muhammad adrevi zaki rukmana, Sasmoko Sasmoko, Novtryananda M.S Ghunu et al.· Mendeley Data· 0 citations
Abstract—Programming activities, such as automatic debugging and program repair, are increasingly being supported by Generative Artificial Intelligence (GenAI). However, a rigorous empirical evaluation of the efficacy of various large language models (LLMs) across several debugging phases is necessary. Using a benchmark of 35 programming issue instances across seven categories—syntax, logic, runtime, boundary, type, performance, and security—this study experimentally assesses the debugging capabilities of ChatGPT and Gemini. Bug detection, localization, diagnosis, repair, compilation, execution, explanation, root-cause identification, and test-case validation were all used to assess each model. The findings demonstrate that ChatGPT and Gemini both attained a 100% success rate in every assessed dimension. No additional introduced defects were found, and all generated corrections passed compilation, execution, and test-case validation. Additionally, no performance difference was found between the two models in any of the seven bug areas that were examined. These results show that, within the parameters of the assessed benchmark, both models showed consistent debugging performance. However, the findings' generalizability is restricted by the small number of test cases and the lack of failure cases. Therefore, larger and more varied benchmarks with more complicated programming bugs and other programming languages should be evaluated in future studies. Keywords—Generative AI, Large Language Models, Program Debugging, Automated Program Repair, ChatGPT, Gemini, Software Engineering.
Muhammad Adrevi zaki, Sasmoko Sasmoko, Adele Mailangkay et al.· Zenodo (CERN European Organi...· 0 citations
Abstract—Programming activities, such as automatic debugging and program repair, are increasingly being supported by Generative Artificial Intelligence (GenAI). However, a rigorous empirical evaluation of the efficacy of various large language models (LLMs) across several debugging phases is necessary. Using a benchmark of 35 programming issue instances across seven categories—syntax, logic, runtime, boundary, type, performance, and security—this study experimentally assesses the debugging capabilities of ChatGPT and Gemini. Bug detection, localization, diagnosis, repair, compilation, execution, explanation, root-cause identification, and test-case validation were all used to assess each model. The findings demonstrate that ChatGPT and Gemini both attained a 100% success rate in every assessed dimension. No additional introduced defects were found, and all generated corrections passed compilation, execution, and test-case validation. Additionally, no performance difference was found between the two models in any of the seven bug areas that were examined. These results show that, within the parameters of the assessed benchmark, both models showed consistent debugging performance. However, the findings' generalizability is restricted by the small number of test cases and the lack of failure cases. Therefore, larger and more varied benchmarks with more complicated programming bugs and other programming languages should be evaluated in future studies. Keywords—Generative AI, Large Language Models, Program Debugging, Automated Program Repair, ChatGPT, Gemini, Software Engineering.
Muhammad Adrevi zaki, Sasmoko Sasmoko, Adele Mailangkay et al.· Zenodo (CERN European Organi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.