Accuracy evaluation of ChatGPT, Gemini, and Perplexity in refining Arabic writing structures based on linguistic error taxonomy
Abstract
The growing demand for Arabic as a professional language, along with the rapid advancement of Artificial Intelligence (AI) technologies such as ChatGPT, Gemini, and Perplexity, has led to their widespread use among students for translating and composing Arabic assignments. However, these AI tools are often used without critically evaluating the accuracy of their output. Rather than serving merely as learning aids, AI applications are increasingly overused, creating excessive dependence that may ultimately hinder the achievement of Arabic-language-learning objectives. This study aimed to examine the capability and accuracy of ChatGPT, Gemini, and Perplexity in improving the grammatical structure of Arabic writing. This study employed a descriptive quantitative approach to measure and describe the performance of the three AI systems in refining Arabic text structures. The research data consisted of one Arabic composition (insyā’) text, which was revised using ChatGPT, Gemini, and Perplexity. The revised outputs were evaluated by three Arabic language experts using an assessment rubric, and the resulting scores were descriptively analyzed to compare the accuracy levels of the three AI systems. The findings revealed that ChatGPT achieved the highest accuracy score (96.5%), followed by Gemini (75%), while Perplexity obtained the lowest score (63%). These results indicate that ChatGPT demonstrates superior and more consistent performance in improving the grammatical structure of Arabic writing compared to Gemini and Perplexity. Therefore, ChatGPT is considered more suitable as a supporting tool for learning Arabic grammar. Nevertheless, AI-generated corrections should be critically reviewed by users, particularly in formal academic and professional contexts.