A Case Study on the Characteristics of Writing-Evaluation Outputs Generated by ChatGPT-5.2 : A Comparison of Prompt Conditions Using the Web Interface
Abstract
This exploratory study analyzes the characteristics of writing-evaluation outputs generated by ChatGPT-5.2, a generative AI model, from the perspective of Learning with AI. Rather than treating instructor evaluations as definitive “correct answers,” the study uses them as comparative benchmarks to examine the agreement between the instructors’ consensus ratings and those produced by ChatGPT-5.2, the consistency and variability of the model’s evaluations, and the distributional characteristics of its ratings across different prompt configurations. The dataset consisted of 60 first drafts of argumentative essays collected from a university writing course at K University. Consensus ratings were established by three instructors, each with more than ten years of experience in writing assessment. ChatGPT-5.2 was then tested under ten progressively configured prompt conditions (G1-G10), which combined six prompt components: question, input, example, response format, context, and instruction. Each manuscript was independently evaluated three times under each condition, with the session reset to a stateless condition before every run. The resulting high, medium, and low ratings, scores on a 10-point scale, and evaluation rationales were collected and analyzed.The results showed that the average agreement between the instructors’ consensus ratings and ChatGPT-5.2’s evaluations was 55.3% for categorical ratings and 55% for numerical scores. A certain proportion of cases also exhibited divergent judgments across repeated evaluations of the same manuscript. Although the overall rating distributions tended to remain within a relatively stable range despite changes in prompt configuration, different combinations of prompt components produced varying tendencies toward greater severity or leniency. The case analysis further revealed differences between the instructors and ChatGPT-5.2 in their interpretation of the evaluation context, application of the same rubric, sensitivity to discriminatory expressions and their effects on readers, and determination of the scope of expression-related criteria. Based on these findings, the study proposes a preliminary revision of the instructors’ assessment criteria.The findings suggest that ChatGPT-5.2 should not be reduced to an automated scoring tool but should instead be employed as an external reference through which instructors can reflect on and adjust their assessment criteria. Rather than using prompt design merely to make AI-generated assessments more closely resemble human judgments, this study conceptualizes writing assessment in the era of Learning with AI as a process in which human evaluators construct and reconsider their own criteria through interaction with generative AI. In doing so, the study contributes to both theoretical and practical discussions of writing assessment in the Learning with AI era.