Background: Large language models are increasingly used for rubric-based grading, but it remains unclear whether they adapt selectively when criterion weights change while the response, criterion definitions, and total score remain fixed. This study examined whether alternative point allocations produce targeted criterion-level adaptation or broader grading instability. Methods: A controlled paired design was applied to 1000 student responses, comprising 500 technical and 500 argumentative answers. Eight local open-weight LLMs evaluated every response under two analytic rubrics totaling 10 points. One point was transferred from Clarity to Completeness for technical responses and from Clarity to Dialecticality for argumentative responses. Of 16,000 expected evaluations, 15,999 were structurally valid, yielding 7999 complete cross-rubric pairs. Model outputs were analyzed for total-score shifts, affected-criterion adaptation, stability of unaffected criteria, model–human alignment, and correspondence with differences between two human evaluation conditions. Results: All models assigned lower mean scores under Rubric B, with mean shifts ranging from −1.239 to −0.182 points. Adaptation mechanisms differed substantially across models and response types. gemma3:4b frequently preserved technical total scores through compensating criterion changes, whereas llama3.1:8b showed extensive spillover into unaffected criteria. The Qwen models generally produced smaller total-score reductions and greater stability in unchanged dimensions. The mean score was 0.637 points higher under the human Rubric B condition than under the human Rubric A condition, although the two conditions were applied by different evaluator pairs; every model shifted negatively, and the lowest overall shift error was obtained by qwen3:4b at 1.498 points. Conclusions: Rubric sensitivity did not consistently imply localized or criterion-consistent adaptation. Reliable evaluation of LLM graders therefore requires separate analysis of total scores, affected criteria, unaffected criteria, human alignment, and cross-condition shift correspondence.
Cătălin Anghel, A. Anghel, Adina Cocu et al.· Informatics· 0 citations
Background: Evaluation of LLM-based automated grading often relies on comparison with a single human score, which can obscure meaningful variability among raters of open-ended answers. This study introduces ConsensusGrade, a consensus-aware framework that treats the human reference as a scoring envelope rather than as a single point. Methods: We analyzed 1000 open-ended student answers from 100 students across 10 questions, each graded by four evaluators. Six previously generated and aligned automated grading configurations from GradeAgentOps were compared with the four-rater human reference. The score sets were generated using Llama 3.3 70B Instruct as the primary grader, with Qwen 2.5 14B Instruct for semantic repair. Results: Human evaluators showed meaningful agreement, with ICC(A,1) = 0.712, but exact four-rater agreement occurred in only 2.2% of records. Broad score dispersion occurred in 59.0%. All automated configurations showed negative bias relative to the human median. FULL achieved 68.5% inside-envelope positioning and a chance-adjusted score of 0.454; under the central-trimmed envelope, this rate decreased to 34.3%, while configuration ordering was preserved. Conclusions: ConsensusGrade provides a diagnostic framework for interpreting automated scores relative to observed human variability; inside-envelope rates should not be interpreted as stand-alone measures of grading accuracy.
Cătălin Anghel, A. Anghel, Mihai Vlase et al.· Applied System Innovation· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.