Skip to content

Author

Cristian Sandu

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

RubricAdapt-LLM: Measuring Criterion-Level Adaptation and Score Shifts Under Alternative Grading Rubrics

Background: Large language models are increasingly used for rubric-based grading, but it remains unclear whether they adapt selectively when criterion weights change while the response, criterion definitions, and total score remain fixed. This study examined whether alternative point allocations produce targeted criterion-level adaptation or broader grading instability. Methods: A controlled paired design was applied to 1000 student responses, comprising 500 technical and 500 argumentative answers. Eight local open-weight LLMs evaluated every response under two analytic rubrics totaling 10 points. One point was transferred from Clarity to Completeness for technical responses and from Clarity to Dialecticality for argumentative responses. Of 16,000 expected evaluations, 15,999 were structurally valid, yielding 7999 complete cross-rubric pairs. Model outputs were analyzed for total-score shifts, affected-criterion adaptation, stability of unaffected criteria, model–human alignment, and correspondence with differences between two human evaluation conditions. Results: All models assigned lower mean scores under Rubric B, with mean shifts ranging from −1.239 to −0.182 points. Adaptation mechanisms differed substantially across models and response types. gemma3:4b frequently preserved technical total scores through compensating criterion changes, whereas llama3.1:8b showed extensive spillover into unaffected criteria. The Qwen models generally produced smaller total-score reductions and greater stability in unchanged dimensions. The mean score was 0.637 points higher under the human Rubric B condition than under the human Rubric A condition, although the two conditions were applied by different evaluator pairs; every model shifted negatively, and the lowest overall shift error was obtained by qwen3:4b at 1.498 points. Conclusions: Rubric sensitivity did not consistently imply localized or criterion-consistent adaptation. Reliable evaluation of LLM graders therefore requires separate analysis of total scores, affected criteria, unaffected criteria, human alignment, and cross-condition shift correspondence.

Cătălin Anghel, A. Anghel, Adina Cocu et al. · 0 citations
Open access Aug 2026

GradeDrift-LLM: Measuring Student-History-Induced Score Drift in LLM-Based Automated Grading

Background: Large language models (LLMs) are increasingly explored for automated educational assessment, while future educational platforms may combine grading, feedback, learner analytics, and personalization. The objective of this study was to determine whether student-history metadata can influence the numerical score assigned to the same answer. Methods: This study introduces GradeDrift-LLM, a controlled framework for measuring student-history-induced score drift in LLM-based automated grading. We evaluated 1000 Computer Science answers from 100 students across six student-history conditions and eight open-weight LLMs. For each grading instance, the submitted answer, question, reference answer, rubric-related information, scoring scale, and grading instruction were kept constant; only the student-history condition varied. Results: Across 39,997 valid paired comparisons, 83.92% showed no drift, 9.40% showed upward drift, and 6.68% showed downward drift. Mean absolute drift was 0.2137 points, and the 95th percentile absolute drift was 1 point. Positive-history frames tended to increase scores, whereas negative-history frames tended to decrease them. Drift was model-dependent, not uniformly explained by approximate scale, and present in both technical and argumentative answers; rare extreme deviations reached 10 points. Conclusions: Student-history metadata can influence LLM-generated grading scores despite explicit instructions to ignore it. Future LLM-based grading systems should separate answer-based scoring from learner-context-based personalization and validate score invariance under controlled learner-context variations.

Cătălin Anghel, A. Anghel, M. Craciun et al. · 0 citations
Review Open access Jul 2026

SafetyJudge-LLM: Auditing Local Open-Weight LLMs as Semantic Safety Judges for Boundary-Failure Detection

Background: LLM-as-a-judge workflows are increasingly used to evaluate open-ended model outputs, but the judge model can itself become a source of error in safety assessment. SafetyJudge-LLM audits local open-weight LLMs as semantic safety judges. Methods: This study reused a fixed set of previously reviewed safety-boundary responses and their hidden reference labels. Two independent human evaluations (R1 and R2) quantified reference-layer ambiguity. Seven local open-weight judge models were evaluated under a common Ollama inference protocol. A paired C6 sensitivity analysis reran llama3.2:3b and qwen3:8b through Hugging Face Transformers. Results: The final judge-output matrix contained 10,612 retained outputs. R1–R2 agreement was 95.45% (Cohen’s κ = 0.612) overall but 47.80% (κ = 0.341) in secondary cases. Several judge models detected more than 90% of confirmed safety-boundary failures, but high detection was not always accompanied by low false-unsafe behavior on control cases. Output-format reliability also varied across models: overall label parseability was 98.11%, while strict JSON schema compliance was 92.55%. The llama3.2:3b schema-failure rate persisted across engines (52.06% under Ollama; 59.60% under Transformers), whereas qwen3:8b maintained complete compliance. Conclusions: SafetyJudge-LLM shows that local open-weight LLMs can support semantic safety judging, but their reliability must be evaluated across multiple dimensions.

Cătălin Anghel, A. Anghel, M. Craciun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.