Skip to content

Author

Rommel Gutierrez

We have 2 of 37 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

EduFairBench: reproducible evaluation of large language models for educational assessment

Large language models (LLMs) are increasingly used to evaluate open-ended educational responses. However, their performance is often assessed using aggregate metrics that provide limited insight into prediction stability, uncertainty, error patterns, and feedback quality. This study presents EduFairBench, a reproducible evaluation protocol designed to characterize LLM behavior across short-answer assessment and automated essay scoring using open educational benchmarks. The protocol combines repeated inference, majority-vote consolidation, uncertainty estimation, error analysis, and structural evaluation of generated feedback within a unified experimental framework. Experiments were conducted on SciEntsBank, Beetle, and ASAP2, comprising 2,000 student responses and 10,000 independent LLM inferences. The results showed moderate predictive agreement with human assessment while revealing substantial differences between nominal and ordinal evaluation tasks. Repeated inference demonstrated high internal stability across benchmarks, although systematic errors remained in semantically adjacent categories, indicating that prediction consistency does not necessarily imply correctness. Feedback quality varied by task type, with longer textual contexts yielding more specific and pedagogically structured explanations. These findings demonstrate that evaluating educational LLMs requires complementary analyses beyond conventional performance metrics. EduFairBench provides a reproducible methodology for jointly analyzing predictive performance, robustness, uncertainty, and feedback quality, providing a comprehensive methodological framework for the rigorous evaluation of LLM-based educational assessment systems.

W. Villegas-Ch., Aracely Mera-Navarrete, Fernando Zúñiga-Tello et al. · 0 citations
Open access Jul 2026

Cross-model evaluation of phishing detectors against LLM-generated emails

This work reframes cross-model phishing detection from a problem of model incompatibility to one of practical calibration, and provides two deployable solutions, threshold recalibration on a small target slice and aggregated-pool training, along with a publicly released multi-LLM corpus.

Rommel Gutierrez, W. Villegas-Ch., Jaime Govea · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.