Step-level reasoning evaluators are commonly based on autoregressive language models, whose causal attention restricts each step representation to the problem, previous steps, and the current step. Yet, when the complete solution is available, the validity of an earlier step may become clearer only through its downstre...
Yi-Ming Feng, Naihao Deng, Yu-Long Chen et al.· 0 citations
It is revealed that the energy divide persists across models, hardware, and tasks, suggesting a systemic energy inequity in multilingual LLM deployment and is recommended that the community treat energy as a first-class evaluation axis, extend reporting checklists and model cards to include it, and adopt deployment-sid...
It is argued that BBQ-style multiple-choice abstention benchmarks measure a single structural cue, and a model that solves them does not thereby become fair, and is called for evaluation suites that cover a broader spectrum of fairness alignment.
Naihao Deng, Samee Arif, Shuai-Chen Chang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.