Skip to content
Open access

From Rubrics to Recipe: Principle-Centric Benchmark for Evaluating Large Language Models

2026 · Proceedings of the Workshop on Evaluating Evaluations (EvalEval) · pp. 82-99 · 0 citations · 14 references

TL;DR

This study proposes a framework to automatically extract and generate task-level principles to automatically extract and generate task-level principles for data generation and evaluation, enabling controllable data creation and fine-grained, interpretable assessment of LLM capabilities.

Abstract

Large language models (LLMs) are often evaluated on benchmarks that rely on surface-level instructions, obscuring what defines high-quality performance. We argue that tasks can be more precisely characterized through principles : human-readable rules that specify what matters for a good response to the task. Our study proposes a framework to automatically extract and generate task-level principles for data generation and evaluation. Using this approach, we build a benchmark of over 20K principle-aligned instances, enabling controllable data creation and fine-grained, interpretable assessment of LLMs. Experiments show that principles both improve output quality and scale evaluation beyond manual curation, offering a new recipe for principled assessment of LLM capabilities. 1

Read PDF

Similar papers

Open access

Evaluation and Distillation of Source Code Generation Tasks by Large Language Models

Two novel contributions are introduced: CodeEval and CodeQual, an open-source execution framework that provides researchers with a ready-to-use evaluation pipeline for evaluating and improving LLMs in software engineering contexts, encompassing both functional correctness assessment and subjective code quality evaluation.

Danny Brahman · 0 citations
Preprint Jul 2026

Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench

PredicateLongBench is proposed, a benchmark that stress-tests long-context reasoning by asking models to identify the longest contiguous subsequence of words in a long input that satisfies given predicates/constraints drawn from a broader predicate class.

Siddhartha Jain, A. Velingker · 0 citations
#small language model Preprint Aug 2026

StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models

This work proposes StrategyBench, which selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility, and experiments show that explicit strategy utility differs substantially across task categories and depends on both strategy generation and execution conditions.

Jinghan Tan, Yuanzhe Wang, Lu Chen et al. · 0 citations
#artificial intelligence Preprint Aug 2026

ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation

The results suggest a novel way of approaching automated evaluation, by offering a faster, more explainable, and less ambiguous alternative to black-box rubric evals, particularly in high-stakes domains such as healthcare and banking where precision and auditability are critical.

Kaustubh D. Dhole, Charles L. A. Clarke, E. Agichtein · 0 citations
#software testing Preprint Aug 2026

XREPOTEST: Benchmarking Multilingual Repository-Level Unit Test Generation for Large Language Models

XREPOTEST is introduced, a multilingual repository-level benchmark for unit test generation spanning five underexplored languages: Rust, Go, Julia, PHP, and Ruby, and Invocation Rate is proposed to assess whether generated tests meaningfully exercise the intended functionality.

L. Dung, Dong Cao Van, Nam Le Hai et al. · 0 citations

Capacity vs. architecture: an evaluation of SLMs for automated docstring generation

A reproducible, human-validated evaluation framework applied to 13 strategies—four architectural families crossed with four reasoning variants crossed with four reasoning variants—across three SLMs spanning 3B–14B parameters, plus targeted ablations.

Balaji Venktesh, Amsaprabhaa M, G. Sundaram · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.