From Rubrics to Recipe: Principle-Centric Benchmark for Evaluating Large Language Models
This study proposes a framework to automatically extract and generate task-level principles to automatically extract and generate task-level principles for data generation and evaluation, enabling controllable data creation and fine-grained, interpretable assessment of LLM capabilities.