Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
AutoSciRub is presented, an evaluation-first framework that induces a task-specific executable rubric before research execution and uses it to guide execution, criterion-level verification as well as iterative revision.