Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation
These results favor paired audits of current outputs over judge-only release decisions or transported calibration, while a task-solvability prediction reverses sign across domains.