Are We Measuring Scientific Intelligence? Rethinking the Evaluation of AI Scientists
AI agents can now carry out data-driven scientific analyses end to end, and benchmarks assess them by giving an agent a question and a dataset and scoring its final answer against a fixed key. These benchmarks assume that a correct answer was derived from the supplied data, a property we call evidence grounding. Howeve...