Skip to content

CruxBench: A Benchmark of Information Discovery

Sep 2026 · 0 citations · 59 references
Computer Science

TL;DR

CruxBench is introduced, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed crux updates beliefs about a target forecasting question and finds that VOI correlates highly with independent measures of model capability and captures cruxes'usefulness for answering target questions.

Abstract

Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions -- which we call cruxes -- whose answers provide key steps on the path toward solving the target problem. To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed crux updates beliefs about a target forecasting question. CruxBench enjoys a rare combination of three key properties: it is (1) contamination-resistant by construction, since ground truth is generated by future world events; (2) open-ended, admitting unbounded and complex text-based submissions rather than one correct numeric answer; and (3) grounded, with informativeness measured against quantified changes in real-world beliefs. We evaluate a diverse set of eight models on 293 target forecasting questions and find that VOI correlates highly with independent measures of model capability (r=0.90) and captures cruxes'usefulness for answering target questions. However, information discovery remains challenging even for frontier LLMs, which only narrowly outperform a random-timing baseline.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Evaluating and Benchmarking the System One Model Jev

Jev answers MMLU's calculation-heavy questions more accurately than other MMLU questions (94% vs. 91%), whereas both open models, and all three on C-Eval, find them harder.

Tobias Deußer, L. Sparrenberg, R. Sifa · 4 citations
Preprint Aug 2026

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.

V. Rodionov, Shamil Assylbekov · 0 citations
#artificial intelligence Preprint Sep 2026

WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation

This work introduces an automated pipeline for generating synthetic datasets of emails reflecting realistic workplace scenarios, along with long- and short-form questions and gold answers grounded in the data, and underscores the importance of realistic, high-complexity evaluation data for developing stronger real-worl...

Amey Varhade, Ananya Sutradhar, Ravishankar Krishnaswamy et al. · 0 citations
#artificial intelligence Preprint Aug 2026

PROOF: Profiling Reliability of Object-Level Facts in Large Language Models

Aggregate factuality scores hide where a language model succeeds, which relations it confuses, and whether an answer survives innocuous changes to the question or decoder. We introduce PROOF, a profile-oriented benchmark for factual coverage in instruction-tuned language models. PROOF converts a frozen Wikidata snapsho...

A. Chetvergov, Mikhail Solovev, Timofei Sivoraksha et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Path Matters: Evaluating Small Language Models Beyond Answer Accuracy in KGQA

This work isolates SLM graph reasoning from navigation and traceability capability by employing the THESEUS navigation and traceability framework and using frozen, off-the-shelf SLMs as local action policies, and results motivate evaluating SLM graph reasoning beyond endpoint accuracy alone.

Eduin E. Hernandez, Sergio A. Diaz, Luis F. Garcia et al. · 0 citations
#artificial intelligence Preprint Sep 2026

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

This work systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026, using staged screening and automated full-text coding to examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.

Chao Wang · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.