The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs
A diagnostic benchmark for open-domain, open-form, long-horizon counterfactual causal reasoning, containing 220 what-if questions across STEM, HSS, and Hybrid scenarios, and finds that WhatIfBench remains far from saturated: even the strongest model reaches only a 64.62% final score.
Yucheng Wang, Yuetian Du, Zheng Liu et al.
· 0 citations