Skip to content

An Empirical Study of Counterfactual Self-Explanations in LLMs

Sep 2026 · 0 citations · 22 references
Computer Science

TL;DR

This work evaluates ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales and shows that model scale is the strongest determinant of explanation quality.

Abstract

Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior. We study this issue through counterfactual self-explanations, where a model minimally edits an input so that its own prediction changes. Across sentiment analysis and natural language inference, we evaluate ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales. Our results show that model scale is the strongest determinant of explanation quality: larger models are substantially more likely to generate counterfactuals that flip their own predictions and target decision-relevant evidence. In contrast, the rationale-guided condition produces edit-minimal counterfactuals that are also more human-aligned. However, it does not consistently improve faithfulness. Overall, counterfactual self-explanations can provide useful behavioral evidence about model decisions, but their reliability depends strongly on model capacity and should be empirically validated rather than assumed.

View source

Similar papers

Preprint Aug 2026

Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

This work introduces CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits, enabling us to evaluate and improve methods for explaining LLM behaviors.

Adam Karvonen, Euan Ong, Subhash Kantamneni et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Counterfactual Tests for Measuring Chain-of-Thought Faithfulness in Visual Language Models

The analysis shows that CoTs do not reliably track visual evidence that influences model predictions, and it is found that Predict-then-Explain explanations align more strongly with perturbation-induced probability shifts than pre-answer CoTs, while binary vCT scores are often nearly saturated.

Bayar Menzat, Max Süss, Rui-Zhi Wang et al. · 0 citations
#artificial intelligence Conference Open access Sep 2026

Self-Reports Are Not Verification

An environment-grounded audit is introduced in which every intermediate proposal receives an exact outcome in an evolutionary Contexto search whose feedback function assigns every valid guess an exact rank without human annotation.

En-Rong Pan, Ryan Zhou, Ting Hu · 0 citations
#natural language process... Preprint Sep 2026

Chronologic: Measuring Language Models'Ability to Represent the Past

Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct an...

Ted Underwood, Zi-Liang Qiu, Sarah Griebel et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning

Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verifi...

Xing Han, Yu-Xin Wang, Chen Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Strangers to Themselves: What Language Models Say About Themselves Is Generic

Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a prediction test. Across nine behavioral evaluations, we measure how a model behaves under different...

Phil Blandfort, Urja Pawar · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.