Skip to content
Conference Open access

I-EDI: Robust Self-Evolution Agents via Verifiable Counterfactual Simulation

Runze Fan Yong Li
Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · pp. 90-98 · 0 citations · 39 references

TL;DR

This work argues that robust evolution implies Structural Invariance: a reasoning path is valid only if its core dependency graph remains isomorphic under counterfactual perturbations, and enforces this with Verifiable Counterfactual Simulation (VCS).

Abstract

Current self-evolution methods mostly chase recall, they learn from any trajectory that ends with the right answer, even when the path relies on a "lucky guess". Such reasoning appears valid in-distribution but often fails the moment the task shifts slightly. We argue that robust evolution implies Structural Invariance: a reasoning path is valid only if its core dependency graph remains isomorphic under counterfactual perturbations. I-EDI enforces this with Verifiable Counterfactual Simulation (VCS). Instead of treating data generation as mere "sample and filter," we treat it as an intervention. Unlike standard rejection sampling, VCS acts as a structural stress test: it accepts a reasoning trace only if it remains valid across generated counterfactuals (e.g., numerical perturbations, context shifts). While this strict filtering reduces the volume of training data (trading recall for precision), our experiments on MATH and GSM8K show it effectively prevents hallucination accumulation, achieving superior out-of-distribution generalization compared to baselines.

Read PDF

Similar papers

#natural language process... Preprint Aug 2026

When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents

It is shown the contamination is structurally irreversible: removing a source skill after the fact cannot erase the flawed reasoning its descendants have already inherited, so post-hoc rollback recovers only a small fraction of the lost performance, making skill admission a pre-commit necessity rather than a post-hoc f...

Lin-Fang Shang, Ming Xu, Yi-Ding Sun et al. · 5 citations · ⚡1
#machine learning Preprint Aug 2026

CoEvo: Oracle-Grounded Self-Evolution of a Single Model for Multi-Step Causal Reasoning

Multi-step causal reasoning requires chaining inferences where each step constrains the next. An early error propagates silently, and a correct answer reached via flawed logic evades outcome-level detection. In specialized domains, teacher LLMs err on intermediate steps, safety constraints restrict cloud distillation,...

Jian Zhang, Bing-Yi Wang, Yi-Zhi Liu · 0 citations
#artificial intelligence Preprint Aug 2026

The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

The mechanism is a halt vector: a difference-of-means direction at layer 18 of this model whose steering strength controls how long it thinks, while a replicated value axis does nothing, and what works is reconstructing the whole steered activation with those dimensions pinned to their natural values.

Dylan Jayabahu, Tinuade Adeleke · 0 citations
#artificial intelligence Preprint Sep 2026

Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning

Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verifi...

Xing Han, Yu-Xin Wang, Chen Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations

Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against...

Marcin Kostrzewa, Maciej Ziȩba · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.