Skip to content
Preprint

FailForge: Distilling Procedural Competence from Persistent Failures into Code Agents

Aug 2026 · 1 citation · 32 references
Computer Science

TL;DR

FailForge is proposed, an agentic framework that converts failed rollouts into training signal, and recovers over 26% of previously failed instances at marginal additional cost, and training Qwen3.5-4B on the augmented corpus improves the SWE-bench Verified resolve rate by 6.6 points over a strong RFT baseline.

Abstract

Rejection sampling fine-tuning (RFT) is widely used to train code agents by generating trajectories on verifiable software engineering tasks, retaining those that pass the tests, and fine-tuning on the successful rollouts. However, even strong code agents repeatedly fail on a substantial fraction of such tasks, and standard RFT simply discards these failures. The discarded samples are precisely the hardest and most informative ones, drawn from verifiable instances that are costly to curate. Stronger base models may reduce the number of failures, but the remaining hard cases still define the frontier for further improvement. We propose FailForge, an agentic framework that converts failed rollouts into training signal. For each failed instance, an agent diagnoses the failure from error feedback and execution traces, distills the diagnosis into a concise and actionable skill, and injects the skill into the agent context for a guided second attempt. Trajectories that succeed under skill guidance are folded back into the RFT corpus. Crucially, the skill is removed at training time, so the model internalizes the recovered behavior rather than relying on external hints at inference. FailForge recovers over 26% of previously failed instances at marginal additional cost, and training Qwen3.5-4B on the augmented corpus improves the SWE-bench Verified resolve rate by 6.6 points over a strong RFT baseline, with gains concentrated on the hardest problems.

View source

Similar papers

Preprint Aug 2026

Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction

DARC is proposed, a diagnosis-guided recovery harness that profiles task-family failure modes, prunes mismatched interventions from a shared recovery library, and freezes a verifier-selected success-cost policy for deployment, providing a practical route toward more reliable agents in domains where compiler-like feedback is absent.

Pan Wang, Yihao Hu, Hang Wang et al. · 0 citations
Jul 2026

Structured Feedback Improves Repair in an LLM Agent Loop

VeriHarness is introduced, a code-controlled agent loop in which models generate candidates while external validators control acceptance, budgets, and traces, and it is used to compare raw diagnostics with feedback that identifies the failure location, observed value, and admissible alternatives.

Jaideep Ray, Ankit Goyal · 1 citation
Preprint Aug 2026

Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems

This study introduces SymTrace, a controlled evaluation framework that records the MAS execution trajectory and establishes intervention anchors and explores the effectiveness of MAS repair methods, revealing that existing unguided rerun methods are highly unreliable.

Zhong-Wen Luan, Xiaoyan Zhang, Ming Hu et al. · 2 citations
#machine learning Preprint Sep 2026

How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method

LLM agents deployed for software engineering fail expensively: they act confidently wrong, and bad actions are recognized only after costly execution and retry. We present Speculative Uncertainty (SU), a method that recovers a predictive failure signal for a black-box agent from its output tokens alone, with no access to logits, weights, activations, or repeated sampling. Inverting speculative decoding, a small open-weight draft model scores the agent's already-generated trajectory in a single forward pass. From these speculative cross-likelihoods we extract phase-aware features by separating the reasoning and action spans, and calibrate them against a verifiable objective. SU produces a failure-likelihood score that any downstream policy, such as routing, human intervention, or extra test-time compute, can consume directly. To show the signal is actionable, we instantiate one such policy, a pre-execution veto gate, on software engineering agents Qwen3-Coder-480B and closed-source Claude 3.5 Sonnet, cutting execution error rate by 6-8 percentage points and token cost by 14-19% in deployment, transferring to out-of-distribution benchmarks without retraining, and generalizing across agent models.

Konstantin Grotov, Valentin Malykh · 0 citations
Jul 2026

AI Agents Do Not Fail Alone:The Context Fails First

These findings establish context measurement as a validated preflight signal for agent reliability and position context engineering as an auditable layer of agent evaluation and governance.

Fouad Bousetouane · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.