Skip to content

Failure-Guided Co-Evolution of Prompts and Training Data

Sep 2026 · 0 citations · 43 references
Computer Science

TL;DR

ForGE is introduced, a failure-guided framework that co-evolves prompts and training data and establishes failures as a shared interface between prompt optimization and data synthesis, and shows the benefit of jointly adapting what a model is instructed to do and what it learns from

Abstract

Automatic prompt optimization (APO) improves language-model programs by revising prompts from task feedback, yet it typically holds its training data fixed. Repeatedly optimizing against the same instances confines feedback to weaknesses already represented in those data, leaving related failure conditions unexplored. We therefore view each failure as a dual signal: it indicates both how the prompt should be revised and what new training evidence should be synthesized. We introduce FORGE, a failure-guided framework that co-evolves prompts and training data. FORGE abstracts imperfect executions into reusable failure modes and synthesizes new training data through four complementary mutation strategies. Verified instances are fed back into prompt search, allowing updated prompts to expose the next data needs. Across eight heterogeneous benchmarks, FORGE improves the aggregate score over the unoptimized baseline by 16.52 percentage points and outperforms all evaluated APO baselines. The synthesized data also transfer beyond FORGE: in a transfer study, they improve all nine APO comparisons by 2--9 points and all three GRPO comparisons by 4--8 points under matched optimization budgets. These results establish failures as a shared interface between prompt optimization and data synthesis, and show the benefit of jointly adapting what a model is instructed to do and what it learns from.

View source

Similar papers

Conference Aug 2026

When Iterative Prompting Fails: An Empirical Study of Unit Test Generation with Open-Source LLMs

Large language models (LLMs) have shown promise in automated unit test generation, yet the effectiveness of prompt engineering for small, locally-deployed open-source models remains poorly understood. Following growing interest in local LLM deployment to mitigate data exposure risks, this paper presents a controlled em...

M. Tran, Khang Mai · 0 citations
#artificial intelligence Preprint Sep 2026

STEVE: Stabilizing Textual Gradient-Based Prompt Optimization via Error-Driven Refinement and Regularized Verification

Textual-gradient methods automate prompt optimization through natural-language feedback, but their iterative updates can be unstable. We identify two sources of this instability: noisy gradients produced from already-correct examples and over-specialization to hard cases that degrades performance on simpler inputs. We...

Yi-Fan Xu, Yi-Xuan Li, Xin-Zhuo Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Prompt or Skill? Attribution-Guided Optimization of Modular LLM Programs

Large language models can solve increasingly diverse reasoning tasks, yet their performance remains highly sensitive to task prompts, intermediate instructions, and the way reusable problem-solving knowledge is incorporated. Existing optimization methods usually focus on only one part of this design space: they either...

Hao-Ran Shou, Hao-Yue Liu, Yun-Chen Huo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize

ESO (Error-Structured Prompt Optimization), which decomposes prompt optimization into three phases: Diagnose clusters all training errors into structural patterns in one round; generate candidates via four complementary strategies with independent biases; Select applies bootstrap stability selection.

Li-Hao Liu, Peng Tang, Kunwar Yashraj Singh et al. · 0 citations
Review Aug 2026

DynaContext: Self-Improving Dynamic Contextualization of Optimized Prompts for Heterogeneous Parameter Extraction

DynaContext is introduced, a framework that combines an offline-optimized extraction core, learned with GEPA or SkillOpt, with inference-time contextual adaptation and validation-gated self-improvement, with inference-time contextual adaptation and validation-gated self-improvement.

Joe Yu, Shibin Thomas Stanley Paul, Sven Mayer · 0 citations
Preprint Aug 2026

Aging of Prompt Engineering Techniques Across LLM Versions

It is shown that prompt engineering "ages" in a model-family-specific way: Newer GPT models exhibit diminishing or even negative marginal gains from structured prompting, suggesting that instruction-following and reasoning scaffolds are increasingly internalized, whereas Qwen models continue to benefit substantially fr...

A. Rudyk, Julian Oertel, Regina Hebig · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.