ForGE is introduced, a failure-guided framework that co-evolves prompts and training data and establishes failures as a shared interface between prompt optimization and data synthesis, and shows the benefit of jointly adapting what a model is instructed to do and what it learns from
Abstract
Automatic prompt optimization (APO) improves language-model programs by revising prompts from task feedback, yet it typically holds its training data fixed. Repeatedly optimizing against the same instances confines feedback to weaknesses already represented in those data, leaving related failure conditions unexplored. We therefore view each failure as a dual signal: it indicates both how the prompt should be revised and what new training evidence should be synthesized. We introduce FORGE, a failure-guided framework that co-evolves prompts and training data. FORGE abstracts imperfect executions into reusable failure modes and synthesizes new training data through four complementary mutation strategies. Verified instances are fed back into prompt search, allowing updated prompts to expose the next data needs. Across eight heterogeneous benchmarks, FORGE improves the aggregate score over the unoptimized baseline by 16.52 percentage points and outperforms all evaluated APO baselines. The synthesized data also transfer beyond FORGE: in a transfer study, they improve all nine APO comparisons by 2--9 points and all three GRPO comparisons by 4--8 points under matched optimization budgets. These results establish failures as a shared interface between prompt optimization and data synthesis, and show the benefit of jointly adapting what a model is instructed to do and what it learns from.
Large language models (LLMs) have shown promise in automated unit test generation, yet the effectiveness of prompt engineering for small, locally-deployed open-source models remains poorly understood. Following growing interest in local LLM deployment to mitigate data exposure risks, this paper presents a controlled em...
M. Tran, Khang Mai· International Conference on...· 0 citations
Textual-gradient methods automate prompt optimization through natural-language feedback, but their iterative updates can be unstable. We identify two sources of this instability: noisy gradients produced from already-correct examples and over-specialization to hard cases that degrades performance on simpler inputs. We...
Yi-Fan Xu, Yi-Xuan Li, Xin-Zhuo Li et al.· 0 citations
Large language models can solve increasingly diverse reasoning tasks, yet their performance remains highly sensitive to task prompts, intermediate instructions, and the way reusable problem-solving knowledge is incorporated. Existing optimization methods usually focus on only one part of this design space: they either...
Hao-Ran Shou, Hao-Yue Liu, Yun-Chen Huo et al.· 0 citations
ESO (Error-Structured Prompt Optimization), which decomposes prompt optimization into three phases: Diagnose clusters all training errors into structural patterns in one round; generate candidates via four complementary strategies with independent biases; Select applies bootstrap stability selection.
DynaContext is introduced, a framework that combines an offline-optimized extraction core, learned with GEPA or SkillOpt, with inference-time contextual adaptation and validation-gated self-improvement, with inference-time contextual adaptation and validation-gated self-improvement.
Joe Yu, Shibin Thomas Stanley Paul, Sven Mayer· 0 citations
It is shown that prompt engineering "ages" in a model-family-specific way: Newer GPT models exhibit diminishing or even negative marginal gains from structured prompting, suggesting that instruction-following and reasoning scaffolds are increasingly internalized, whereas Qwen models continue to benefit substantially fr...
A. Rudyk, Julian Oertel, Regina Hebig· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.