Skip to content

Retrieval-Guided Fine-Tuning as Noisy Estimation: Risk bounds and Architectural Analysis

Sep 2026 · 0 citations · 26 references
Computer Science Mathematics

TL;DR

Under homoscedastic retrieval noise, it is shown that retrieval failure decays exponentially with task separation relative to noise, and explicit finite-sample conditions under which RAG-FT achieves lower risk than both target-only and full-corpus training are derived.

Abstract

Retrieval-Guided Fine-Tuning (RAG-FT) incorporates retrieved data directly into the training objective, but the statistical consequences of noisy retrieval during training remain theoretically undercharacterized. We study this question by modeling RAG-FT as an estimation problem in a multi-task linear regression framework, using an OLS proxy for single-layer linear self-attention to obtain finite-sample risk bounds. Under homoscedastic retrieval noise, we show that retrieval failure decays exponentially with task separation relative to noise, and derive explicit finite-sample conditions under which RAG-FT achieves lower risk than both target-only and full-corpus training. We then introduce a Distance-Proportional Noise (DPN) model, in which retrieval quality degrades with rank, and compare two estimators under the same retrieval process: the OLS proxy and the literal, uniform-weight forward pass of linear self-attention. We prove that the attention estimator's bias diverges as $\Theta(n^{2q})$ even under exact retrieval, while OLS risk remains $\Theta(d/n)$ for every noise exponent $q>0$. These results locate the instability not in noisy retrieval itself, but in the fixed, unweighted aggregation of the literal LSA forward pass, which reweighting by reliability empirically removes. We validate the predicted rate separation through direct simulation of the DPN model.

View source

Similar papers

Preprint Aug 2026

The RAT: A Unified Bayesian Model for RAG Evaluation

A Bayesian evaluation framework is introduced that jointly models retrieval success, abstention behavior, and answer correctness, factorized according to the pipeline's information flow, and extends to incorporate LLM-as-a-judge annotations as calibrated noisy observations, enabling practitioners to combine limited hum...

Pius von Däniken, Felix Matthias Saaro, Mark Cieliebak et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Learning from What You Retrieve: Online RL Fine-Tuning for Semantic Retrieval

This work proposes PAO (Positive-Advantage-Only), a selective RL optimization method that selectively applies gradient updates only to retrieved items with positive advantages, effectively pulling query embed- dings toward high-reward regions while preserving global topo- logical stability.

Shao-Wei Wei, Chong Huang, Songtao Fang et al. · 0 citations
Preprint Aug 2026

Noise-Aware Shrinkage for Differentially Private Zeroth-Order Fine-Tuning of Large Language Models

SAGE is proposed, a noise-aware shrinkage method that adaptively attenuates privatized estimates according to their estimated signal quality, and shows that shrinkage reduces the quadratic update-risk term faster than the linear descent term, preserving useful descent while limiting the influence of noise-dominated upd...

Le-Le Zheng, Wei-Feng Kong, Xinyi Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

Online Self-Weighted Fine-Tuning

Online Self-Weighted Fine-Tuning is proposed, a simple method that augments SFT with online, trajectory-level weighting and offers a favorable compute-performance trade-off as a practical approach for fine-tuning small-to-medium LLMs on binary-verifiable reasoning tasks with only 2 online rollouts.

Hai-Quan Wen, Yiwei He, Bei Peng et al. · 0 citations

STAR: Structure-Aware Adaptive Retrieval for RAG

STAR is presented, a structure-aware adaptive retrieval framework for RAG that treats this mismatch as a problem of diagnosing evidence sufficiency and benefits from a control signal that preserves structurally distinct insufficiency patterns rather than collapsing them into a single scalar confidence estimate.

Yeowon Jeon, Chong-kwon Kim, Y. Choi · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.