Skip to content

PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents

Sep 2026 · 0 citations · 33 references
Computer Science

TL;DR

Pseudo Self-Distillation is presented, a framework that enables small language models to construct hierarchical memory representations by distilling behavior from a strong black-box oracle through a multi-stage training pipeline, with off-policy PSD achieving the strongest results across most conditions.

Abstract

Memory systems are becoming a core component of LLM agents, but constructing and maintaining memory remains expensive because it relies on repeated calls to large proprietary language models. This cost creates a major barrier to deploying memory-enhanced agents at scale. In this paper, we present Pseudo Self-Distillation (PSD), a framework that enables small language models (SLMs) to construct hierarchical memory representations by distilling behavior from a strong black-box oracle through a multi-stage training pipeline. Standard distillation methods require access to teacher logits or hidden states, which closed models do not expose. Unlike conventional self-distillation settings, where supervision is derived from a model's own predictions, sampled rollouts, or aggregated outputs, PSD enables a single-model distillation setup while channeling external oracle knowledge through the prompt. PSD uses a single small model in two roles: a teacher that sees a privileged prompt containing the oracle's answer as reference context, and a student that sees only the task prompt. The student learns to reproduce the teacher's output distribution, absorbing oracle-guided behavior into its own weights without accessing the oracle's internals. On LoCoMo, PSD-trained Qwen3-0.6B, 1.7B, and 4B match or exceed GPT-4.1-mini on downstream retrieval at a fraction of the deployment cost, with off-policy PSD achieving the strongest results across most conditions. We further show that this memory-construction capability transfers out-of-distribution to LongMemEval, despite the students being trained exclusively on LoCoMo with no exposure to LongMemEval data.

View source

Similar papers

Preprint Aug 2026

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

Agent Memory Distillation (AMD), a training-free framework that transfers structured knowledge from a large teacher agent to a small student agent through hierarchical memory, is proposed, which consistently outperforms existing memory-based baselines.

Taeil Kim, Kangsan Kim, S. Hwang · 3 citations
#artificial intelligence Preprint Sep 2026

Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models

SOLID is proposed, a novel framework for self-improving OR language models without verified answers or external evaluators that improves solution accuracy for both general-purpose and OR-tuned models over outcome-only group-relative training.

Rui-Chen Zhu, Ming-Long Cao, Chen-Yu Zhou et al. · 0 citations
#natural language process... Preprint Sep 2026

Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models

CaRE-KD is proposed, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization and provides a gradient-level analysis showing how this dual-granularity design induces a conditional calibration mechanism that prior static divergences cannot reproduce.

Ayan Sengupta, Vaibhav Seth, Tanmoy Chakraborty · 0 citations
#machine learning Preprint Sep 2026

MemoryWalker: Stop Training Agents on Contexts They Never Saw

This work introduces two exact, gradient-equivalent corrections: LogitTree, a segmented K-forward traversal, and a packed 4D attention mask, and proposes SDCC (Self-Distillation for Conditioning Consistency), a single-backward-pass variational relaxation.

J. Zinco, Xun-Jie Zhu, Shen Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models

Large language models increasingly serve as long-horizon assistants and agents, where they must both accumulate information across interactions and make the relevant parts available when later requests depend on them. Existing compact online memories typically use a single persistent state both to accumulate history an...

Xiang-Long Shi, Rui-Jie Yang, Si-Rui Zhao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beam Search as Test-Time Self-Distillation via Counterfactual Contexts

Self-Distillation Fine-Tuning (SDFT) enables a language model to act as its own teacher: by conditioning on a demonstration, the model produces an implicit reward via pointwise mutual information, which guides on-policy learning without external supervision. However, SDFT operates at training time: it requires gradient...

Su-Ee Tan, Xiao-Tong Ji, Rasul Tutunov et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.