Skip to content

PREreview of "Memory Control Signals Emerge Before Action in Long Horizon Agents"

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling

Abstract

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23197372. ## Summary This paper asks whether a language model already represents its memory needs — when to compress history, when to recall earlier evidence — in its hidden state before acting. Using linear probes on pre-action hidden states across 72,912 decision points from 2,521 long agent trajectories, the authors find compression and recall needs are predictable at AUROC 0.831 and 0.765, well above observable controls (context length, turn count, tool type), with compression signals strengthening with depth and recall peaking at intermediate layers. They further show that keeping just the task prefix plus the two most recent interaction blocks (~26% of full context) preserves most decision information, while selective restoration of raw historical evidence recovers long-range dependencies. These findings motivate PaMER: a frozen Qwen3.5-9B feature extractor reads the pre-action state to decide when to compress, older history moves to retrievable external memory, and PaMER+ adds step-level evidence selection. On WorkBuddyBench (260 tasks), token consumption drops dramatically — e.g., −88% on GPT-5.6-Luna — with task performance described as "competitive" but varying by model and domain: MiMo-V2.5 gains +8.8 points average while GPT-5.6-Luna loses −2.3, and the Office domain collapses by −14 points in places. The authors are careful to say the aggregate should not be read as an improvement on every task type, and they disclose LLM assistance including an LLM annotator for the memory-decision corpus. ## Strengths 1. The probing study is carefully scoped. The authors compare hidden states against observable metadata controls, analyze signal formation across layers, and explicitly frame claims as predictive rather than causal — citing the probe-interpretation literature. This is how mechanistic analysis should be reported. 2. The 26% finding is practically useful. That a task prefix plus two recent blocks preserves most memory-decision information gives practitioners a concrete, cheap heuristic independent of the full PaMER machinery. 3. Honest about heterogeneity. Rather than hiding the model- and domain-dependence, the paper reports it: PaMER helps some backbones and hurts others, and says so. The cross-model evaluation with a shared frozen extractor is the right experimental design for a controller paper. 4. Good disclosure hygiene. The AI-use statement covers language editing, literature discovery, and — importantly — the LLM annotator. Reviewers can calibrate accordingly. ## Major comments 1. The probe's ground truth is an LLM annotator's judgment. The 72,912 decision labels come from an LLM identifying "states where compression or recall is needed" — so the probe may be decoding the annotator model's notion of memory need rather than a property of the trajectories themselves. The paper says labels were reviewed, but without human inter-annotator agreement numbers on a sample, the headline AUROCs rest on a circular foundation: an LLM's representation predicts another LLM's labels. A human-adjudicated validation subset is needed. 2. The memory controller's own compute is missing from the cost accounting. PaMER runs a forward pass through a frozen 9B model at every agent decision point to extract the pre-action representation. That is real GPU compute on every step of every trajectory — yet the token-consumption tables report only the agent's context tokens. For a paper whose practical claim is "substantially reduces context consumption," the controller overhead must be in the ledger: tokens are not the only cost, and a 9B forward pass per step is not free. 3. Results are strongly model- and domain-dependent in ways the design doesn't explain. GPT-5.6-Luna degrades under both PaMER variants; the Office domain drops 14 points; MiMo-V2.5 gains 15.8. If the memory-need signal is a general property of pre-action states, why does acting on it hurt some models? The paper needs at least a hypothesis — e.g., whether the frozen Qwen3.5-9B extractor's representations transfer poorly to certain backbones' trajectory distributions. 4. Single benchmark. All downstream evaluation is WorkBuddyBench. Memory management that works on office-workflow tasks may not transfer to coding agents or long-horizon research tasks. One additional domain would materially strengthen the generality claim. ## Minor comments 1. "Qwen3.5-9B" as the extractor naming is confusing against the Qwen3.8-Flash backbone also evaluated — clarify the model lineage. 2. Recall is evaluated only when valid external memory exists — the resulting selection effect on the recall AUROC should be quantified. ## Overall assessment Recommend with revisions. The probing analysis is a genuine contribution to understanding memory needs in long-horizon agents, and the honest reporting of heterogeneous results builds trust. But the LLM-labeled ground truth, the unaccounted controller compute, and the unexplained model-dependence need addressing before the "predictable therefore actionable" leap can be recommended to practitioners managing real context budgets. Competing interests The author declares that they have no competing interests. Use of Artificial Intelligence (AI) The author declares that they used generative AI to come up with new ideas for their review.

View source

Similar papers

#artificial intelligence Conference Open access Apr 2020

ECCOLA - a Method for Implementing Ethically Aligned AI Systems

The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.

Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson · 64 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Open access Mar 2024

LLM-based agents for automating the enhancement of user story quality: An early report

The use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams is explored, with a reference model for an Autonomous LLM-based Agent System developed and implemented at the company.

Zheying Zhang, M. Rayhan, Tomas Herda et al. · 48 citations · ⚡4
#computer vision Review Mar 2024

System for systematic literature review using multiple AI agents: Concept and an empirical evaluation

This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.

Abdul Malik Sami, Z. Rasheed, Kai-Kristian Kemell et al. · 44 citations · ⚡2
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 41 citations

Related blog posts

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.