TRACE is presented, a training-free layer that treats re-entry as an eligibility decision rather than a storage or retrieval operation, reconciling a departure checkpoint against absence-period updates, resolving explicit and implicit invalidation, and releasing a bounded Return View only when it covers the returning role's open obligations.
Abstract
Persistent memory lets language-model agents carry information across long-running collaborations, but leaves a lifecycle question open: what may a returning agent still act on once the shared state has changed? A memory can be correctly retrieved, relevant to the current task, and faithful to its source, and nonetheless be inadmissible for action: an itinerary saved before a pause still names the hotel the team has since replaced. We formalize this as temporal memory admission and present TRACE, a training-free layer that treats re-entry as an eligibility decision rather than a storage or retrieval operation, reconciling a departure checkpoint against absence-period updates, resolving explicit and implicit invalidation, and releasing a bounded Return View only when it covers the returning role's open obligations. We evaluate TRACE under three actor models on Memora, STALE Type II, and a derived ManBench-Return setting, each recast as return episodes: one agent departs, four teammates change the shared state, and the agent rejoins. What separates methods is not overall accuracy but whether one can retain valid memory and reject stale memory at once, and no single-policy baseline can: Restore (reinstate the departure checkpoint in full) admits stale state, Reset (start the return from an empty memory) discards valid state, each bottoming out at 0% on one of the two. TRACE is the only method high on both, reaching 92.6-98.3% valid-information availability with 98.4-99.5% invalid-information rejection on ManBench-Return, within 3.8 points of the best baseline's overall accuracy. On STALE Type II it improves Overall over the strongest comparison policy by 22.3 (Qwen), 18.5 (Gemini), and 27.5 (DeepSeek) points at roughly 2.3 times their tokens, while a write-time consolidation pipeline is more accurate still at 3.99 times TRACE's.
A large language model (LLM) agent that inherits a plan through shared memory can hold the latest requirement yet act on a plan derived from an older one: fresh memory, stale plan. Freshness checks miss this failure because they compare local copies with current state (observation currency) rather than the inputs the p...
Evan Chen, Shi-Qiang Wang, Christopher G. Brinton· 0 citations
GPM is introduced, an auditable bitemporal state-transition model with source-bound admission, derived lifecycle state, current public barriers, and fail-closed structured release with bounded contract and implementation results, not open-world model accuracy or evidence of world truth.
StateMem is presented, a state-first memory method that explicitly tracks supersession and relational dependencies, and it is shown it improves current-state accuracy over the strongest same-backbone baseline and over the strongest memory system, while remaining competitive with the long-context baselines.
Xin-Yi Fan, Miri Liu, Ruozhen Yang et al.· 4 citations
The introduction of TEPA, a revocable evidence-memory mechanism that makes validity an explicit state of memory, and the results establish lifecycle revocation as a core memory operation for agents that must falsify, audit, and later re-promote evolving knowledge.
Yan Zhou, Yue Ouyang, Kaiyang Zheng et al.· 3 citations· ⚡1
Large language model agents that persist across sessions, tools, users, and changing environments do more than answer isolated prompts; they accumulate state. When new evidence arrives, the central question is where that change should live: transient context, external memory, tool or workflow definitions, activation st...
Gabriel Chavira-Juárez, Eder Jahir Gonzalez Bravo, G. Rivera-García et al.· Frontiers in Big Data· 0 citations
C COUNTERMEM is introduced, a reinforcement-learning framework for constructing and using verified counterfactual memory across tasks, and it is shown that removing verification or persistent storage weakens the gains, while applying verified corrections to unsuitable decisions can reverse them.