PM-Bench, a text-based benchmark for measuring prospective memory capabilities in modern LLM agents, is introduced, Inspired by the Virtual Week paradigm from cognitive science, which evaluates how well LLM agents maintain user intentions, execute delayed intentions, and monitor latent environment changes.
Abstract
A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while other activities are ongoing. We introduce PM-Bench, a text-based benchmark for measuring prospective memory capabilities in modern LLM agents. Inspired by the Virtual Week paradigm from cognitive science, PM-Bench evaluates how well LLM agents maintain user intentions, execute delayed intentions, and monitor latent environment changes. Over the course of a simulated seven-day week, agents must continue an ongoing activity while deciding whether any deferred task is due. We compare eight state-of-the-art LLMs on PM-Bench under eight different agent configurations. PM-Bench proves challenging across all settings: the best method, a GPT-5.4 agent, reaches only 65.1\% F1 score under our evaluation. Furthermore, no single strategy for improving prospective memory dominates across models. We release PM-Bench as a controlled testbed for diagnosing these failures and developing training or inference-time interventions that support reliable prospective behavior.
PIS sets a new state of the art on this benchmark and enables small models to surpass the published large-model scaffold and enables small models to surpass the published large-model scaffold.
Lifelong LLM agents increasingly adapt through external learning states that store past interactions as retrievable memories or reusable skills, yet existing benchmarks rarely account for how the path of accumulated experience shapes what agents transfer and retain. In this work, we establish PATH-Bench, a benchmark fo...
Xi-Dong Yang, Xing-Yi Zhang, Wen-Hao Li et al.· 0 citations
A benchmark and harness that measures the marginal utility of memory for task-executing agents under explicit cost accounting, and provides episodic tool-use tasks in three domains whose dependence on earlier-episode facts is verified by an automated leak check.
This study compares MemoryLake, a structured multi-track memory backend, with Mem0, text-embedding-3-small vector RAG, and a long-context control across all five MemoryArena domains, to support a workload-dependent view of memory backends and an observed lead among the four evaluated systems.
Chao-Shun Zhan, Qiang Zhou, Guannan Li et al.· 0 citations
The PAST-Bench benchmark is introduced, a benchmark designed to isolate how persistent agents can progress from retaining experience to systematically improving through it, and Hermes+ is developed, which raises the average gain from retained experience and provides clearer pathway evidence.
Shu-Han Xue, Zixin Ding, Yi-Jun Shen et al.· 2 citations
Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retention or r...
Ben Wang, Kang Zhou, Li-Fan Guo et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.