Skip to content

PM-Bench: Evaluating Prospective Memory in LLM Agents

Jul 2026 · arXiv.org · Vol abs/2607.12385 · 1 citation · 55 references
Computer Science

TL;DR

PM-Bench, a text-based benchmark for measuring prospective memory capabilities in modern LLM agents, is introduced, Inspired by the Virtual Week paradigm from cognitive science, which evaluates how well LLM agents maintain user intentions, execute delayed intentions, and monitor latent environment changes.

Abstract

A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while other activities are ongoing. We introduce PM-Bench, a text-based benchmark for measuring prospective memory capabilities in modern LLM agents. Inspired by the Virtual Week paradigm from cognitive science, PM-Bench evaluates how well LLM agents maintain user intentions, execute delayed intentions, and monitor latent environment changes. Over the course of a simulated seven-day week, agents must continue an ongoing activity while deciding whether any deferred task is due. We compare eight state-of-the-art LLMs on PM-Bench under eight different agent configurations. PM-Bench proves challenging across all settings: the best method, a GPT-5.4 agent, reaches only 65.1\% F1 score under our evaluation. Furthermore, no single strategy for improving prospective memory dominates across models. We release PM-Bench as a controlled testbed for diagnosing these failures and developing training or inference-time interventions that support reliable prospective behavior.

View source

Similar papers

Preprint Aug 2026

PATH-Bench: Path-Dependent Evaluation of Lifelong Agents

Lifelong LLM agents increasingly adapt through external learning states that store past interactions as retrievable memories or reusable skills, yet existing benchmarks rarely account for how the path of accumulated experience shapes what agents transfer and retain. In this work, we establish PATH-Bench, a benchmark fo...

Xi-Dong Yang, Xing-Yi Zhang, Wen-Hao Li et al. · 0 citations
#artificial intelligence Preprint Jul 2026

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

A benchmark and harness that measures the marginal utility of memory for task-executing agents under explicit cost accounting, and provides episodic tool-use tasks in three domains whose dependence on earlier-episode facts is verified by an automated leak check.

Shweta Mishra, Shashank Mishra · 0 citations
Preprint Aug 2026

MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends

This study compares MemoryLake, a structured multi-track memory backend, with Mem0, text-embedding-3-small vector RAG, and a long-context control across all five MemoryArena domains, to support a workload-dependent view of memory backends and an observed lead among the four evaluated systems.

Chao-Shun Zhan, Qiang Zhou, Guannan Li et al. · 0 citations
Preprint Aug 2026

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

The PAST-Bench benchmark is introduced, a benchmark designed to isolate how persistent agents can progress from retaining experience to systematically improving through it, and Hermes+ is developed, which raises the average gain from retained experience and provides clearer pathway evidence.

Shu-Han Xue, Zixin Ding, Yi-Jun Shen et al. · 2 citations
Preprint Aug 2026

FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retention or r...

Ben Wang, Kang Zhou, Li-Fan Guo et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.