Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration
It is argued that scientific memory should be evaluated as budgeted, modality-aware context restoration rather than as an unconstrained architecture leaderboard, and the datasets, harness, raw outputs, judgments, and scripts are released to reproduce the results and serve as tools for such evaluation.