The results show that replay retention depends on change magnitude and on how the dynamics evolve, and an estimator built from interaction data can still provide the quantities needed to choose a replay strategy after permanent changes.
Everest Yang, Skye Thompson, G. Konidaris· 0 citations
This work aims to identify the benefits of HRL from the perspective of the fundamental challenges in decision-making, as well as highlight its impact on the performance trade-offs of AI agents.
Martin Klissarov, Akhil Bagaria, Zi-Yan Luo et al.· arXiv.org· 30 citations· ⚡1
This work considers agent memory as a temporally-extended abstraction over the agent’s observation-action history, and derives POMDP classes by applying traditional MDP state abstractions, such as model-preservation, optimal value Q ∗ preservation, and optimal policy π ∗ preservation, to the POMDP setting.
Aaron Kirtland, Alexander Ivanov, Cameron S. Allen et al.· 1 citation
This work introduces tentative object pruning, an approach that leverages this property by constructing multiple simplified tasks with reduced object sets and searching them in parallel until a valid, satisficing solution is found.
Anita de Mello Koch, Naman Shah, Cameron S. Allen et al.· 0 citations
We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configur...