Skip to content

PreSIST: Vision-Language-Informed Object Persistence Prediction in Open-World Scenes

Jul 2026 · arXiv.org · Vol abs/2607.04057 · 0 citations · 41 references
Computer Science

TL;DR

This work proposes PreSIST (Predictive Scene-conditioned Instance Survival over Time), a method for predicting whether an observed object will remain in its last seen pose at arbitrary future times and develops two interchangeable variants: PreSIST-Lang, which estimates persistence priors using a VLM, and PreSIST-Vis, a novel vision-only model trained using PreSIST-Lang pseudo-labels for efficient deployment.

Abstract

Robots deployed over long periods must reason about environments that change over time. Existing long-term perception systems often address object change reactively, updating their maps only after revisiting a scene and observing that an object has moved. Instead, robots should reason proactively about how long objects are likely to persist using the context in which they appear. For example, a car at a traffic light and a car in a parking spot share the same semantic class, but their contexts imply different persistence durations. We propose PreSIST (Predictive Scene-conditioned Instance Survival over Time), a method for predicting whether an observed object will remain in its last seen pose at arbitrary future times. PreSIST estimates instance-level persistence priors from object properties and scene context, then integrates these priors with a probabilistic persistence filter as observations become available. Its key insight is that the reasoning capabilities of vision-language models (VLMs) can relate scene context to likely object use and human activity, enabling persistence prediction before long-term observations are available. We develop two interchangeable variants: PreSIST-Lang, which estimates persistence priors using a VLM, and PreSIST-Vis, a novel vision-only model trained using PreSIST-Lang pseudo-labels for efficient deployment. Experiments on a new dataset of in-the-wild object persistence annotations show that PreSIST-Lang and PreSIST-Vis outperform baselines on open-world persistence prediction.

View source

Similar papers

Open access Jul 2026

SuperMap: A Spatio-Temporal SLAM System for Visual-Language Navigation

This work presents SuperMap, a 4D spatio-temporal mapping framework for language-guided navigation that integrates high-frequency geometric SLAM with asynchronous open-vocabulary perception and releases the full system as open-source to provide the community with a deployable baseline for open-vocabulary spatio-tempora...

Shibo Zhao, Guofei Chen, Honghao Zhu et al. · 2 citations
Aug 2026

Towards a Causally-inspired Evolving World Model for Vision-and-Language Navigation in Continuous Environments.

A causally-inspired evolving world modeling framework that formulates VLN-CE as a sequence of causal partially observable Markov decision processes that learns unified latent states that integrate vision, language, and action, and strengthens generalization across diverse navigation contexts.

Xuan Yao, Junyu Gao, Chang-Sheng Xu · 0 citations
Preprint Aug 2026

Hierarchical Fast-Slow ReAct Agent for Zero-Shot Object-Goal Navigation

This paper presents a hierarchical fast-slow agent that turns what the robot has already seen into the object of deliberation in zero-shot object-goal navigation, and reaches the highest success rate among the zero-shot methods compared here.

Zhaochen Lan, Zhi Yang, Yuxiang Fu et al. · 0 citations
Preprint Sep 2026

Map the Possibilities: Spatial Belief Fields for Language-Goal Aerial Navigation

Language-goal aerial navigation requires an agent to local- ize a potentially unobserved target from relational instruc- tions and partial observations, and translate this inference into metric actions in large-scale continuous environments. Existing methods often reduce language grounding to one single waypoint or act...

Hao-Tian Xu, Yue Hu, Zheng-Qiu Zhu et al. · 0 citations
Preprint Sep 2026

HINT: Human-Intent Inception for Long-Horizon Robot Manipulation

Humans can perform complex manipulations given a simple intent through an overall instruction, while continuously adapting to evolving visual observations. However, current vision-language action (VLA) models and other action policies struggle to realize this high-level intelligent behavior under dense, evolving visual...

Ming-Yu Mei, Haojie Xu, Shi-Hao Jin et al. · 0 citations
Jul 2026

Memory for Attention: Language-Conditioned Re-Perception with a Vision-Language-Motion Map

A persistent map's memory yields the best schedule (held-out), matching an oracle, while the memoryless VLM prior is poor, showing a persistent map's memory earns its keep not by telling the robot how to walk around a room, but by telling it what to pay attention to.

Dibyendu Ghosh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.