Depth estimation from thermal images is highly valuable for robotic applications in adverse conditions, such as nighttime and rainy weather. Recent studies have sought to transfer knowledge from RGB-based foundation models to thermal modalities, yet the rich hierarchical representations these models encode remain under...
Jie Hong, Ting-Tian Li, Xue-Song Li et al.· 0 citations
EvoScene-VLA is introduced, which uses compact scene tokens to unify within-chunk scene prediction, cross-chunk state propagation, and observation-based correction and shows that a recurrent scene state alone does not necessarily improve performance, whereas state propagation can improve control when combined with geom...
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on 40 simulated robotic manipulation tasks, paired closed-loop...
Chu-Shan Zhang, Jin-Guang Tong, Xue-Song Li et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.