Skip to content

Author

Sangbeom Lee

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

Robust 3D-Aware Video Object Tracking for Mobile Robots Based on 2D Vision Foundation Models

Egocentric cameras are widely used in robotic navigation and manipulation, yet conventional 2D Video Object Tracking (VOT) methods suffer from severe performance degradation under rapid viewpoint changes and frequent frame-out events. Because most existing trackers rely solely on 2D appearance cues, they often fail to recover object identities once targets temporarily disappear. We propose R3DVOT, a 3D-aware tracking framework that augments 2D vision foundation models with spatial geometric reasoning. Its core component, the Position-Aware Memory Selection (PAMS), lifts mask candidates into a canonical 3D world coordinate system and maintains a persistent world-frame state for position-consistent hypothesis selection. By curating the memory bank with spatially consistent anchors, R3DVOT improves robustness to occlusion and frame-out events. On the VOT benchmark, R3DVOT achieves an AUC improvement of 5.1% over SAM 2 and 1.4% over SAMURAI. Furthermore, on the VOS benchmark, R3DVOT increases the J&F score by 4.6% compared to SAM 2 and 9.0% compared to SAMURAI. These results highlight the effectiveness of 3D spatial continuity in enhancing tracking, segmentation, and long-term identity consistency in robotic perception.

Hayeoung You, Sangbeom Lee, Huisu Kim et al. · 0 citations
Conference Jul 2026

Physical Reasoning VLA Models for Enhancing Robotic Manipulation

Vision-Language-Action (VLA) models have demonstrated strong performance in robot manipulation by leveraging pre-trained vision-language models to map observations directly to actions. However, existing approaches reason primarily at the visual or semantic level, lacking explicit understanding of the physical interactions that fundamentally govern manipulation tasks. In this paper, we propose Physics Reasoning VLA, a method that enables VLA models to explicitly reason about physical interactions prior to acting, grounded in two fundamental quantities: contact points, which specify where the target object interacts with the robot or surrounding environment, and contact forces, which describe the magnitude and direction of force applied at those locations. Rather than directly mapping observations to actions, our model first predicts contact points and forces at the pixel level via learnable physics queries, then incorporates the resulting physics-aware features alongside visual and language inputs to generate actions. To prevent physics reasoning from disrupting pre-trained visual and linguistic representations, we further introduce a hybrid attention mechanism that applies full attention over image, language, and proprioceptive tokens, while applying causal attention over physics query and action tokens. We evaluate our method on the RoboCasa simulation benchmark, demonstrating that physics reasoning consistently improves performance over the vanilla π0 baseline, with an average success rate improvement from 36.2% to 43.6%.

Kangmin Kim, Geonhyup Lee, Sangbeom Lee et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.