Jul 2026· Journal of Instrument Engineering· 0 citations· 2 references
Abstract
For embodied agents to navigate and reason indoor spaces, they need object-level 3D representations that stay consistent over time as new frames arrive from a monocular camera. Current online 3D instance segmentation methods either depend on posed RGB-D input with ground-truth depth or couple tightly to the internal representations of specific foundation models, sacrificing modularity. We observe that appearance-based and geometry-based object matching exhibit complementary failure modes: appearance is ambiguous among spatially separated duplicates, while geometry is unreliable for visually distinct objects at similar locations. This motivates SAM3R, a training-free pipeline that fuses spatial overlap, 3D centroid displacement, and visual-semantic similarity into a single assignment cost solved via bipartite matching. The cost is constructed entirely from the outputs of frozen foundation models without accessing internal representations. Object tracks are classified through a cascaded decision tree that detects scene changes via field-of-view gated temporal voting. On ScanNet200 and Replica, SAM3R performs competitively with methods that require architecture-specific features or additional training, despite operating in a fully online, monocular setting. Qualitative evaluation on the Aria Digital Twin dataset further demonstrates that the pipeline maintains correct object identities through physical object manipulation, including hand occlusion and large spatial displacement.
Map-Det3D is an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video.
Yung-Hsu Yang, Luigi Piccinelli, S. R. Bulò et al.· 0 citations
This work proposes three learned matching heads: a LightGlue-style attention head with DoubleSoftmax scoring on frozen MASt3R descriptors; a DPT-style multi-scale fusion module that exposes layered spatial detail from the VGGT foundation model before pooling; and a multi-view extension that performs joint self-attention over segments drawn from several views at once, recovering transitive correspondences that strictly pairwise matchers cannot reach.
Denis Fatykhoph, Timur Akhtyamov, Konstantin Pakulev et al.· arXiv.org· 0 citations
3D scene graphs provide structured environmental representations that enable robots to perform language-grounded tasks such as navigation and manipulation. A key challenge in constructing 3D scene graphs is preserving consistent object identity. During robot navigation, newly observed 3D segments need to be associated with previously accumulated instances. Existing methods perform this association using geometric overlap and feature similarity, but these cues alone progressively fragment or merge object instances falsely. To address this challenge, we present GORI, an image-guided selective 3D object re-association framework that combines 2D multi-object tracking with selective 3D re-association. GORI employs 2D temporal tracking as the primary association mechanism and performs 3D re-association selectively to account for tracking discontinuity. To mitigate erroneous merges of spatially adjacent objects, GORI enforces a co-detection constraint that prevents merging 3D instances observed as distinct detections within the single image. The resulting 3D scene graph provides consistent object instances that serve as a grounding space for language-conditioned task planning. We evaluate GORI on HM3DSem dataset over existing 3D scene graph baselines, and assess its performance on real-world indoor scans. We demonstrate improved panoptic quality, F1 score, and average precision on HM3DSem; showing that improved object consistency supports more robust downstream task planning.
Jei Kong, Seungjae Lee, Jeewon Kim et al.· 2026 23rd International Conf...· 0 citations
A mask-aware tri-modal framework that improves the quality of superpoint representations by retrieving a scene-level structural context from a pretrained PointSAM encoder to enhance object-centric evidence and predicting a soft mask weight to suppress unreliable superpoints.
Feng Zhou, Hui Wang, Kaida Ning et al.· The Visual Computer· 0 citations
Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-based frameworks often rely on backbones pretrained for semantic recognition, and introduce 3D geometry through downstream task-specific modules. As a result, their shared representations may fail to preserve explicit metric geometry and consistent 3D scene structure. In this paper, we present a Geometry-grounded Unified 3D Perception (GeoUP) framework that adapts the reconstruction-oriented latent of VGGT to calibrated, streaming multi-camera driving scenes. GeoUP factorizes cross-image interaction into self, temporal, and view attention to capture structurally distinct temporal and cross-view correspondences. It further injects calibration-aware raymap encodings to provide metric scale and camera geometry. The resulting geometry-grounded latent is decoded for metric depth estimation, 3D object detection, and semantic occupancy prediction, corresponding to surface-, instance-, and volume-level readouts of the same 3D scene. Through joint multi-task and multi-dataset training, GeoUP effectively leverages heterogeneous annotations and generalizes across diverse sensor configurations and perception ranges. Extensive experiments on nuScenes, Argoverse 2, Waymo, KITTI, and DDAD demonstrate that GeoUP achieves SOTA performance across detection, occupancy, and depth estimation. These results validate the effectiveness of geometry-grounded representations for unified 3D driving perception.
Longfei Xu, Xiaohui Wang, Zehao Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.