A mask-aware tri-modal framework that improves the quality of superpoint representations by retrieving a scene-level structural context from a pretrained PointSAM encoder to enhance object-centric evidence and predicting a soft mask weight to suppress unreliable superpoints.
For embodied agents to navigate and reason indoor spaces, they need object-level 3D representations that stay consistent over time as new frames arrive from a monocular camera. Current online 3D instance segmentation methods either depend on posed RGB-D input with ground-truth depth or couple tightly to the internal representations of specific foundation models, sacrificing modularity. We observe that appearance-based and geometry-based object matching exhibit complementary failure modes: appearance is ambiguous among spatially separated duplicates, while geometry is unreliable for visually distinct objects at similar locations. This motivates SAM3R, a training-free pipeline that fuses spatial overlap, 3D centroid displacement, and visual-semantic similarity into a single assignment cost solved via bipartite matching. The cost is constructed entirely from the outputs of frozen foundation models without accessing internal representations. Object tracks are classified through a cascaded decision tree that detects scene changes via field-of-view gated temporal voting. On ScanNet200 and Replica, SAM3R performs competitively with methods that require architecture-specific features or additional training, despite operating in a fully online, monocular setting. Qualitative evaluation on the Aria Digital Twin dataset further demonstrates that the pipeline maintains correct object identities through physical object manipulation, including hand occlusion and large spatial displacement.
M. Mohrat, E. Derevyanka, I. Obrubov et al.· Journal of Instrument Engine...· 0 citations
LiDAR-based 3D object detection is sensitive to sparse point support and occlusion-induced incomplete bird’s-eye-view (BEV) representations, especially for pedestrians, cyclists, and distant objects. This paper asks how much accurate, target-aligned occlusion guidance can help BEV feature completion and why a geometrically estimated region-of-occlusion map (ROM) fails to reproduce that benefit. We introduce OccBEV-Oracle, a lightweight completion neck inserted between the BEV backbone and dense head of a CenterPoint-style detector. Given an occlusion mask and a density map, it selects low-density occluded tokens, aggregates visible-token context through density-weighted cross-attention, and applies spatially constrained weak-residual fusion. On the KITTI validation split, full-map completion at a strong residual coefficient reduces mean Moderate 3D AP_R40 from 59.42% to 56.65%, whereas oracle GT-ROM-guided masked completion raises mean Moderate and Hard AP_R40 to 64.31% and 60.63%. A matched-strength control shows that weakening the full-map residual recovers only part of this gain (60.97% mean Moderate), so spatial restriction contributes a further 3.34 points that residual strength alone cannot supply. The gains concentrate on Pedestrian and Cyclist and on partly occluded and middle/far-range objects. Raycasting-based estimated ROM variants remain below the baseline. Mismatch, shifted/shuffled, zero-context and mask-only controls, and a direct measurement of where the neck edits the BEV map, show that the aligned GT-ROM mask itself supplies a strong localization prior. A mask-target analysis on nuScenes confirms the failure mode is dataset-independent. OccBEV-Oracle is therefore an upper-bound analysis, not a deployable detector: accurate target-related occlusion localization remains the main bottleneck.
Jun Wang, Quanxin Zheng, Jian-Ping Yu· IEEE Access· 0 citations
This work investigates whether a frozen, self-supervised point transformer already contains the structural information required to isolate object instances without any handcrafted geometric prior, and develops a training-free segmenter that groups points via connected components on a key-similarity graph, using neither density-based clustering nor proximity priors.
Ted Lentsch, Santiago Montiel-Mar'in, Holger Caesar et al.· 0 citations
OVR-GS (Open-Vocabulary Removal in Gaussian Splatting), an instruction-driven object-removal framework for pre-trained 3D Gaussian Splatting (3DGS) scenes, demonstrates the effectiveness of localized Gaussian optimization for instruction-driven cleanup of reconstructed environments before visual inspection, presentation, or reuse as renderable virtual-scene assets.
Yongpeng Ding, Feng Ouyang, Jiawei Fan et al.· Italian National Conference...· 0 citations
This work enhances the existing iterative object-basesd visual localization approach with an additional semantic feature derived from a pretrained semantic segmentation model and conducts a systematic baseline study of contemporary feature matching techniques on such cross-domain query-reference image pairs.
Yasmin Loeper, Markus Gerke, P. Fanta-Jende· The International Archives o...· 0 citations
Results indicate that incorporating textual semantic priors can effectively enhance high-level semantic representations of point clouds, providing a feasible solution for indoor 3D scene understand.
Jinyu Tan, Juntao Yang, Yutao Zhang et al.· The International Archives o...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.