Jul 2026
Mask-aware tri-modal learning for indoor 3D object detection
A mask-aware tri-modal framework that improves the quality of superpoint representations by retrieving a scene-level structural context from a pretrained PointSAM encoder to enhance object-centric evidence and predicting a soft mask weight to suppress unreliable superpoints.
Feng Zhou, Hui Wang, Kaida Ning et al.
· The Visual Computer · 0 citations