GORI: Image-Guided Selective 3D Object Re-Association for 3D Scene Graphs and Task Planning
Abstract
3D scene graphs provide structured environmental representations that enable robots to perform language-grounded tasks such as navigation and manipulation. A key challenge in constructing 3D scene graphs is preserving consistent object identity. During robot navigation, newly observed 3D segments need to be associated with previously accumulated instances. Existing methods perform this association using geometric overlap and feature similarity, but these cues alone progressively fragment or merge object instances falsely. To address this challenge, we present GORI, an image-guided selective 3D object re-association framework that combines 2D multi-object tracking with selective 3D re-association. GORI employs 2D temporal tracking as the primary association mechanism and performs 3D re-association selectively to account for tracking discontinuity. To mitigate erroneous merges of spatially adjacent objects, GORI enforces a co-detection constraint that prevents merging 3D instances observed as distinct detections within the single image. The resulting 3D scene graph provides consistent object instances that serve as a grounding space for language-conditioned task planning. We evaluate GORI on HM3DSem dataset over existing 3D scene graph baselines, and assess its performance on real-world indoor scans. We demonstrate improved panoptic quality, F1 score, and average precision on HM3DSem; showing that improved object consistency supports more robust downstream task planning.