Aug 2026· IEEE Robotics and Automation Letters· Vol 11, pp. 10050-10057· 0 citations· 29 references
Computer Science
Abstract
Existing neural SLAM and 3D Gaussian Splatting (3DGS) SLAM systems often suffer from insufficient observations during indoor turn-arounds and large viewpoint changes, leading to incomplete coverage and missing details in corner regions. We propose IMGS-SLAM, a monocular Gaussian SLAM system tailored for indoor reconstruction, which improves mapping completeness and visual detail fidelity using only RGB images while maintaining competitive tracking accuracy. The method leverages a learning-based dense SLAM frontend to provide camera poses and dense geometric priors, and adopts 3DGS as the map representation. We propose a coverage-aware Dual-Cue mapping-frame selection strategy that decouples tracking keyframes from mapping frames and selects views with high expected coverage gain. This design improves map completeness under indoor turn-around motions and large viewpoint changes, while online-to-offline refinement further improves visual consistency and local details. In addition, we introduce quadtree-guided structured initialization and a high-frequency weighted loss to enhance textures and edge details, and incorporate co-visibility-constrained densification and pruning to reduce artifacts. Experiments on Replica, TUM RGB-D, and ScanNet demonstrate improved map completeness and competitive rendering quality in both synthetic and real indoor scenes.
Recent advances in SLAM have leveraged 3DGS for photorealistic reconstruction and novel view synthesis. However, most methods rely on RGB-D input, which is unavailable on consumer-grade smartphones, and few integrate 3DGS within a collaborative framework. Therefore, we present CGS-SLAM, a hybrid decentralized/centralized system enabling multi-agent 3DGS SLAM using only RGB and inertial data. Each agent performs local tracking with inertial data as a motion prior and reconstructs a scaled map using a metric monocular depth estimator (Depth Pro). Keyframe encodings are shared among agents, enabling dynamic keyframing in regions of spatial overlaps with other agents, enhancing submap alignment. Afterwards, a central server aligns submaps using VGGT as a view alignment model. This bidirectional communication keeps communication cost low during mapping and global reconstruction in difficult GNSS-denied environments. Experiments on multiple datasets demonstrate competitive tracking performance, improved rendering quality over state-of-the-art methods, and accurate submap alignment.
Jean-Daniel de Ambrogi, Aladine Chetouani, Vincent Nguyen et al.· 0 citations
RGB-D SLAM systems based on 3D Gaussian Splatting (3DGS) often suffer from map degradation caused by diminishing historical supervision, noisy depth observations, and local-window optimization during online reconstruction. To address these issues, we propose DUDG-SLAM, a Dynamic Replay and Depth-Uncertainty-Guided Gaussian SLAM framework. The proposed method selectively replays historical keyframes according to reconstruction error, forgetting degree, and Gaussian visibility, thereby restoring supervision in previously reconstructed regions. In addition, RGB-D sensor depth is retained as the primary geometric constraint, while scale-aligned Depth Pro predictions are introduced only in regions with missing or unreliable sensor depth. A pixel-wise uncertainty-weighted log-depth loss is further designed to reduce the influence of noisy depth observations and unreliable depth boundaries. Experiments on Replica and TUM RGB-D show that DUDG-SLAM improves rendering quality, trajectory accuracy, and reconstruction robustness. Compared with the baseline, DUDG-SLAM consistently improves PSNR and SSIM while reducing LPIPS and ATE.
Jing-Wen Liu, Tao Zuo, Aibo Tian et al.· Italian National Conference...· 0 citations
SLAM systems based on 3D Gaussian Splatting (3DGS) have recently demonstrated promising reconstruction accuracy for dense 3D scene representations. However, current 3DGS systems struggle to meet the strict demands of real-world deployments due to severe limitations in operational performance and map adaptability. To this end, we propose LightSplat, a hybrid-representation RGB-D SLAM framework. It synergizes local sparse features for robust and fast tracking with a dual-thread backend that progressively constructs dense Gaussian submaps. Crucially, we enable online loop closure through feature-accelerated 3DGS registration, refining overall map consistency through pose graph optimization. Ultimately, LightSplat achieves the online reconstruction of high-fidelity Gaussian map. Extensive experiments on multiple datasets and real-world robotic platform demonstrate that our method achieves near state-of-the-art reconstruction quality and the capability to accommodate practical camera motions, maintaining an average framerate of 8 FPS. Overall, LightSplat provides an efficient and robust foundation for deploying high-fidelity 3DGS in real-world environments.
Jun-Ze Bao, Ye Gao, Yi-Ming Huang et al.· 0 citations
This paper proposes LV-GS SLAM, a novel system that integrates LiDAR and visual data for incremental, large-scale reconstruction with real-time tracking, and develops a keyframe-based submap management framework that dynamically adjusts memory allocation based on both primitive density and inter-frame overlap ratio, effectively preventing GPU memory overflow.
This work revisits 3D Gaussian Splatting heuristics in a decoupled 3DGS-SLAM setting and proposes three geometry-aware methods that operate in the mapping thread: transmittance-preserving densification, camera-aware scale initialization from depth and intrinsics, and error-guided densification that focuses new primitives on high-residual regions.
Thai Luu, Quan Tran, Hieu Phan et al.· 0 citations
To address pose estimation drift and dynamic artifacts in dense mapping caused by moving-object interference in visual SLAM, this paper proposes SGA-SLAM, a robust RGB-D SLAM system tailored for complex indoor dynamic environments. The system is built upon the ORB-SLAM3 framework and incorporates the lightweight instance segmentation model YOLO11n-seg to obtain semantic priors of dynamic objects and perform preliminary feature screening within potentially dynamic regions. Furthermore, a cascaded geometric screening mechanism that integrates epipolar geometric constraints with adaptive depth consistency verification is designed to further improve the accuracy of dynamic feature discrimination, thereby effectively removing dynamic features while retaining stable static features. Experimental results demonstrate that SGA-SLAM substantially reduces the absolute trajectory error in highly dynamic indoor scenes from the TUM RGB-D and Bonn RGB-D Dynamic datasets. Specifically, the ATE RMSE is reduced by more than 94% across all TUM RGB-D Walking sequences, and the proposed method achieves higher localization accuracy with lower trajectory-error dispersion on most Bonn RGB-D Dynamic sequences. Moreover, the proposed method effectively suppresses dynamic artifacts and point-cloud contamination in dense mapping, generating static environment maps with clearer structures and improved geometric consistency.
Xiao-Xuan He, Xiao-Hui Zhang, Jin-Feng Zheng et al.· Engineering Research Express· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.