LIO-YOLO: A Lightweight Object Detection Framework for Indoor Robotic Perception
Abstract
Indoor medium-to-large indoor object detection is a fundamental capability for mobile robot perception and visual navigation, yet it remains challenging in complex indoor environments due to occlusion, overlap, cluttered backgrounds, and the need for efficient deployment on resource-constrained devices. To address these challenges, this paper presents LIO-YOLO, a lightweight object detection framework developed from YOLO11n. A new MDA-C3k2 module is integrated into the backbone to improve multi-scale feature extraction and contextual information modeling. The Slim-neck design is further employed to achieve more efficient feature fusion with less redundant computation. In addition, SEAM-Head is introduced to improve the detection of occluded and overlapping objects, which enhances the model’s performance in complex indoor environments. To better reflect the characteristics of large-object detection in real indoor environments, an Indoor Large Object Dataset (ILOD) was constructed in this study. Experimental results show that LIO-YOLO achieved 84.9% precision, 80.7% recall, 88.6% mAP@0.5, and 58.6% mAP@0.5:0.95. Furthermore, the proposed model maintained only 5.4 GFLOPs, demonstrating a good trade-off between detection accuracy and computational cost. Moreover, deployment on the RK3588 edge platform reduced inference time from 11.8 ms to 9.5 ms and improved throughput from 84.8 FPS to 105.3 FPS. These results indicate that LIO-YOLO demonstrates its effectiveness for real-time indoor robotic perception applications.