CDU-YOLO: A Scene-Aware Real-Time Smoke and Flame Detection Framework for High-Rise Building Fire Safety
Abstract
Reliable optical sensing of smoke and flames in high-rise buildings is challenging due to weak early cues, vertical smoke diffusion, facade occlusions, nighttime illumination, and fire-like urban interferences. We propose CDU-YOLO, a scene-aware real-time detection framework built upon YOLOv8n. Rather than relying on indiscriminate network scaling, task-oriented integration of existing modules is introduced: dynamic point-sampling (DySample) to preserve blurred boundaries of distant micro-targets, an enlarged receptive field (UniRepLKNet) to capture large-scale vertical propagation, and a dynamic bounding-box regression loss (WIoU) to handle occlusions. Experiments on a custom high-rise fire dataset and two public datasets demonstrate 94.9% mAP@0.5 and 56.7% mAP@0.5:0.95. In a dedicated flame-only size-stratified evaluation, CDU-YOLO improves AP@0.5 for small flames from 79.6% to 91.7% and reduces their miss rate from 25.2% to 11.3% relative to YOLOv8n. Under a unified desktop protocol (RTX 3080, PyTorch FP16, 640×640, batch size 1, no TensorRT), end-to-end throughput increases from 41 FPS to 55 FPS. A separate Jetson Orin NX deployment benchmark reaches 92 FPS using TensorRT FP16. The explicit introduction of an “others” category during training contributes to reducing false positive predictions against fire-like distractors. These results support the use of CDU-YOLO as a supplementary visual sensing component for early situational awareness. Nevertheless, residual misses on small and ultra-distant flames, continuous video-stream validation, and long-term field testing remain to be addressed before safety-critical online deployment.