Stabilizing YOLOv8-Based Video Fire Detection Using Diff ROI Preprocessing and Tracking-Hysteresis Postprocessing
Abstract
This paper proposes a video-based fire-detection method that adds fire-detection capability to existing conventional CCTV systems without additional hardware, mitigating the temporal instability of single-frame YOLOv8s detection by combining difference-based region-of-interest preprocessing, spatial filtering, and tracking hysteresis post-processing. Using 23 videos from the Federal University of Rio Grande (FURG) dataset and the AIHUB image-sequence dataset (318 sequences), we evaluate fire-only, smoke-only, and normal-sequence false-alarm rate at IoU 0.30 and 0.50, and report AP/mAP, FPS, sequence-level Wilcoxon signed-rank tests and bootstrap confidence intervals. For the fire-only class the proposed method improves micro-average F1 (FURG 0.6789 → 0.7060), but the per-sequence F1 gain is not statistically significant (p > 0.05). In contrast, the frame-level false-alarm rate is considerably reduced and the reduction is statistically significant over the full set of sequences (AIHUB fire p < 10⁻⁷). The smoke-only setting degrades significantly; therefore, smoke is treated as an auxiliary early warning cue rather than an independent fire criterion. Overall, the method suppresses false alarms without harming detection accuracy and serves as a lightweight add-on to the existing surveillance infrastructure.