1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

An integrated YOLOv11 action recognition and vision-language model framework for intelligent kindergarten safety monitoring

Children's safety is the primary and core issue that kindergartens face. Traditional video surveillance is the main way for kindergartens to ensure safety, but this method highly relies on manual monitoring. It is not only inefficient but also prone to overlooking risks due to human negligence, and it is even more difficult to achieve real-time risk early warning. To solve these practical problems, we propose a real-time intelligent monitoring, safety early warning, and analysis framework based on a multi-modal large model. This framework integrates computer vision technology and large vision-language models, enabling both multi-dimensional scene perception and intelligent safety early warning and scene analysis. Specifically, for the needs of personnel identification and tracking, we fine-tuned the YOLOv11 model with a dedicated dataset to achieve high-precision real-time personnel detection; paired it with the ByteTrack algorithm to complete multi-target tracking, and then used the fine-tuned S3D network to identify children's dangerous actions or abnormal behaviors-grade these behaviors according to their danger levels and trigger corresponding preliminary early warnings. Then,we transmit these early warning results and scene images to the Qwen3-VL model for scene-level risk reasoning and finally generate a complete safety analysis report. We conducted tests in real kindergarten scenarios, and the results show that this framework can quickly and accurately detect personnel, identify various actions, issue safety early warnings in a timely manner, and complete intelligent scene analysis, meeting the actual usage needs of kindergartens.

Xiaojian Rao, Lin Fan, Yong Tian et al. · 0 citations