Early Recognition of Workplace Hazards Using Data-Efficient Multimodal Learnable Prompting and Parameter-Efficient Fine Tuning
Abstract
Timely and accurate recognition of workplace hazards is critical for ensuring safety in dynamic and high-risk environments such as construction sites. However, existing video-based approaches often rely on extensive annotations and fully supervised training, and many are tailored to specific hazard types, which limits scalability to rare or diverse scenarios. This study investigates whether a data-efficient vision–language framework can support early-stage hazard recognition from short preincident video segments under limited supervision. The proposed method builds on a video-adapted contrastive language-image pre-training (CLIP)-based vision–language backbone and introduces a parameter-efficient adaptation strategy that combines learnable visual prompting with low-rank adaptation (LoRA) applied to the text encoder. Learnable visual prompts capture global, summary, and local spatiotemporal cues from short video clips, and LoRA enables lightweight semantic adaptation without fine-tuning the full backbone. This design preserves most pretrained parameters and introduces only a small number of additional trainable parameters, making it well suited to few-shot learning in data-scarce safety scenarios. To evaluate early recognition capability, this paper further examines performance under shorter temporal observation windows by reducing the amount of video available for prediction. Experiments on a curated real-world hazard video data set show that the proposed method substantially improved few-shot recognition compared with a prompt-free baseline across five hazard categories and 5-, 10-, and 15-shot supervision. Overall, the findings suggest that combining temporal visual prompting with parameter-efficient text adaptation offers a scalable engineering pathway toward assistive hazard recognition in data-scarce environments, with promising but preliminary performance.