Retrieval-Augmented Multimodal Large Language Models for Visual Question Answering of Construction Occupational Health and Safety Hazards
TL;DR
A visual knowledge enhancement framework for construction OHS visual question answering (VQA) based on multimodal large language models (MLLMs) and retrieval-augmented generation (RAG) is proposed, augmenting managerial capacity for reliable and objective OHS hazard prevention.
Abstract
Construction occupational health and safety (OHS) hazard oversight is a critical pillar of engineering management, requiring the complex integration of dynamic visual evidence with rigorous regulatory standards. Traditional oversight, heavily reliant on manual inspections, is labor intensive and prone to cognitive omissions. While automated hazard detection has evolved, existing paradigms remain constrained by closed-set recognition, failing to simulate the open-ended, heuristic reasoning of safety experts. To bridge this gap, this study proposed a visual knowledge enhancement framework for construction OHS visual question answering (VQA) based on multimodal large language models (MLLMs) and retrieval-augmented generation (RAG). It can reliably respond to site managers’ open questions about OHS hazards in construction images. The method’s primary innovation lies in a tailored RAG framework with a knowledge base for construction OHS hazard VQA, addressing three critical challenges: cross-modal semantic misalignment, knowledge demand variability, and information overload–induced cognitive bias. To enable systematic evaluation of the framework and mitigate the lack of public benchmarks, we designed experiments across three question types and built the Construction Hazard VQA Dataset (ConHazard-VQA), the first dedicated dataset for construction OHS hazard VQA, featuring 1,034 high-quality image-question-answer pairs. Experiments across classification, counting, and open-ended question types confirmed consistent and significant performance gains over the baseline. By introducing a tailored, practical framework that translates a general-purpose MLLM into a domain-specific expert for construction OHS hazard VQA, a new paradigm for data-driven safety management was established. This research advances the body of knowledge in engineering management by transitioning automated oversight from rigid pattern matching to expert-like, diagnostic decision support, thereby augmenting managerial capacity for reliable and objective OHS hazard prevention.