A VR-based Human Robot Interaction framework for studying how VLM-assisted robots can support situational awareness in simulated hazardous environments suggests that combining VLM-based robotic perception with immersive visualization is a promising approach for supporting situational awareness in hazardous settings.
Abstract
In high-risk environments such as disaster response, situational awareness depends not only on detecting hazards but also on communicating them clearly to human operators. Vision Language Models (VLMs) have shown strong potential for scene understanding in safety-critical settings, yet their value as part of human-facing robotic systems remains underexplored. We present a VR-based Human Robot Interaction framework for studying how VLM-assisted robots can support situational awareness in simulated hazardous environments. In our system, a robot explores a virtual scene and queries a VLM to identify potential hazards and annotate user-facing points of interest. These annotations are presented to a human operator through an immersive VR interface. This framework enables controlled evaluation of both robotic hazard identification and the communication of safety-critical information to users. Results from our study indicate that the annotated VR interface was preferred over the unannotated baseline and that participants reported high clarity, usefulness, and comfort when interacting with the system. These findings suggest that combining VLM-based robotic perception with immersive visualization is a promising approach for supporting situational awareness in hazardous settings.
This work evaluates several state-of-the-art VLMs across two datasets and multiple prompting strategies to test whether an explicit distinction between hazard and anomaly changes model behavior, and shows that explicitly separating anomaly from hazard provides a more informative evaluation of VLM safety reasoning and exposes failure modes that binary safety judgments can obscure.
M. Indukuri, Mohammad Eskandari, Sree Nitya Kollu et al.· arXiv.org· 0 citations
It is argued that agentic AI should be approached as a socio-technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms.
Timothy Merritt, Alejandro Jarabo-Peñas, Juan Bravo-Arrabal et al.· 0 citations
Mobile phones and other smart displays, such as in-car entertainment units, are widely used to display and dictate navigation routes. These navigation systems offer alerts about road closures, construction zones, and traffic conditions. However, these navigation applications often fail to dynamically update safety information in response to rapidly changing threats, such as public safety incidents. The current study explores virtual cues as an intermediate step toward developing augmented reality (AR) safety navigation aids. In a virtual reality (VR) environment simulating a safety–critical scenario, we evaluate two types of navigation cues: (1) a world-fixed localized, egocentric ground-area (GA) cue that visually highlights danger zones and (2) a screen-fixed allocentric minimap (MM) offering a global bird’s-eye view of the danger areas and surroundings. The cues were evaluated using a free-navigation goal directed wayfinding task, in which participants were asked to navigate to a target beacon, answer affective questions providing trial-by-trial evaluation of affective states (i.e., spatial anxiety), and then asked to point to the last known threat location to evaluate spatial memory. Results show both cues were associated with minimal time in danger zones, but the minimap significantly decreased spatial anxiety ratings and increased participants’ reports of willingness to walk the same route in the real world, compared to the GA cue. Further, higher self-reported navigation ability predicted better spatial memory within the virtual environment. These findings represent an initial step toward understanding how virtual cue design affects navigation behavior, affective states, and spatial memory in safety–critical environments.
Ashley M. Buzard, Yu Zhao, Monika Lohani et al.· Cognitive Research· 0 citations
Climate change is increasing the severity and unpredictability of natural disasters. In time-critical crises such as wildfires, traditional monitoring practices remain limited by coverage, cost, and personnel risk, paving the way for autonomous and adaptive monitoring solutions. Within this context, this paper introduces D3ARC, an asynchronous distributed hierarchical framework for time-aware and reliable wildfire detection. D3ARC integrates multiple robotic agents that cooperate under uncertainty through distributed perception, shared situational awareness and coordinated actions. A remote controller asynchronously decides upon each robot's motion, while each robotic agent senses the environment and decides where and how to execute the wildfire detection. All robotic operations require time, and as time progresses, wildfires continue to spread, reducing the opportunity for early intervention. As such, all agents share a common objective: to detect a wildfire with a certain performance threshold as fast as possible and within a time limit. D3ARC integrates mechanisms for safe navigation, coverage efficiency, cooperation and reliability. It introduces a forward-looking capability that allows agents to anticipate the future by evaluating candidate strategies before execution. The framework is evaluated through realistic robotics simulations, ablation studies, and baseline comparisons, achieving an overall mission success up to 94% with 89.4% detection confidence.
Nikolaos Koursioumpas, Lina Magoula, Nancy Alonistioti et al.· 0 citations
This paper analyzes key technological developments, system design approaches, and operational frameworks in areas such as disaster response, autonomous surveillance, firefighting, and emergency medical support and proposes a staged autonomy framework incorporating perception, cognition, control, and coordination modules.
Hiroshi Tanaka· International Journal of Int...· 0 citations
Automated vehicles may encounter situations at system limits that require external human support without continuous remote driving. We present a lab-based remote assistance setup for investigating planning-oriented interaction methods. The setup combines real-world driving video stimuli, contextual dashboard information, a pen-based planning display, and instrumentation for logging interaction and attention data in a multi-screen operator workplace. It supports two assistance concepts: collaborative planning, in which operators select system-generated maneuver suggestions, and trajectory guidance, in which operators draw a desired path for the vehicle. We conducted a formative within-subjects evaluation with 35 participants to assess the setup and explore perceived usability and user experience of two implemented interaction conditions. No statistically significant differences were found between the two implemented conditions in perceived usability or user experience. The findings suggest that both concepts warrant further investigation and motivate future setup iterations that combine interaction methods according to scenario demands and operator preferences.
Alexandra Nick, Leon Lamesic, Barbara Deml· Message Understanding Confer...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.