As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon trajectories. In multi-turn interactions, malicious intent can be decomposed across seemingly harmless turns and gradually reconstructed through interaction trajectories, eventually resulting in safety failures. Existing safeguards remain largely reactive, detecting manifested violations while lacking the ability to predict latent risk evolution and enable preemptive prevention. To address this limitation, we propose Recast, a safety risk forecasting framework that advances LLM safeguarding beyond turn-level violation detection to trajectory-level risk prediction. Recast first retrieves risk-relevant evidence from both short-term dialogue progression and long-term historical context via a dual-scale trajectory view. It then models compositional risk evolution by capturing the current risk configuration and its temporal dynamics. Finally, a causal temporal encoder learns latent risk evolution patterns and predicts the distribution of future risk emergence turns. Extensive experiments across 7 risk categories show that Recast predicts 88.3% of future safety failures with an average lead time of 2.41 turns, while maintaining a false alarm rate of 12.3%, showcasing the effectiveness of trajectory-level forecasting in identifying emerging risks before safety violations occur.
Shi Lin, Peng Qian, Ding-Hao Liu et al.· arXiv.org· 0 citations
This paper presents a semantic mapping system that integrates Vision-Language Models (VLMs) with Simultaneous Localization and Mapping (SLAM) to enhance the understanding of indoor environments. While traditional SLAM systems primarily focus on geometric occupancy, they often lack the semantic context needed for high-level robotic tasks. The proposed framework bifurcates the mapping process into two synchronized subsystems: Geometric SLAM and VLM Recognition. The first system uses LiDAR and odometry data to perform real-time localization and mapping, providing a stable and accurate environment for semantic data integration. Simultaneously, the second system leverages a VLM with a ZED2 depth camera to identify furniture and perform precise object size estimation, mapping these semantic entities into the global coordinate frame. A data fusion layer then synthesizes these inputs into a comprehensive Semantic Map that includes both geometric structures and physical object dimensions. Experimental results demonstrate that our integrated approach achieves precise spatial anchoring of semantic entities and improves overall mapping stability, providing a robust foundation for robot environmental understanding in complex indoor scenarios.
HalluProp, a Propagation-aware Hallucination inference framework that estimates individual agent failures and emergent system-level hallucination risks before inter-agent interaction, and effectively complements post-hoc methods, highlighting the potential of pre-hoc risk inference for building more reliable multi-agent systems.
Shi Lin, Chenpei Wang, Peng Qian et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.