Autonomous Driving Systems (ADS), such as Apollo and Autoware, are complex cyber-physical platforms that fuse perception, planning, and control to ensure safe vehicle operation in dynamic environments. Achieving high safety assurance demands rigorous testing across diverse and challenging driving conditions. However, due to the inherent incompleteness of scenario-based testing, it remains essential to evaluate how effectively scenario datasets exercise system behavior and expose potential failures. This paper introduces a set of scenario-level coverage metrics for ADS testing that characterize temporal combinations of driving maneuvers. Our key insight is that failures often emerge from specific sequences of maneuvers rather than from individual actions in isolation. Using a large-scale scenario dataset, we assess the capability of these metrics to reveal system violations and unsafe behaviors. Experimental results show that 4-way sequential coverage achieves a 100% detection rate for violations and threats, significantly outperforming both 4-way combinatorial coverage (35%) and a state-of-the-art metric, ComOpt (≈ 30%). Overall, our findings highlight that incorporating temporal maneuver sequences yields a more rigorous and sensitive measure of test adequacy for autonomous driving systems.
The rapid integrationof artificial intelligence (AI) into everyday life has brought transformative benefits, but it has also sharpened concerns surrounding safety, security and robustness, particularly in safety-critical domains. This special issue seeks to address these interconnected challenges in a holistic manner, recognizing that dependable AI systems must be not only performant, but also trustworthy across a wide range of operational conditions. The articles bring together contributions in the fields of computer science and mathematics, while also addressing concerns about ethics, accountability and regulation.
This article is part of the theme issue ‘Safe, secure and robust AI for safety-critical systems’.
Ajitha Rajan, D. Higham· Philosophical Transactions o...· 0 citations
This work surveys real-world TUI applications, turns them into a headless benchmark spanning ratatui/Rust, bubbletea/Go, textual/Python, and ink/TypeScript, and compares four frontier LLMs with random exploration, finding no model dominates.
Chao Peng, Ruida Hu, Ajitha Rajan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.