Skip to content
Conference

Improving Robustness of Semantic Segmentation for Autonomous Driving: A Case Study

Jul 2026 · International Conference on Artificial Intelligence Testing · pp. 156-163 · 0 citations · 19 references

Abstract

Deep neural networks (DNNs) have achieved remarkable success in recent years and are increasingly integrated into safety-critical systems such as autonomous driving vehicles. However, when deployed in real-world environments, their robustness to common input corruptions remains a major challenge for safety assurance. Corruptions such as motion blur can change the outputs of DNN-based semantic segmentation models and, more importantly, cause unsafe system-level decision inconsistencies, for example by failing to identify ground obstacles that are correctly recognized under clean conditions. In this paper, we present a testing-oriented robustness repair approach for semantic segmentation models in real-world industrial settings. We first use corruption-based testing to reveal decision-level failures under realistic perturbations, and then repair the model through a combination of data augmentation and self-training using only unlabeled data. Rather than focusing solely on pixel-level prediction changes, our method targets the reduction of system violations while preserving decision behavior on clean inputs. We evaluate the approach on a semantic segmentation model used in an industrial product. Experimental results show that our method significantly reduces system violation rates while maintaining system-level decision-making accuracy, demonstrating the practical value of testing-guided repair for safety-critical deployment.

View source

Similar papers

#machine learning Preprint Sep 2026

Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language

Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness, and reliable operation across domains. For example, domain shift between locations could lead to the operating environment being misaligned with the training data, resulting in potentially dangerous performance degradation. Yet, existing data analysis pipelines largely rely on metadata, predefined labels, or manual inspection, which provide limited semantic insight or do not scale. This paper studies set difference captioning: given two subsets of images, the goal is to produce a natural-language hypothesis describing differences between the target and reference set. Building on a two-stage formulation, we adapt the method to autonomous driving by focusing on object-centric patches derived from object detection, which simplifies aggregation and enables attribution of differences to specific object instances or categories. To evaluate this setting in-domain, we introduce a new benchmark, AD-Diff Bench. Low-concentration experiments assess the suitability of set-difference-captioning approaches to sparse, real-world differences. We restrict our experiments to open-weight models to support reproducibility and ease of deployment. The proposed benchmark and analysis provide a step towards practical, human-interpretable dataset introspection for autonomous driving datasets. Our implementation and benchmark dataset are available at https://github.com/KIT-MRT/AD-Diff

Julian Truetsch, Felix Hauser, Christoph Stiller et al. · 0 citations
Preprint Aug 2026

SeFaR: Semantic Feature-aware Robustness Testing of Deep Neural Networks

Deep neural networks are increasingly deployed in safety-critical domains as perception modules, where failures are often caused due to rare and under-represented scenarios. This necessitates the need to evaluate the semantic robustness of perception models; conformance of behavior to high-level requirements over real-world perceptual variability. To address this, we propose SeFaR, a framework for systematic semantic-feature-centric testing of vision models. Given a natural-language requirement and a set of satisfying inputs, SeFaR evaluates robustness with respect to diverse realistic semantic variations that preserve requirement satisfaction. The approach employs a novel hierarchical concept model enabling structured exploration of the feature space and incorporation of domain knowledge via user-defined concepts. State-of-the-art diffusion and vision-language models are leveraged to generate photorealistic semantics-preserving perturbations and identification of previously unknown features impacting behavior. A feedback-driven adaptive process is adopted to generate interpretable failure-inducing semantic concepts along with corresponding test inputs. Evaluation on case studies demonstrates that the proposed framework effectively satisfies requirement preconditions while identifying requirement-independent features that influence model decisions, enabling it to both uncover faults and relate them to such features.

Nusrat Jahan Mozumder, Divya Gopinath, Corina S. Păsăreanu et al. · 0 citations
Preprint Aug 2026

VOLA: Improving Open-World Driving by VLM-Based Semantic Attribute Prediction

Driving in the real world is open-world: a car may encounter a fallen mattress, a deer, or other objects outside its training data. Naming them is not enough. The system must know how to treat each region: can it drive over it, and how severe would a collision be? We therefore shift scene perception from category labels to dense action-relevant attributes, where each pixel is labeled by how it should affect motion rather than by object name. We instantiate this general formulation with two ordered attributes: 7-rank drivability and 5-rank vulnerability. We read Qwen3.5 image-token hidden states directly as a spatial semantic representation. A lightweight boundary-aware decoder then turns this coarse token grid into sharp full-resolution attribute maps. The whole process requires neither autoregressive text generation nor an external mask model such as SAM. We train on dense attribute labels built in CARLA and test transfer to real scenes and to novel obstacles never seen in training. We compare with vision-only segmenters trained on the same attributes and prompted VLM segmenters. Our model matches strong vision-only segmenters on familiar categories and improves transfer to real open-world anomalies, reaching 69.4% mean vulnerability-rank recall versus 57.1% for the best vision-only baseline and 53.9% for the best prompted VLM baseline. These results show that VLM image tokens provide useful semantic cues for transferring driving attributes to objects outside the training vocabulary.

Yuchen Zhang, Yuan Gao, Sebastian Schmidt et al. · 0 citations
Preprint Aug 2026

CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios

CAViAR exposes a practical Perception--Reasoning Gap: current VLMs may recognize salient context, but do not reliably map visible agent actions to annotated rule-relevant responsibility categories in safety-critical driving scenarios, and all models degrade sharply on accident type and responsibility reasoning.

Sparsh Garg, Yi-Wen Chen, Vijay Kumar et al. · 0 citations
Open access Jul 2026

Enhancing Vision-Based Perception in Autonomous Driving: YOLO11–DETR Integration with Selection Model

Abstract. Vision-based object detection is a key component of autonomous driving perception systems; however, models pretrained on large-scale generic datasets usually struggles when implemented in automotive environments due to domain shift. This research introduces a comprehensive evaluation and fusion of YOLO11 and RT-DETR for improving robustness in autonomous driving scenarios using KITTI dataset. Both models are pretrained on COCO dataset and evaluated under a zero-shot transfer setting to assess cross-domain generalization. The results show that RT-DETR-L and RT-DETR-XL experience significant performance degradation, dropping from 53.0 and 54.8 𝑚𝐴𝑃 on COCO to 34.3 and 34.5 on KITTI, respectively. In contrast, YOLO11-Nano and YOLO11-L demonstrate better generalization, achieving 44.0 and 51.3 𝑚𝐴𝑃 on KITTI compared to 40.9 and 55.0 on COCO. Controlled fine-tuning experiments (10 and 100 epochs) are conducted to analyze adaptation dynamics. RT-DETR-L improves to 61.7 and 79.1 𝑚𝐴𝑃, while YOLO11-L reaches 64.9 and 76.2 after 10 and 100 epochs, respectively. To further evaluate robustness under challenging conditions, three degraded data subsets are generated. Building on the strengths of convolutional and transformer-based detectors, this work introduces an image-based selection model that selects the most suitable detector for each input image. Experimental results demonstrate substantial zero-shot degradation, strong recovery after fine-tuning, and consistent performance improvements under degraded conditions using the proposed selection strategy. Our method achieves gains of up to 4 𝑚𝐴𝑃 points over the best standalone detector without incurring the computational overhead. The proposed framework provides a context-aware and computationally efficient perception enhancement strategy suitable for real-world autonomous driving systems.

Ahmed M. Reda, Naser El Sheimy, Adel Moussa · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.