Skip to content
Open access

ST-PaveCLIP: A Spatio-Temporal Vision–Language Framework for Road Anomaly Segmentation in Images and Videos

Sep 2026 · Remote Sensing · 0 citations · 18 references

Abstract

Static images and vehicle-mounted video are the two primary data sources for road inspection. Since cracks, potholes, and patched areas can all be considered anomalies on the road surface, Contrastive Language–Image Pre-training (CLIP)-based anomaly segmentation provides a promising approach under limited labeled data. However, two challenges remain in practical applications: whole-image resizing may weaken fine anomalous structures, while weak and irregular damage regions can exhibit spatially varying prediction difficulty; for video input, frame-wise prediction often produces inter-frame flickering. This paper proposes ST-PaveCLIP for road anomaly segmentation in images and videos. For single images, we introduce a training-time residual-scale auxiliary supervision based on heteroscedastic negative log-likelihood and adopt a local–global dual-scale inference scheme to preserve global road context and fine anomaly structures. For video input, a temporal alignment module based on RoMa v2 dense matching and homography estimation warps the previous fused probability map to the current frame before temporal fusion, without additional video-level training. Multi-seed experiments on public and self-collected datasets show that ST-PaveCLIP improves the principal segmentation metrics over the AA-CLIP baseline and remains competitive with supervised baselines under the same limited annotation budget. Video experiments further show reduced inter-frame inconsistency in Aligned Temporal Consistency Error (TCE) and Aligned Threshold-Crossing Rate (ATCR), with a modest additional improvement in key-frame segmentation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.