SAM-Anchored DINOv3 Soft Calibration for Weakly Supervised Semantic Segmentation
Abstract
Weakly supervised semantic segmentation (WSSS) aims to train dense prediction models from inexpensive supervision such as image-level labels. Recent foundation models provide complementary capabilities: promptable segmentation models can produce high-coverage object masks, while self-supervised vision transformers provide discriminative dense representations. However, directly modifying or filtering pseudo-labels can reduce supervision coverage and degrade downstream student training. This paper proposes DINO-SAM SoftCal, a SAM-anchored DI-NOv3 reliability calibration framework for WSSS. The proposed method preserves pseudo-labels generated by a frozen SAM 3.1 model and uses frozen DINOv3 features to construct class-aware prototypes. Instead of changing semantic targets, DINOv3 prototype agreement is used as a soft loss calibration signal for training a DeepLabV3+ student. On PASCAL VOC 2012, the proposed method improves on average over both SAM-only training and a same-checkpoint SAM continuation control. We also reproduce ToCo as an external WSSS reference baseline. On an exploratory MS COCO 2014 subset, the same calibration strategy also improves over the corresponding SAM continuation control. These results suggest that, in our setting, using DINOv3 as a conservative reliability calibrator is an effective way to integrate dense foundation-model features without sacrificing pseudo-label coverage.