HRFNet: Ground-Truth-Guided Kullback–Leibler Divergence Routing for Semantics-Aware Multi-Branch Segmentation of Urban Driving Scenes
Abstract
Multi-branch networks combine backbones whose inductive biases are complementary, but their fusion modules set the branch weights from feature statistics alone and are therefore blind to what a pixel represents. We supervise fusion in label space instead. Ground-truth labels are converted into per-pixel routing targets, which are imposed on the branch weights through a Kullback–Leibler divergence term. Each pixel is thus routed towards the branch suited to its category: a convolutional branch for fine-grained boundaries, a state-space branch for large homogeneous regions, and a windowed-attention branch for intermediate-scale context. The mechanism is instantiated in HRFNet, a three-branch encoder with branch-specific dilation rates. HRFNet attains 76.85% and 79.88% mean intersection over union (mIoU) on Cityscapes and CamVid, averaged over five seeds, and 58.18% and 46.74% on the 59-class PASCAL Context and the 150-class ADE20K benchmarks, exceeding every baseline retrained under an identical budget. Ablations locate the gain in the routing rather than in added capacity: routed fusion adds 2.29% mIoU over average fusion, of which 1.53% comes from the label-space target itself, for 3.8M parameters and 8.8% of the network’s operations. All nine ablation contrasts remain significant after Holm–Bonferroni correction, and the gain persists in every two-branch configuration.