Skip to content

A Shared Encoder Is Not a Shared Task: Conditional Comparison for Deep Expert Pools

Aug 2026 · 0 citations · 8 references
Computer Science

TL;DR

It is shown that cross-evaluated heads on a frozen shared representation inherit the extrapolation confound of shallow exchange scores: pure input rotations with fixed labels inflate a deep exchange score from about 0 to 0.80, while representation-novelty scores are blind in the complementary direction.

Abstract

Sharing a deep encoder does not, by itself, fix the central confound of task-comparison scores. We show that cross-evaluated heads on a frozen shared representation inherit the extrapolation confound of shallow exchange scores: pure input rotations with fixed labels inflate a deep exchange score from about 0 to 0.80, while representation-novelty scores are blind in the complementary direction (flat under label permutations that change the task completely). Transplanting a conditional two-discriminator discrepancy into the embedding space resolves both blind spots: the functional axis stays within +-0.001 under rotations and tracks label-permutation drift mass monotonically. Built into a mixture-of-heads lifecycle, the two-axis gate attains better decision quality with fewer heads than exchange or novelty triggers at a matched training budget. On generalized category discovery, the same chunk-level functional axis separates semantic novelty from photometric shift with AUROC 0.98-0.99 where per-input OOD scores (MSP, Energy, Mahalanobis, KNN) sit near chance for that distinction. All findings replicate across frozen ImageNet-21k ViT-B/16 and self-supervised DINOv2 backbones on CIFAR-100, and extend to residual adapter pools with recurrence, where a null-calibrated novelty trigger never fires on mechanism changes while the two-axis gate handles them with full recurrence reuse. We state explicitly the common-factoring condition under which embedding-space conclusions transfer to the original mechanism.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

What Changed? Drift Detection with Real, Virtual, and Incomparable Diagnosis

Sharing a deep encoder does not, by itself, fix the central confound of task-comparison scores. We show that cross-evaluated heads on a frozen shared representation inherit the extrapolation confound of shallow exchange scores: pure input rotations with fixed labels inflate a deep exchange score from about 0 to 0.80, w...

Kentaro Oda · 0 citations
Preprint Aug 2026

DiD It in 87 Minutes: A Label-Free Softmax-to-Linear Adaptation of Vision Transformers for Object Detection

DiD is introduced, a label-free conversion method that exclusively trains the linear-attention backbone by aligning detector-facing interface tensors with those of a frozen Softmax teacher, and substantially outperforms established baselines and matches supervised, fully trained linear models.

Huai-Yuan Qin, Gabriel James Goenawan, Zihang Lin et al. · 1 citation
Preprint Sep 2026

BindCLIP: One Balanced Coupling For Compositional Vision Language Scoring

Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which region or how those regions are arranged; a model can recognize every word and object yet prefer a compositionally incorrect caption. We argue that a frozen encoder retains...

Liu-Yang Song, Yi Zhang, Zhong-Yi Deng et al. · 0 citations
#machine learning Preprint Sep 2026

Visual Jev: Accurate and Efficient Decisions from Shared Visual Context

Many vision applications ask several independent, forced-choice questions about the same image. Visual Jev encodes the image and public context once, executes isolated question suffixes as a batch, and reads candidate probabilities from the backbone's language-model head. Across four benchmarks, answer-supervised post-...

Guan-Xu Yu, Yu-Hang Yao · 5 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Which Tasks Survive Self-Supervised Learning?

Same-instance self-supervised learning (SSL) learns representations by enforcing consistency across two views of the same underlying instance. This principle alone, however, does not determine which downstream tasks remain recoverable from the learned representation. We study this question through \emph{semantic recove...

Achleshwar Luthra, Lucas Bryant, T. Zhu et al. · 0 citations
#machine learning Preprint Aug 2026

DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging

This work proposes Decoder-Aware Representation Tuning via Surgery (DARTS), which employs a novel entropy-weighted L1 loss to upweight correction at high-entropy positions where errors most affect generation quality, and a per-position additive bias that captures position-dependent error without overparameterization.

Aaryan Ajay Sharma, Sai Nishanth Padala, Seganrasan Subramanian · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.