Skip to content

Below the Noise Floor: Bimodal Seed Collapse and Distinct Failure Modes in Small-Model Knowledge Distillation

Aug 2026 · 0 citations · 20 references
Computer Science

TL;DR

On a 740-instance healthcare API routing task with a 1.5B Qwen student and a 20B teacher, eight KD variants are compared against supervised cross-entropy, finding single-seed evaluation is unable to detect central failure modes in small-model KD.

Abstract

Function routing -- selecting the correct API call from a fixed catalog given a natural-language request -- is a deployment problem where small students are attractive but knowledge distillation gains are typically reported single-seed, at scales where seed variance is unknown. On a 740-instance healthcare API routing task with a 1.5B Qwen student and a 20B teacher, we compare eight KD variants against supervised cross-entropy, using three to six seeds for key configurations. We find: (i) per-seed standard deviation ranges from 2.8 to 48.7 percentage points, swallowing every claimed KD gain below five points; (ii) three of seven KD variants exhibit bimodal collapse, with at least one in three to five seeds falling below 55% accuracy while the others train normally, and a fourth showing elevated variance; (iii) collapse has distinct modes -- wrong-function selection for ce_kd and ce_paraphrase, and a previously undocumented output-truncation mode for reasoning_kd, where the model emits reasoning but terminates before producing a function name (0.9% accuracy); (iv) only progressive_kd and rank_kd avoid collapse across observed seeds, with sigma<= 3.9 pp; (v) a naive cross-split +3.78 pp gain from input enrichment reverses to -2.70 pp under controlled within-split multi-seed re-testing. Single-seed evaluation is therefore unable to detect central failure modes in small-model KD.

View source

Similar papers

Preprint Aug 2026

One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse

The results support intervention at the shared QK locus rather than separate repair at each fault source, and support intervention at the shared QK locus rather than separate repair at each fault source.

S. Xie, Shuyang Xie, Yuan Cao et al. · 0 citations
#machine learning Preprint Sep 2026

Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation

A compressed student has two shapes that need not agree: the weight it deploys at inference and the weight family its training can reach. We show that a state-of-the-art weight-inheritance distiller, Low-Rank Clone (LRC), deploys a full-width student MLP but ties training to a teacher-induced slice, leaving 62.5-81.4% of each deployed matrix's independent linear degrees of freedom unreachable-paid for at inference, never trainable. Our principle is one line: train what you deploy. From the identical LRC warm start, we make the training object the entire deployed matrix, with no change in deployed shape, deployed parameter count, or inference FLOPs, via two mergeable realizations (Dense-LRC and CORE-LRC) that both collapse to one deployed weight. This recovers stranded capacity: taking the stronger realization per teacher, +2.36/+2.71/+10.45 Avg9 over matched-budget plain-LRC baselines across three teachers (Llama3.2-3B, Llama3.1-8B, Qwen2.5-3B), with the largest gain on the widest teacher (Qwen), where it reaches the original recipe's approx. 20B-token accuracy at 10B tokens (2x token efficiency); there the strictly same-lineage arm still recovers +6.39, the fully controlled figure. Controls strongly support attributing the gain to the enlarged reachable set, rather than to added parameters or the recipe. From approx. 10B distillation tokens plus a short SFT, a half-parameter 1.5B student matches its approx. 9T-token teacher's 9-task macro-average, within evaluation noise and with a residual MMLU deficit, and a 2.7B student beats Meta's own official compression of Llama3.1-8B at ~900x fewer compression tokens (a token count under unmatched recipes, not a compute claim). All results are from single-seed runs on the LRC backbone.

Wen-Hui Chen, Zhi-Feng Li, Jie Zhou et al. · 0 citations
Preprint Aug 2026

Routing Divergence Is Not Evidence of Behavioral Influence in Same-Weight MoE Self-Distillation

Two Mixture-of-Experts (MoE) forward passes can share every weight yet route the same token through different experts. This creates a possible blind spot in same-weight self-distillation, where a demonstration-conditioned teacher supervises a query-only student. We study this mismatch in its single-step form, with frozen weights rather than as a proxy for a full training trajectory. An exact blockwise decomposition separates a routing term, which changes gates at fixed content, from a dense-like content term. Across seven open-weight checkpoints and two domains, the routing term spans only $1.6\times$ as a fraction of block output, while its residual-stream exposure spans $3.2\times$. Exposure is ordered by the routed block's share of the residual. Scaling the always-on backbone in two confirmatory models moves exposure monotonically; common-mode controls support a mass-and-coherence mechanism rather than denominator dilution alone. Preregistered PubMedQA patches on three models show that the full routing term moves outputs by less than half the natural context effect and is largely reproduced by matched-norm noise, whereas the content term is strongly direction-specific. Scale and merged-expert probes show that the narrow block-level range is not universal, although exposure remains small at the tested boundaries. Router movement alone is therefore not evidence of behavioral influence: measure exposure first, and use a behavioral intervention when the decision matters.

Cédric Caruzzo, Donggeun Yoo, Tae Soo Kim · 0 citations
Preprint Aug 2026

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling

PluginEval is introduced, a benchmark constructed through a two-stage framework that systematically mitigates limitations of Large Language Models and evaluates five model families, including proprietary models and models with open weights, analyze their performance across difficulty levels and error categories, and validate the judge through agreement with human annotations.

Dong Xu, Julius, Hanchi Dong et al. · 0 citations
Preprint Jul 2026

When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation

Through mechanistic analysis, a mechanistic pathway consistent with imbalance promoting generalization is characterized: a mechanistic pathway consistent with imbalance promoting generalization in sufficiently capable models.

Cheng-Ting Chou, Duc Hoang · 0 citations
Preprint Aug 2026

FlowNeg: GFlowNet-Guided Diverse Hard Negative Sampling for Knowledge Graph Embedding

FlowNeg is introduced, a context-conditioned hierarchical generative flow network that amortizes reward-proportional sampling without normalizing a composite reward over the entity set: given a positive triple and corruption side, it selects a type, then an entity.

Ibne Farabi Shihab, Naoshin Anzum Hridi, Joyanta J. Mondal · 0 citations

Related blog posts

Google DeepMind Blog Aug 12, 2026

Putting sign language AI into users’ hands

Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.