Skip to content

Covariance Last-Layer Ensembles: Function-Space Diversity for Efficient Uncertainty Quantification

Jul 2026 · arXiv.org · Vol abs/2607.23856 · 0 citations · 35 references
Computer Science

TL;DR

Viewing OC as a last-layer ensemble also organizes detectors into a two-axis taxonomy and exposes the OC score as a magnitude, motivating a scale-invariant, label-free direction score that repairs its near-OOD failure.

Abstract

A Last-Layer Ensemble (LLE), $K$ linear units on one shared frozen feature map, is an efficient single-pass approach to the disagreement-based epistemic uncertainty for out-of-distribution (OOD) detection. Its weakness is that members share the backbone gradient and can converge toward the same function, collapsing the inter-member diversity the signal depends on. Whether last-layer diversity can be restored, and what mitigates the collapse, is an open question. The weight-orthonormality defining Orthonormal Certificates (OC), the weight-orthonormal special case of the LLE, is only an indirect correction; it decorrelates the weights of the members, not their predictions. Here, we instead target the collapse directly in function space, with a Covariance Last-Layer Ensemble (cov-LLE) that places a direct covariance penalty on member activations. Cov-LLE restores the function-space diversity that weight-orthonormality cannot, and at matched $K$ recovers much of the diversity and calibration of a deep ensemble at $1\times$ backbone cost (in-distribution prediction variance $0.05\!\to\!9.3$ vs. $22.1$ ($\times10^{-3}$), and ECE $0.135\!\to\!0.090$ vs. $0.035$, for a $K\times$-cost deep ensemble), at no cost to accuracy. Viewing OC as a last-layer ensemble also organizes detectors into a two-axis taxonomy (by how their units are trained and how their outputs are scored) and exposes the OC score as a magnitude, motivating a scale-invariant, label-free direction score that repairs its near-OOD failure, adding $+0.16$ to $+0.18$ ROC AUC on every backbone.

View source

Similar papers

Jul 2026

Controllable Diversity in Normalization-Based Implicit Ensembles via Softmax-Temperature Modulation

A normalisation-based implicit ensemble that treats each member as a task in a multi-task architecture and modulates the shared backbone through sigmoid-bounded scalers is introduced, which matches or outperforms deep ensembles at a fraction of their parameter cost, scales with ensemble size where partitioning methods collapse, and maintains calibration under distribution shift.

Mihai Suteu, Ovidiu Serban · 0 citations
Open access Aug 2026

DESS: A Robust Uncertainty Layer for Embedding-Space Models

DESS is introduced, a lightweight uncertainty layer that augments an existing embedding model with a predicted mean vector and an independent per-dimension spread vector that provides a modular, geometry-aware uncertainty layer for embedding-space models, provided its spread is calibrated to local embedding geometry.

Morten Grundetjern, J. Voigt, Per-Arne Andersen et al. · 0 citations
#machine learning Preprint Sep 2026

Frozen Cores Need Task Signal: Fisher-Whitened Cross-Covariance for Low-Resource LLM Adaptation

Parameter-efficient fine-tuning is usually framed as a question of how many parameters to update. Under a severe trainable-state budget, however, where those coefficients act is equally consequential. We study this choice through frozen-core adaptation: a calibration pass fixes left and right bases for each weight matrix, and fine-tuning optimizes only an $r\times r$ core. This removes the ability of trainable factors to repair a poor initial span and makes subspace quality directly observable. We introduce FCCA, which estimates the signed input--error cross-covariance, whitens it with diagonal Fisher moments, truncates it in the resulting local metric, maps the selected directions back, and applies thin QR to obtain stable core coordinates. Under a matched $r^2$ budget, we compare eight basis constructors on 11 tasks, four model settings, and three seeds. On Qwen2.5-3B, FCCA reaches an 83.0 macro-average, 2.3 points above the next-best matched-budget constructor, and exceeds its unwhitened RawGrad control on all 11 tasks. It ranks first at all three Qwen scales and finishes within 0.13 points of the best method on Llama-3.2-1B. Controlled ablations show gains of 2.7--17.2 points from whitening and identify QR as necessary for stable core optimization in the tested regime. Finally, FCCA comes within 0.32 and 0.23 average points of LoRA and DoRA while optimizing 36.9K rather than roughly 7.4M parameters. These results show that a carefully selected fixed span can recover most of the benefit of movable low-rank factors at a much smaller trainable and optimizer-state cost.

Wen-song Ye, Zhan-Ming Shen, Zhiqing Xiao et al. · 0 citations
#machine learning Preprint Aug 2026

AdaptNTK: Adaptive Uncertainty Quantification and Active Learning for Neural Network Potentials

Machine learning interatomic potentials bridge the gap between quantum chemical precision and classical computational speed, enabling molecular dynamics simulations with first-principles accuracy. Their reliability is often improved through active learning, which iteratively expands the training set by identifying uncertain, out-of-distribution configurations. Existing uncertainty-quantification methods often involve a trade-off between computational cost and reliability, and generally cannot account for redundancy as an acquisition batch is assembled. Here, we introduce AdaptNTK, a single-model framework that measures uncertainty as a regularized Mahalanobis distance in empirical neural tangent kernel (NTK) feature space. With the NTK features fixed during acquisition, the uncertainty depends on the acquired configurations but not their reference labels. This allows the uncertainty to be updated recursively after each selection without retraining, reducing redundancy within an acquisition batch. On held-out rMD17 data, AdaptNTK achieves the highest mean correlations with force errors (Spearman 0.68, Pearson 0.71) and matches a three-member ensemble in error retention. In active learning experiments, AdaptNTK achieves the lowest force errors across rMD17 and Transition-1X, with particularly strong performance on transition-state configurations in Transition-1X. AdaptNTK provides a 2.6-fold speedup per Transition-1X cycle relative to the ensemble, providing efficient single-model uncertainty estimation with sequential updates for data-efficient active learning.

Prajwal Ananth, Shuwen Yue · 0 citations
Book Open access Aug 2026

Quantized Model Soup Shake-Up: Weight Perturbation for Enhanced Ensemble Diversity

Model soup, averaging the weights of multiple fine-tuned models, delivers ensemble-level accuracy at single-model inference cost, but its success requires both linear mode connectivity (LMC) and sufficient diversity among candidates. We study these two requirements under quantization-aware training (QAT). First, we show analytically and empirically that fine-tuning from a sufficiently converged QAT checkpoint, which we call the QAT anchor state, preserves linear mode connectivity by keeping models within the same loss basin. Second, we identify Ensemble Degeneracy, where the many-to-one mapping of the quantizer collapses independently fine-tuned models into nearly identical quantized representations, eliminating diversity despite intact connectivity. To resolve this, we propose Quantized Model Soup Shake-Up (QMSS), which selectively perturbs low-magnitude weights near quantization bin boundaries to flip their integer assignments, then briefly re-trains each variant via QAT. Because only a small fraction of inherently unstable weights are modified, QMSS induces substantial quantized-domain diversity while preserving LMC. Experiments on CIFAR-100, Tiny-ImageNet, ImageNet, and FSD-Kaggle2018 with both EWGS and LSQ quantizers on CNNs and Vision Transformers show consistent improvements over standard model soup, with gains up to +1.54% Top-1 on CIFAR-100 and +2.09% on CIFAR-100-C, alongside over 6.37× inference speedup on a Jetson Orin Nano at 8/8-bit.

Jinwook Chung, Sungyeop Jung, Weronika Czorapinska et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.