Skip to content

dRAE: Representation Autoencoder with Hyper-Spherical Codes

Jul 2026 · arXiv.org · Vol abs/2607.22148 · 0 citations · 70 references
Computer Science

TL;DR

This work proposes Hyper-Spherical Quantization (HSQ), which decouples semantic content from feature magnitude via angular routing, preventing code assignment from being dominated by scale rather than meaning.

Abstract

In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric mismatch: standard Euclidean codebook objectives are fundamentally misaligned with the anisotropic geometry of representation space, leading to codebook embeddings with high-variance magnitude scales and uneven angular distributions that hinder scalability. To address this, we propose Hyper-Spherical Quantization (HSQ), which decouples semantic content from feature magnitude via angular routing, preventing code assignment from being dominated by scale rather than meaning. The resulting discrete Representation Autoencoder (dRAE) achieves high-fidelity reconstruction while preserving semantic integrity and supporting scalable codebook budget. Extensive experiments demonstrate consistent performance gains as the vocabulary size scales to 131{,}072, along with 100\% codebook utilization, simplified training pipeline, and strong performance across understanding and generation tasks.

View source

Similar papers

Preprint Aug 2026

Hadamard-Domain Model Quantization for Learned Image Coding

Hadamard-Transform-domain Quantization (HaTQ), which uses orthogonal Hadamard reparameterization before quantization to redistribute weight and activation responses in the original domain across channels, and is compatible with integer-only execution.

Jun-Qi Shi, Chongzhi Wang, Yiwen He et al. · 0 citations
2025

Dimensional Collapse in VQVAEs: Evidence and Remedies

This work identifies a surprising yet consistent phenomenon that it is identified: despite using high-dimensional embeddings, VQVAEs tend to compress their representations into a much smaller subspace, typically only 4 to 10 dimensions, and proposes Divide-and-Conquer VQ, which partitions the latent space into multiple low-dimensional subspaces, each quantized independently.

Jiayou Zhang, Yifan Shen, Guan-Hong Chen et al. · 4 citations
Preprint Jul 2026

KronQ: LLM Quantization via Kronecker-Factored Hessian

KronQ, a PTQ framework that challenges the assumption that all output channels contribute equally to the layer-wise reconstruction objective by introducing the gradient covariance into the quantization pipeline, and introduces bidirectional incoherence processing.

Donghyun Lee, Yuhang Li, Ruokai Yin et al. · 0 citations
Preprint Aug 2026

NAE: Normalizing AutoEncoder

This work proposes Normalizing Autoencoder (NAE), which employs a novel conditional loss that aligns the surrogate loss gradient with that of reconstruction loss, directly improving upon the current standard.

Muhammad Abdur Rafae, Niels Landwehr · 0 citations
Jul 2026

The Scalable Tensor-based Codebook Product Quantization for Multi-Label Image Retrieval.

Scalable product quantization has recently attracted considerable attention in large-scale image retrieval, as it avoids the need to train multiple models for generating quantization codes of varying lengths. However, most existing approaches primarily concentrate on approximating the ground-truth similarity between image pairs, while overlooking the correlations among sub-codebooks and among codewords. In addition, limited work has addressed the challenge that increasing the number of subspaces or codewords substantially raises memory consumption. To address these limitations, we propose a novel scalable product quantization framework within an end-to-end network, termed Tensor-based Codebook Product Quantization (TCPQ). This work represents an innovative attempt to integrate tensor theory with product quantization methods. The framework leverages tensor-based methods to capture spatial correlations among sub-codebooks and among codewords, and adopts lightweight codebooks for efficiency. For optimization, a subspace-wise unbiased supervised contrastive loss is proposed to bring embeddings of the same class closer together, push embeddings of different classes farther apart within quantization subspaces, and precisely regulate the minimum distance between positive and negative samples. In addition, the ArcFace loss is incorporated to enhance the discriminative power of the learned features, and an orthogonality constraint is imposed on the factor matrix along the codebook dimension to avoid overfitting. Extensive experiments on three large-scale real-world benchmarks demonstrate that TCPQ achieves state-of-the-art retrieval performance.

Bin Luo, Laurence T. Yang, Debin Liu et al. · 0 citations
Review Aug 2026

Transforms for LLM Quantization: The Great Inversion and Format Co-Design

This work identifies and formalizes the principle that organizes the Great Inversion, the Great Inversion: allocation-flexible coding rewards energy concentration, whereas the grouped shared-scale quantization a deployed matrix instruction performs rewards within-group flattening.

E. Jokar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.