Skip to content

Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling

Jul 2026 · arXiv.org · Vol abs/2607.19437 · 0 citations · 66 references
Computer Science Engineering

TL;DR

A unified generative framework that leverages pre-trained Diffusion Transformer priors to achieve high perceptual quality at extremely low bitrates, achieving state-of-the-art perceptual fidelity with rich spatial details and robust temporal consistency.

Abstract

Most existing video compression algorithms follow a paradigm of transformation and quantization, optimizing the trade-off between distortion and bitrate. However, extremely low-bitrate compression remains an underexplored frontier where perceptual quality optimization under severely constrained coding resources has not been adequately addressed. In this paper, we propose a unified generative framework that leverages pre-trained Diffusion Transformer (DiT) priors to achieve high perceptual quality at extremely low bitrates. We first introduce a flexible Group-of-Latents (GoL) strategy within the latent space of a causal tokenizer, explicitly partitioning the latent stream into intra $I$-latents and inter $P$-latents. The Deep Compression Module (I-DCM) then encodes key $I$-latents to preserve perceptual anchors with minimal overhead. Building upon these anchors, the DiT-based Unified Latent Denoising Module (U-LDM) refines intra-frame textures and synthesizes $P$-latents from noise, reconstructing temporal dynamics at zero additional bitrate cost. Extensive experiments demonstrate that our method uniquely operates in the extreme-low-bitrate regime (e.g., (<0.005) bpp), achieving state-of-the-art perceptual fidelity with rich spatial details and robust temporal consistency. The code will be made publicly available.

View source

Similar papers

Preprint Aug 2026

GVC-RT: Towards Real-Time Generative Video Compression at Ultra-Low Bitrates

This work systematically identifies the computational bottlenecks and proposes GVC-RT, which redesigns the generative latent coding framework to realize real-time video coding without sacrificing compression performance, and introduces a lightweight de-tokenizer architecture to resolve the final latency bottleneck during decoding.

Tianjian Dang, Sixian Wang, Lei Luo et al. · 0 citations
Preprint Aug 2026

DiffVC-ONE: Diffusion-based Generative Video Compression with One-Step Video Diffusion Transformer

DiffVC-ONE, a diffusion-based generative video compression framework built on a one-step Video Diffusion Transformer, is proposed and a Unified Unidirectional Latent Compressor that uses a shared model to efficiently and uniformly compress compact latent slices is introduced.

Wenzhuo Ma, Zhenzhong Chen · 0 citations
Preprint Aug 2026

Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation

KATok (Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation), a transformer-based VAE that incorporates an adaptive token selector which is jointly learned with latent tokens that achieves strong reconstruction and generation quality at a state-of-the-art compression ratio.

Yeonkyeong Lee, Hyun-Young Go, Jongmin Kim et al. · 0 citations
Preprint Aug 2026

FLM: Frequency-Aware Language Models for Generative Image Compression

Generative models have significantly improved the performance ceiling of image lossy compression at low bitrates by exploiting learned priors. However, the generated textures and semantic details may deviate from the source content, thereby affecting the fidelity of image reconstruction. To solve these challenges, we propose FLM, a frequency-aware language model that improves compression efficiency through frequency-domain probabilistic modeling while retaining deterministic reconstruction. At the encoder, the input image is transformed into quantized DCT coefficients, which are organized into discrete sequences using macroblock-based coefficient tokenization. FLM then performs next-coefficient prediction to autoregressively estimate token-wise conditional probability distributions for arithmetic coding, thereby generating a compact bitstream. At the decoder, the LLM and arithmetic decoder jointly recover the frequency-domain data, followed by inverse transformations for image reconstruction. A task-specific frequency-domain dataset and a two-stage fine-tuning strategy are further developed to enable the model to operate across multiple bitrate settings. FLM is a versatile compressor that is compatible with both lossy compression and lossless JPEG recompression frameworks. Experiments show that FLM exceeds conventional and generative lossy compression methods in rate-distortion performance. FLM achieves BD-PSNR gains of 3.30 dB, 3.83 dB, and 3.80 dB than JPEG baseline on Kodak, Tecnick, and CLIC2020, respectively. Better qualitative quality of FLM can be achieved in improving semantically high fidelity and suppressing blocking artifacts. FLM is also validated to be applicable to the lossless recompression task with competitive performance.

Jia-Run Chen, Kejun Wu, Li Li et al. · 0 citations
Preprint Aug 2026

V-RAE: Rethinking Video Latent Spaces for Generation

V-RAE, a video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations, and tFVD, a temporal-coherence diagnostic that correlates more reliably with downstream generation quality are introduced.

Minghui Guo, Shengqiong Wu, Hao Fei · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.