This work proposes GVCCTurbo, a BPP-driven scheduler that separates expensive prior refreshes from codebook corrections, and supports BPP-to-compute scheduling as a controllable extension of sampler-length tuning, without requiring the allocated point to dominate every boundary point.
Abstract
Codebook-driven generative compression uses a pretrained image or video generator as a zero-shot visual prior and transmits compact codebook indices to guide reconstruction at ultra-low bitrate. Current codecs tie each finite-rate correction to a fresh prior evaluation, so shortening the sampler also removes correction slots that carry target-dependent information. We propose GVCCTurbo, a BPP-driven scheduler that separates expensive prior refreshes from codebook corrections: after calibrating an atom-count operating point and skip-gap ratio once per protocol, it maps a target codebook-payload bitrate to a trajectory length and refresh period, making BPP a schedule input instead of a fixed consequence of sampler length. The same endpoint-prediction and finite-rate steering interface covers GVCC-style rectified-flow video and DDCM-style diffusion image compression, preserving zero-training deployment and compatibility with future distilled priors. Native 1080p curves position the complete zero-shot codec in the ultra-low-bitrate regime. In a controlled 720p Wan-GVCC study, the scheduler cuts prior evaluations from 20 to 9 for a $\sim\!44\%$ measured decoding-time reduction shared across the whole schedule family, at a small shared LPIPS cost on high-motion content; within that family, uniform refresh thinning (pure-skip) is a boundary point, and the BPP-aware interior point trades $2.9\%$ fewer codebook-payload bits for consistently higher PSNR at comparable LPIPS. These results support BPP-to-compute scheduling as a controllable extension of sampler-length tuning, without requiring the allocated point to dominate every boundary point.
ReGenVC is the first end-to-end generative video codec to combine ultra-low-bitrate encoding with real-time decoding on an 8-GPU system and three model-preserving system techniques are presented.
This work systematically identifies the computational bottlenecks and proposes GVC-RT, which redesigns the generative latent coding framework to realize real-time video coding without sacrificing compression performance, and introduces a lightweight de-tokenizer architecture to resolve the final latency bottleneck during decoding.
Tianjian Dang, Sixian Wang, Lei Luo et al.· 0 citations
This work proposes MixCompress, a unified VBR framework based on sparse structural specialization that not only matches individually optimized single-rate baselines but can even surpass them, establishing a new Pareto frontier for computationally efficient image coding.
Calvin-Khang Ta, Praneet Singh, Tong Shao et al.· arXiv.org· 0 citations
Implicit neural representations (INRs) encode video frames as network weights, offering a new compression paradigm. A persistent limitation is that existing INR codecs train one model per target bitrate, so multi-rate deployment needs separate runs, separate checkpoints, and model reloading, costs that grow with the number of rate points. We propose a Generalized Slimmable Framework that replaces standard layers with width-configurable counterparts in most INR decoders, letting a single checkpoint serve multiple bitrates through nested weight tensors without topology changes. To recover the quality lost in shared-weight training, we introduce Slimmable Conditional Decoder Modulation (SCDM), which blends slimmable expert convolutions via width- and frame-conditioned gating. For encoder-based backbones, encoder output caching separates encoder computation from multi-width training. Across four INR backbones on DAVIS and Bunny, the framework cuts multi-rate storage by about 2.3–2.5×, and SCDM recovers substantial quality at all rate points while surpassing independently trained fixed-width models at narrow widths.
Qingyu Mao, Jiacong Chen, Shuai Liu et al.· Electronics· 0 citations
Ultra-low bitrate video compression still faces critical challenges: traditional neural video compression inevitably introduces blurring artifacts, while diffusion-based generative video compression suffers from excessive decoding latency and poor temporal consistency. To address these issues, we propose $\mathtt{VoRTeC}$, a Video Compression framework built upon a foundational flow model (Wan2.1). By compactly encoding latent video representations, predicting the positions of compressed representations along flow trajectories, and integrating multi-scale priors, $\mathtt{VoRTeC}$ enables the compressor to harness generative video flow priors effectively. Without accessing the parameters or gradients of flow matching networks, our framework achieves one-step decoding and reconstructions with high perceptual fidelity. Meanwhile, we maintain consistency across frame groups via tail-frame reuse and prior caching. Extensive experiments demonstrate that our method reduces bit consumption by 58\% compared to prior diffusion-based approaches, with decoding speed boosted by 3 to 197 times: $\mathtt{VoRTeC}$ achieves a decoding speed of 13 FPS at 720p and 32 FPS at 480p.
Yichong Xia, Qin-Hong Wu, Jin-Peng Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.