Skip to content
Open access

Context-Aware Semantic Video Coding via Content-Adaptive Parameter Overfitting

2026 · IEEE Access · Vol 14, pp. 144815-144839 · 0 citations · 77 references

TL;DR

A context-aware framework that specializes a pretrained semantic decoder to each group of pictures by optimizing only its channel-wise scale factors and transposed-convolution biases, enabling the receiver to reproduce the adapted decoder without full-model retraining is introduced.

Abstract

Hybrid semantic video coding combines learned representations with standardized residual coding, but existing approaches either retrain an entire model for each group of pictures or use a fixed decoder that cannot adapt to local content. This paper introduces a context-aware framework that specializes a pretrained semantic decoder to each group of pictures by optimizing only its channel-wise scale factors and transposed-convolution biases. The compact update is differentially compressed and transmitted with the semantic representation, enabling the receiver to reproduce the adapted decoder without full-model retraining. A hierarchical bidirectional prediction structure supplies motion-aligned temporal context, while a standardized residual pathway preserves reconstruction fidelity. Unlike prior hybrid semantic codecs, the proposed method jointly provides lightweight content adaptation, explicit accounting of parameter-transmission cost, and compatibility with a shared pretrained backbone. Experimental evaluation across natural and out-of-distribution video demonstrates consistent coding gains over the fixed-decoder baseline and the standardized reference codec under diverse conditions.

Read PDF

Similar papers

Preprint Sep 2026

Resolution-Flexible Decoding for Hybrid Neural Video Representations

Neural video representations (NVRs) represent videos using neural network parameters and, in hybrid formulations, frame-wise latent embeddings. Although hybrid NVRs can improve reconstruction quality by using content-adaptive latent embeddings, their latent spatial sizes and decoder upsampling schedules are tied to the...

Taiga Hayami, Masaya Takabe, Hiroshi Watanabe · 0 citations
Sep 2026

DeepJSCC for video semantic communication with general semantic preservation.

Video-based intelligent applications such as autonomous driving, video understanding, and telemedicine rely heavily on the availability of robust and transferable semantic representations. In practical systems, however, video signals transmitted over wireless links are often exposed to severe and unpredictable distorti...

Jun-Ting Li, Xuechen Chen, Xiao-Heng Deng · 0 citations
Preprint Sep 2026

Video Compression with Graph-inspired Neural Representation

The results show that the G-NeRV codec outperforms the state-of-the-art INR-based codec, NVRC, and the latest standard video codec, VVC VTM, by 8.86\% and 14.68\% (in BD-rate), respectively, measured by PSNR on the UVG dataset.

Chang-Qiu Wang, Ge Gao, Fan Zhang et al. · 0 citations
Preprint Aug 2026

Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation

KATok (Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation), a transformer-based VAE that incorporates an adaptive token selector which is jointly learned with latent tokens that achieves strong reconstruction and generation quality at a state-of-the-art compression ratio.

Yeonkyeong Lee, Hyun-Young Go, Jongmin Kim et al. · 0 citations
Preprint Sep 2026

tcnerv:dual-domain temporal context modeling for implicit neural video compression

TCNeRV is proposed, which exploits reconstructed context in both feature and embedding domains and reduces BD-rate by 22.06%, 66.73%, and 29.85% relative to HM, DCVC, and HiNeRV, respectively, demonstrating competitive rate-distortion performance with limited model capacity.

Xue-Lian Xiang, Yi-Xin Zhao, He-Qi Xiang et al. · 0 citations
Sep 2026

AdaCompVL: Adaptive Compression of Spatiotemporal and Cross-Modal Redundancy for Efficient Video–Language Learning

Multimodal large language models (MLLMs) have recently extended from static image understanding to video comprehension, but representing videos as frame-level token sequences incurs substantial computational overhead. Existing visual token compression methods typically rely on uniform sampling or single-dimension redun...

Jian-Xin Ma, Shi-Bo Jin, Lu-Juan Dang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.