TCNeRV is proposed, which exploits reconstructed context in both feature and embedding domains and reduces BD-rate by 22.06%, 66.73%, and 29.85% relative to HM, DCVC, and HiNeRV, respectively, demonstrating competitive rate-distortion performance with limited model capacity.
Abstract
Video compression aims to minimize reconstruction distor tion under a constrained bit rate. Existing video implicit neural representations (INRs) often decode frames independently, leaving intermediate features unconditioned on previous reconstructions and content embeddings without explicit temporal prediction. We propose TCNeRV, which exploits reconstructed context in both feature and embedding domains. Its multi-scale temporal-context fusion (MTCF) module injects gated historical features at multiple decoder scales, while temporal embedding-residual coding (TERC) predicts each content embedding and codes only its residual. With approximately 3M parameters, TCNeRV achieves an average PSNR of 36.08 dB on the UVG dataset, outperforming HNeRV-Boost by 2.20 dB. It reduces BD-rate by 22.06%, 66.73%, and 29.85% relative to HM, DCVC, and HiNeRV, respectively, demonstrating competitive rate-distortion performance with limited model capacity.
A Video Compression framework built upon a foundational flow model that enables the compressor to harness generative video flow priors effectively, which reduces bit consumption by 58\% and achieves one-step decoding and reconstructions with high perceptual fidelity.
Yichong Xia, Qin-Hong Wu, Bin Chen et al.· 1 citation· ⚡1
DiffVC-ONE, a diffusion-based generative video compression framework built on a one-step Video Diffusion Transformer, is proposed and a Unified Unidirectional Latent Compressor that uses a shared model to efficiently and uniformly compress compact latent slices is introduced.
Neural video representations (NVRs) represent videos using neural network parameters and, in hybrid formulations, frame-wise latent embeddings. Although hybrid NVRs can improve reconstruction quality by using content-adaptive latent embeddings, their latent spatial sizes and decoder upsampling schedules are tied to the...
A context-aware framework that specializes a pretrained semantic decoder to each group of pictures by optimizing only its channel-wise scale factors and transposed-convolution biases, enabling the receiver to reproduce the adapted decoder without full-model retraining is introduced.
P. Samarathunga, Yasith Ganearachchi, Thanuj Fernando et al.· IEEE Access· 0 citations
Implicit neural representations (INRs) have emerged as a promising paradigm for video compression, providing compact neural representations with flexible spatial and temporal reconstruction. Hierarchical grid-based architectures such as HiNeRV achieve strong rate--distortion performance, but require extensive per-video...
Naser Alizada, Farhang Baghban, Hashem Pishkar et al.· 0 citations
Implicit neural representations (INRs) encode video frames as network weights, offering a new compression paradigm. A persistent limitation is that existing INR codecs train one model per target bitrate, so multi-rate deployment needs separate runs, separate checkpoints, and model reloading, costs that grow with the nu...
Qing-Yu Mao, Jia-Cong Chen, Shuai Liu et al.· Electronics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.