Skip to content
Book Open access

SmoothMotionVectors: Optimizing Your Content for Video Codecs in Free View Video Compression

Jul 2026 · International Conference on Computer Graphics and Interactive Techniques · 0 citations · 54 references
Computer Science

TL;DR

An empirical study of free-view video compression for dynamic scenes reconstructed with 3D Gaussian Splatting for dynamic scenes reconstructed with 3D Gaussian Splatting, examining how practical pipeline design choices affect reconstruction fidelity and storage efficiency.

Abstract

We present an empirical study of free-view video compression for dynamic scenes reconstructed with 3D Gaussian Splatting, examining how practical pipeline design choices affect reconstruction fidelity and storage efficiency. Rather than introducing new representations, we analyze how commonly used components, including temporal chunking, deformation-based reconstruction, and quantization-aware training, interact in practice. We observe that partitioning long sequences into shorter temporal segments, such as GOPs, simplifies optimization and improves reconstruction fidelity, but can introduce additional storage overhead. We further show that encouraging smooth motion vectors across both space and time produces deformation signals that are easier for standard video codecs to compress, leading to improved rate–distortion performance. When integrated into a unified pipeline, these design choices consistently benefit different deformation-based reconstruction methods. Across multiple datasets, our approach achieves 20% storage reduction compared with state-of-the-art methods while preserving or improving visual quality, and we discuss sources of variability and ambiguity in current training and evaluation protocols.

Read PDF

Similar papers

Preprint Aug 2026

QuARC-GS: Quantized Anchored Residual Coding for Compact Dynamic Scene Streaming with Gaussian Splatting

Quantized Anchored Residual Coding Gaussian Streaming (QuARC-GS), a quantization-aware 4D scene optimization framework for online dynamic scene reconstruction that achieves ultra-high compression while maintaining reconstruction speed and quality, is proposed.

V. Nguyen, Yuchen Wang, Kyung Chul Lee et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VoRTeC: Taming Foundation Flow for One-step Real time Video Compression

Ultra-low bitrate video compression still faces critical challenges: traditional neural video compression inevitably introduces blurring artifacts, while diffusion-based generative video compression suffers from excessive decoding latency and poor temporal consistency. To address these issues, we propose $\mathtt{VoRTeC}$, a Video Compression framework built upon a foundational flow model (Wan2.1). By compactly encoding latent video representations, predicting the positions of compressed representations along flow trajectories, and integrating multi-scale priors, $\mathtt{VoRTeC}$ enables the compressor to harness generative video flow priors effectively. Without accessing the parameters or gradients of flow matching networks, our framework achieves one-step decoding and reconstructions with high perceptual fidelity. Meanwhile, we maintain consistency across frame groups via tail-frame reuse and prior caching. Extensive experiments demonstrate that our method reduces bit consumption by 58\% compared to prior diffusion-based approaches, with decoding speed boosted by 3 to 197 times: $\mathtt{VoRTeC}$ achieves a decoding speed of 13 FPS at 720p and 32 FPS at 480p.

Yichong Xia, Qin-Hong Wu, Jin-Peng Wang et al. · 0 citations
Preprint Aug 2026

Struct-GStream: Towards Efficient Free-Viewpoint Video Streaming at Low-Bitrates with Structured 3D Gaussians

Struct-GStream is proposed, which can achieve efficient FVV streaming using structured 3D Gaussians (3DGs) and introduces dynamic anchor points to generate structured 3DGs to construct basic scenes and model approximate scene movements based on the assumption of local rigidity in object motion.

Han Jiao, Jiakai Sun, Lei Zhao et al. · 0 citations
2026

Person-Prioritized Restoration for High-Compression 360° Video

High-Compression videos suffer from severe distortions, among which degradation in person regions has the greatest impact on viewers’ immersive experience. Existing quality enhancement techniques usually focus on overall image denoising or super-resolution, often overlooking the crucial recovery of fine structures in these essential person regions. To address these challenges, the research introduces a novel framework titled Person Region Restoration Driven by Perceptual Fidelity (PRRDPF), which combines long-range dependency features with perceptual structure loss for enhanced generative restoration. Specifically, first, the research constructs a high-fidelity distorted person-region dataset via a closed-loop degradation pipeline, addressing the lack of paired datasets. Secondly, a Temporal Gated Fusion (TGF) block is designed to use gated convolutions for selectively recovering high-frequency features while capturing local and global dependencies. Finally, a Structural Similarity Index Measure (SSIM)-based dynamic weighted adversarial loss is proposed to prioritize the restoration of visual texture details. Experimental results validate that PRRDPF significantly outperforms the best models in Peak Signal-to-Noise Ratio (PSNR), SSIM, and Learned Perceptual Image Patch Similarity (LPIPS), effectively mitigating artifacts and enhancing clarity in person visuals. This framework presents a promising approach for intelligent video coding integrated with generative artificial intelligence and holds significant potential for practical applications.

Linyun Liu, Li Yu, Jiaxin Zeng et al. · 0 citations
Review Open access Aug 2026

A survey of implicit neural representations for video compression

Neural Video Representations (NVRs) have recently been proposed as a novel approach to the video compression problem. NVRs consist of one or multiple small neural networks that are overfitted on one specific video sequence, thereby encoding the video within the weights and biases of the network(s). In contrast to other learned video coding approaches, NVR-based codecs do not rely on large datasets and can achieve lower decoding complexity by using compact, video-specific models instead of large shared encoder-decoder architectures. Many works have focused on improving the compression performance of NVR-based codecs by enhancing the overall codec design, devising more performant and parameter-efficient model architectures, and incorporating more advanced model compression schemes such as weight pruning, quantization, and entropy minimization. We provide a systematic overview of representative work in the field and discuss common weaknesses and opportunities for future work, with a focus on practical deployment for video streaming. This paper serves as both an introduction for newcomers and a reference for existing researchers, highlighting the potential of neural video representations as an alternative to traditional codecs in video compression.

Hannes Keunen, Maarten Wijnants, J. Liesenborgs · 1 citation
Jul 2026

PE-Field 4D: Video Generation Models as Canvas

This work revisits the role of positional encoding in video diffusion transformers and shows that it provides a useful spatial bias for geometry-aware control, and introduces a geometry-aware cross-attention mechanism that enables target video latent tokens to attend to structured context tokens derived from reference images or frames.

Yunpeng Bai, Haoxiang Li, Qixing Huang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.