Skip to content

TOM-GS: Editable Video Representation via Temporal Opacity Modulation of Static 3D Gaussians

Jul 2026 · arXiv.org · Vol abs/2607.22717 · 0 citations · 25 references
Computer Science

TL;DR

TOM-GS is introduced, an editable video representation that forgoes complex deformations in favor of regular 3D Gaussians equipped with a continuous temporal opacity formulation, which enables static 3D spatial components to fade smoothly in and out of the scene.

Abstract

While Implicit Neural Representations (INRs) and dynamic 3D Gaussian Splatting (3DGS) achieve impressive results in video processing, they often fall short of producing representations that are easily editable. Recent methods address this by introducing complex spatial deformations or folded distributions, which constrain optimization and reduce flexibility for downstream editing. In this paper, we introduce TOM-GS, an editable video representation that forgoes complex deformations in favor of regular 3D Gaussians equipped with a continuous temporal opacity formulation. By assigning a learnable temporal mean and scale to the opacity of each Gaussian, our model enables static 3D spatial components to fade smoothly in and out of the scene. Grounded by robust, off-the-shelf pose estimation, our approach maintains a static spatial geometry that naturally supports a wide range of manual and physics-based edits. TOM-GS outperforms prior editable video representations in visual fidelity, while its reliance on standard 3D Gaussians ensures seamless compatibility with established 3D editing tools.

View source

Similar papers

Jul 2026

Consistent 4D Appearance Editing with Gaussian Splatting.

Editing dynamic scenes with 4D Gaussian Splatting (4DGS) is often hampered by spatiotemporal inconsistencies, or "Gaussian drifting", which degrades edit quality and temporal coherence. We identify that these artifacts stem from two distinct sources: foundational inaccuracies in the initial scene reconstruction, and the disruption of learned trajectories during the editing process itself. To address this, we propose a comprehensive framework that systematically tackles both sources of inconsistency. To solve reconstruction-induced errors, we introduce a novel prior-guided, multi-stage reconstruction pipeline that fuses geometric and motion priors to build a physically plausible and temporally stable foundation. To solve editing-induced errors, we further apply a universal trajectory-preserving technique, which safeguards high-quality motion by decoupling the appearance optimization from the learned deformation. Experiments demonstrate that by systematically addressing both the reconstruction and editing phases, our method achieves state-of-the-art, temporally consistent editing on a wide range of dynamic scenes where previous monolithic approaches fail.

Xiaosheng He, Feng-Lin Liu, Lin Gao et al. · 0 citations
Preprint Aug 2026

FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors

This work presents FixAnything, a single model for fixing a wide range of rendering artifacts by repurposing a pretrained video generative model, leveraging its implicit multi-view priors with only minimal modification and lightweight finetuning.

Khiem Vuong, D. Ramanan, Srinivasa Narasimhan · 0 citations
#diffusion models Preprint Sep 2026

Rethinking 3D Noise: Learning 3D-Aware Video Priors via Optimization-Free Morphological Perturbations

3D scene representations like NeRF and 3D Gaussian Splatting (3DGS) suffer severe artifacts in sparse-view settings. Recent generative 3D artifact fixers attempt to address this, but rely on paired corrupted and clean renders requiring costly, per-scene reconstructions across varying view configurations. While 2D image augmentations act as instant regularizers, no explicit equivalents exist for 3D representations to preserve spatial consistency across views, an essential property for 3D-aware training. We propose 3D Morphological Perturbations as an optimization-free regularizer that preserves spatial consistency. Leveraging explicit 3DGS, we treat each Gaussian as a fundamental building block - analogous to a 2D pixel - and apply perturbations across its morphological parameter space via scale, rotation, and pruning. Our method eliminates per-scene 3DGS optimization loops from dataset curation while enabling models to learn stronger geometric priors than sparse-view baselines in diagnostic ablations conducted on a lightweight video diffusion sandbox. Scaled to a 14B-parameter video model via ControlNet, our approach maintains visual fidelity while reducing mean depth error by 12.5% over state-of-the-art image-to-image 3D artifact refiners, ultimately boosting downstream robotics policy success rates by up to 8.0% across 3 of 4 manipulation tasks.

Onat Şahin, Mohammad Altillawi, George Eskandar et al. · 0 citations
Preprint Aug 2026

GaussVid: Sparse-View Gaussian Splatting with 3D-Aware Video Diffusion Priors

This work proposes a novel 3D-aware video restoration framework designed to enhance the quality of sparse 3DGS reconstruction and introduces a camera-conditioned geometric prior that guides the network toward geometrically grounded restoration that remains coherent across viewpoints.

Xinhui Liu, Can Wang, Wei Jiang et al. · 0 citations
Preprint Aug 2026

DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion

DiGS-Avatar is proposed, which reformulates this task as an efficient, diffusion-based UV-latent completion task, ensuring 3D consistency by design, and introduces a teacher-student framework where a multi-view teacher provides geometrically aligned pseudo-ground-truth latents to supervise a single-view diffusion student.

Jiakun Li, Li Fang, Hao Zhu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.