Skip to content

AniGS: Bridging Rendering and Diffusion Prior for 3D Scene Animation

Jul 2026 · arXiv.org · Vol abs/2607.18539 · 0 citations · 42 references
Computer Science

TL;DR

AniGS is presented, a method for scene-level animation of 3D Gaussian Splatting (3DGS) reconstructions that adds subtle, distributed dynamics, e.g., vegetation motion, while preserving rigid structures in reconstructed environments.

Abstract

Novel view rendering of large and complex reconstructed scenes is becoming increasingly photorealistic. However, most reconstructions remain static and lack the ambient motion that makes environments immersive. We present AniGS, a method for scene-level animation of 3D Gaussian Splatting (3DGS) reconstructions that adds subtle, distributed dynamics, e.g., vegetation motion, while preserving rigid structures. Unlike existing 3D animation techniques which are limited to object-centric subjects or small regions, AniGS is designed for large, cluttered, navigable scenes. AniGS represents the scene with a canonical 3DGS and models motion using a time-conditioned deformation field. To animate the entire scene, we leverage a pretrained video diffusion model and introduce an iterative dataset--model update strategy that progressively expands viewpoint coverage and repeatedly updates camera-fixed training videos using a render-and-refine scheme. To prevent artifacts from unintended motion in static areas, we further introduce a composed video-to-video refinement scheme that restricts motion to desired regions. Experiments on five real-world, large-scale outdoor scenes demonstrate that AniGS produces natural ambient dynamics and high-quality novel view videos, enabling more immersive viewing experiences of reconstructed environments.

View source

Similar papers

Preprint Jul 2026

Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh

D, a reference-guided renderer that extends Wan2.2 camera control from Plucker rays alone to a joint camera-plus-geometry interface and projects a neural 4D G-buffer from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior, supporting tracking+world-position correspondence as a practical 4D rendering condition.

Junhao Chen, Mingjin Chen, He Zhang et al. · 0 citations
Preprint Aug 2026

FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors

This work presents FixAnything, a single model for fixing a wide range of rendering artifacts by repurposing a pretrained video generative model, leveraging its implicit multi-view priors with only minimal modification and lightweight finetuning.

Khiem Vuong, D. Ramanan, Srinivasa Narasimhan · 0 citations
Aug 2026

Dynamic View Synthesis from Monocular Videos via Motion-aware Gaussian Splatting.

This paper proposes a semantics-guided scene decoupling module that separates Gaussian primitives into static and dynamic components based on motion vectors, and introduces a motion-aware densification module for motion compensation, which alleviates the incomplete rendering of dynamic objects caused by insufficient spatio-temporal information.

Chulin Zhao, Xue Wang, Guoqing Zhou et al. · 0 citations
Jul 2026

Consistent 4D Appearance Editing with Gaussian Splatting.

Editing dynamic scenes with 4D Gaussian Splatting (4DGS) is often hampered by spatiotemporal inconsistencies, or "Gaussian drifting", which degrades edit quality and temporal coherence. We identify that these artifacts stem from two distinct sources: foundational inaccuracies in the initial scene reconstruction, and the disruption of learned trajectories during the editing process itself. To address this, we propose a comprehensive framework that systematically tackles both sources of inconsistency. To solve reconstruction-induced errors, we introduce a novel prior-guided, multi-stage reconstruction pipeline that fuses geometric and motion priors to build a physically plausible and temporally stable foundation. To solve editing-induced errors, we further apply a universal trajectory-preserving technique, which safeguards high-quality motion by decoupling the appearance optimization from the learned deformation. Experiments demonstrate that by systematically addressing both the reconstruction and editing phases, our method achieves state-of-the-art, temporally consistent editing on a wide range of dynamic scenes where previous monolithic approaches fail.

Xiaosheng He, Feng-Lin Liu, Lin Gao et al. · 0 citations
Jul 2026

Motion-driven 4D scene generation

This paper presents an innovative method that leverages user-specified action paths to guide the 4D scene generation that dynamically synchronizes motions in the action path domain with their corresponding contents in the time domain.

Guo-Wei Yang, Qun-Ce Xu, Zhao Wei et al. · 0 citations
Preprint Aug 2026

SPVC: Structured and Panoptic Video Fixing for Cross-Dataset Driving Scene Rendering

Driving scene reconstruction and rendering, especially with 3D Gaussian Splatting, has become an important component of autonomous driving simulation. However, rendered views often degrade under extrapolated ego trajectories and scene edits, producing blurry structures, temporal flicker, and foreground-background misalignment. Existing refinement methods are commonly designed for a specific setting, such as image-level novel-view repair or object-editing correction. In this paper, we introduce SPVC, a structured and panoptic video fixing framework for cross-dataset driving scene rendering. The name summarizes four design principles. (1) Structured fixing denotes the use of explicit spatial conditions, including camera pose, 3D bounding boxes, and HD maps, to guide the repair process and reduce uncontrolled hallucination. (2) Panoptic fixing refers to correcting both background rendering artifacts, such as distorted roads, buildings, and lanes, and foreground vehicle artifacts introduced by scene editing, such as inconsistent object appearance. (3) Video fixing means that the model operates on driving sequences rather than isolated frames, allowing temporal cues to be used during artifact correction. (4) Cross-dataset fixing means that a single shared network is trained and applied across multiple driving datasets, reducing the need for dataset-specific or scene-specific fixers. Concretely, we construct paired degraded-clean training data by simulating under-constrained 3DGS rendering and foreground vehicle insertion artifacts, and train a two-stage controllable video diffusion model that first addresses video-level appearance and then refines scene layout with structured controls.

Gen Li, Shu Han, Y. Qiao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.