Skip to content

4DWeaver: Bridging Reconstruction and Generation Via Compact Autoregressive Priors.

Aug 2026 · IEEE Transactions on Pattern Analysis and Machine Intelligence · Vol PP, pp. 1-16 · 0 citations
Medicine

TL;DR

This work proposes Compact Autoregressive Latent Prior (CALP), which regularizes low-dimensional latent variables with a history-conditioned autoregressive prior, achieving compact, long-range temporally coherent, and spatially structured latent representations through a more reasonable latent capacity allocation and a latent-space organization that is better suited for diffusion-based 4D generation.

Abstract

Large-scale 4D scene generation aims to synthesize dynamic 3D environments and provides a critical intermediate representation for downstream tasks such as autonomous driving simulation, embodied agent training, and scene forecasting. Existing methods typically adopt a two-stage latent diffusion paradigm, which improves computational efficiency by modeling and generating scenes in a compressed latent space. However, this paradigm suffers from a reconstruction-generation trade-off: increasing the latent dimensionality improves reconstruction fidelity, but substantially increases the computational burden of diffusion modeling and makes generative optimization more challenging. This issue becomes particularly pronounced in 4D occupancy generation, where complex spatial layouts and long-range temporal dynamics must be jointly preserved within compact representations. To alleviate this problem, we advance a central principle: low-dimensional latent spaces should not rely solely on unconstrained compression, but should instead be structurally regularized to preserve sufficient 4D spatio-temporal information while maintaining compactness. To instantiate this principle, we propose Compact Autoregressive Latent Prior (CALP), which regularizes low-dimensional latent variables with a history-conditioned autoregressive prior, achieving compact, long-range temporally coherent, and spatially structured latent representations through a more reasonable latent capacity allocation and a latent-space organization that is better suited for diffusion-based 4D generation. Built upon CALP, we further introduce 4DWeaver, a compact 4D scene generation framework that enables high-quality spatio-temporal occupancy synthesis in a low-dimensional latent space. Extensive experiments on multiple large-scale 4D occupancy benchmarks demonstrate that 4DWeaver achieves superior reconstruction and generation performance while substantially reducing memory consumption and computational cost.

View source

Similar papers

Preprint Aug 2026

GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors

GeoFlow is a novel framework designed to achieve efficient driving video generation by harnessing explicit geometric priors, using a Geometry-Aligned Prior (GAP) distribution as starting point, and can achieve remarkable efficiency of both training and inference.

Jiazheng Liu, Hangbiao Li, J. Zhang et al. · 0 citations
Preprint Aug 2026

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

SpatialCrafter is presented, a novel two-stage framework that addresses explorable image-to-scene generation issues by introducing a global 3D proxy for high-fidelity image-to-scene generation and appearance refinement and introduces Parallel Geometry Injection and Proxy-Aware Corruption training strategies.

Chuan Fang, Lingteng Qiu, Yixun Liang et al. · 0 citations
Jul 2026

MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction

The key idea is to progressively densify anchored VecSet latents via hierarchical point-shuffle upsampling, increasing spatial capacity for fine-grained geometry modeling and replacing global cross-attention with AVS-Conv, a geometry-aware local aggregation operator operating within local neighborhoods rather than the exhaustive latent set.

Dehao Hao, Kaiyi Zhang, Tanghui Jia et al. · 0 citations
Jul 2026

Structure-Detail Decoupled Autoregressive Generation for Fast and High-Fidelity Virtual Try-On

This work introduces VAR-VTON, a VAR-based VTON model that incorporates garment conditioning and structural guidance for efficient latent-space VTON, and proposes STAR-VTON, a Two-Stage AutoRegressive framework that builds upon VAR-VTON by decoupling latent-space structural synthesis from pixel-space detail recovery.

Lu Yang, Xiaonan Hu, Yanan Li et al. · 0 citations
Jul 2026

ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning

Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that depth can be recovered as a globally smooth field through iterative denoising, which does not explicitly reflect the piecewise and scale-dependent organization of scene geometry. In practice, geometric structure emerges progressively across spatial scales, where coarse layout, surfaces, and boundaries are constructed in a hierarchical manner. Motivated by this observation, we introduce ARDepth, which formulates depth estimation as structured auto-regressive generation. Instead of recovering depth through global refinement, ARDepth progressively constructs depth representations as spatial resolution increases. To support this generative process, we introduce Scale-Progressive Conditioning (SPC) to inject multi-scale visual features at each generation stage, and Semantic-Aware Guidance (SAG) to provide scene-level semantic priors that enhance global structural consistency. Together, these designs enable the model to capture fine-grained local details while maintaining coherent global geometry. Empirical results demonstrate that our approach achieves strong performance and produces structurally consistent depth predictions across scales, validating auto-regressive generation as a promising alternative paradigm for geometric modeling.

Zijie Wang, Wei Zhang, Wei-Ming Zhang et al. · 0 citations
Preprint Aug 2026

MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

MotionCraft is presented, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface to deliver temporally consistent, high-quality reconstructions under streaming constraints.

Rong Fu, Chunlei Meng, Yangcheng Zeng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.