Skip to content
Preprint

GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation

Sep 2026 · 0 citations · 64 references
Computer Science

TL;DR

Results show that the latent space is central to geometry-consistent generation and can serve as a shared interface between perception and generation.

Abstract

We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically rich space that encodes cross-view structure. Rather than adding geometry as another output, we reparameterize a geometry foundation model's features into a compact latent space for generation. We realize this shift with the geometry-native autoencoder (GAE), whose latent is jointly decodable to appearance, depth, cameras, and point maps. With this state, a standard conditional flow supports diverse generation tasks. In controlled comparisons that hold the generator and training protocol fixed, replacing the latent with GAE improves both visual quality and independently measured 3D coherence: FVD falls by $12.7\%$ and $23.1\%$ on RealEstate10K and DL3DV, and camera-trajectory error is halved on RealEstate10K. Together, these results show that the latent space is central to geometry-consistent generation and can serve as a shared interface between perception and generation.

View source

Similar papers

Preprint Aug 2026

WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild

A generative appearance model where a $\beta$-VAE learns a structured and continuous manifold of global appearance is introduced, Conditioned on the latent code, a 3D neural appearance field is constructed that generates dynamic Tri-Plane features to encode spatially-varying local illumination effects.

Yu Bai, Qian-Qiu Tan, Li-Long Chen et al. · 0 citations
Preprint Sep 2026

GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space

Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unseen regions, whereas...

Ke-Rui Ren, Tao Lu, Lin-Ning Xu et al. · 0 citations
Preprint Aug 2026

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

SpatialCrafter is presented, a novel two-stage framework that addresses explorable image-to-scene generation issues by introducing a global 3D proxy for high-fidelity image-to-scene generation and appearance refinement and introduces Parallel Geometry Injection and Proxy-Aware Corruption training strategies.

Chuan Fang, Lingteng Qiu, Yixun Liang et al. · 1 citation
Preprint Aug 2026

Luce: Relightable Gaussians for 3D Asset Generation

High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. However, preserving fine detail across the physically based rendering (PBR) modalities needed for relighting remains challenging. To address this, we propose Luce, a 3D representation that unifies geometry and...

M. Singh, Michele Stoppa, Alvise Memo et al. · 0 citations
Preprint Aug 2026

GaussianWAM: Distilling Geometry and Semantics from 3D Gaussian Fields into World-Action Models

GaussianWAM is proposed, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field and improves performance on standard LIBERO and shows positive transfer trends on RoboTwin and real-world manipulation.

Zi-Jian Zhang, Yu-Qing Jiang, Wei-Tao Zhou et al. · 1 citation
Preprint Sep 2026

LatentReRig: An SDF-Based VAE with Dual Decoders for Latent-Space Deformation Conditioning

Transferring deformation between characters with different geometry and topology is challenging because conventional rigs encode behaviour through character-specific structures and correspondences. We present LatentReRig, an experimental framework that investigates whether pose-associated changes can instead be represe...

D. Dolci, Fabrizio Poggioni, Carlo Melchiorri · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.