Skip to content
Preprint

What Makes a 3D Scene Editable? A Factorized Benchmark of Fidelity, Locality, Consistency, and Preservation

Sep 2026 · 0 citations · 69 references
Computer Science

Abstract

Neural 3D scene editing is often evaluated by semantic alignment alone, although a convincing result may alter unrelated content or become inconsistent across views. We introduce EditBench3D, a representation-agnostic benchmark that treats editing as controlled information replacement. It evaluates four complementary properties: instruction fidelity, spatial locality, cross-view consistency, and preservation of non-target content. The protocol combines visibility-aware 3D target supports, paired descriptions, held-out cameras, and five edit families covering appearance, material, geometry, and object-level changes. We evaluate eight representative NeRF, 3D Gaussian Splatting, hybrid, and proxy-based editors on 240 scene-edit pairs. The study shows that semantic fidelity is only weakly associated with the other editing properties, and that no single method is optimal across all dimensions. Explicit Gaussian editors offer a strong overall balance, whereas direct proxy manipulation provides the most conservative edits at the cost of open-ended fidelity. These findings support reporting editability as a multi-objective profile rather than a single semantic score.

View source

Similar papers

Preprint Oct 2026

View Matters: Keyframe-Guided Text-Driven 3D Gaussian Editing

Text-driven 3D Gaussian editing commonly does not distinguish the editing reliability of rendered views, although different viewpoints provide supervision of substantially different quality. Views that clearly show the scene and match the edit instruction provide reliable guidance, while less informative views may weak...

Kai-Zhe Zhang, Yi-Jie Zhou, Wei-Zhan Zhang et al. · 0 citations
#artificial intelligence Review Oct 2026

EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

3D editing methods are usually tested on a single edit, yet an asset is built through a long sequence of revisions, each of which must implement the requested change while leaving everything else unchanged. We introduce EditHero, to our knowledge the first benchmark for long-horizon, part-level 3D editing, with natural...

Rui-Han Yu, Yu-Ju Tsai, Mu-Yao Niu et al. · 0 citations
Review Open access Aug 2026

Semantic 3D Gaussian Splatting: A State-of-the-Art Review

A unified multi-axis taxonomy is introduced that enables us to classify the available methods in 3D Gaussian splatting methods in terms of five complementary categories: semantic vocabulary space, representation form, functional role, knowledge source, and query mechanism.

J. Flotyński · 0 citations
Preprint Aug 2026

ReRef-3D: A Benchmark for Spatial Referring Expression-Guided 3D Scene Rearrangement

We introduce ReRef-3D, a benchmark for language-guided placement in 3D scenes. It contains 33,826 instructions across 998 CLEVR-derived scenes, spanning 16 placement families and direct, one-hop, and two-hop references. Each instruction must be resolved into a valid new placement position. Given that an instruction def...

Mary Lynn Martin, Yifei Zhang, Martha Palmer et al. · 0 citations
#computer vision Preprint Sep 2026

InfoEdit: Probing Global Layout Reasoning in Infographic Editing

Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element often requires surrounding elements to be adapted. We refer to...

Cheng Yang, Chu-Fan Shi, Hui-Juan Wang et al. · 0 citations
Preprint Aug 2026

ES3D: Embedding Semantics into 3D Space for Component-Aware Editing

Existing 3D editing methods have made notable progress in controllability, yet they remain limited in several important ways. Most approaches rely on text-driven editing, which struggles to express fine-grained visual changes intended by the user. Moreover, many methods require manually supplied 3D masks or introduce u...

Xuancheng Jin, Ren-Gan Xie, Jiayuan Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.