Skip to content

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

Jul 2026 · arXiv.org · Vol abs/2607.27616 · 0 citations · 45 references
Computer Science

TL;DR

This work introduces MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities, and proposes MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction.

Abstract

Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.

View source

Similar papers

Preprint Sep 2026

Revisiting Avatar-As-Image: High-Fidelity Registration is All You Need

The representation of 3D clothed humans as standardized 2D UV texture and displacement maps over an underlying body model has long been studied. This compact representation is enticing as it enables pretrained image networks to process, generate, and edit 3D avatars, but is only useful if scans are accurately aligned and brought into correspondence via high-fidelity registration. This prerequisite has never been met, which we argue explains the limited quality of prior UV-based methods for clothed humans. Despite its significance, no public method produces high-fidelity SMPL(-X)+D registrations with UV texture from arbitrary clothed scans. We present AvaImg, a multi-stage optimization pipeline, to close this gap: it enforces body-inside-clothing constraint via signed winding numbers, made viable by a three-level efficiency cascade (~10x runtime reduced, ~95% storage saved), and recovers fine surface detail using coarse-to-fine displacement optimization. AvaImg outperforms all baselines in body fitting, shape estimation, and surface registration across six datasets, yielding textured registrations near-indistinguishable from scans (PSNR=34.48dB). For validation of AvaImg's Avatar-as-Image representation as imminently compatible with image foundation models, we auto-encode our UV maps via the frozen FLUX VAE. This achieves only 0.76mm added Chamfer error relative to scan and shows that the resulting maps lie within natural-image distributions, supporting the use of 2D generative priors for 3D avatar generation. Code, data, and Singularity containers will be at https://yuxuan-xue.com/avaimg.

Margaret Kostyrko, Yuxuan Xue, Garvita Tiwari et al. · 1 citation
Preprint Aug 2026

RigidBench: Evaluating Rigid-Body Physics in Video Generation Models

RididBench is introduced, a simulator-grounded benchmark that compares a generated continuation with a reference rollout from the same initial frame and motion description, with per-frame masks, depth, 6-DoF trajectories, and contacts available for scoring.

Swarnim Jain, Shangzhe Wu · 2 citations
Preprint Aug 2026

Astrolabe: Spherical-Map Guidance Across Diffusion Pipelines for Full-Body Capture from Unconstrained Images

Astrolabe is a host-portable adapter built on frozen viewpoint-guided spherical maps (SPH), a host-portable adapter built on frozen viewpoint-guided spherical maps (SPH) that follows one SPH--shift--adapt--guide process without dense warping or a learned control branch.

Shuliang Zhu, Qi Wang, Ryugo Morita et al. · 0 citations
Preprint Aug 2026

DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion

DiGS-Avatar is proposed, which reformulates this task as an efficient, diffusion-based UV-latent completion task, ensuring 3D consistency by design, and introduces a teacher-student framework where a multi-view teacher provides geometrically aligned pseudo-ground-truth latents to supervise a single-view diffusion student.

Jiakun Li, Li Fang, Hao Zhu et al. · 0 citations
Jul 2026

FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility

Reconstructing animatable 3D human avatars from monocular video is a fundamental problem in computer vision with broad applications in AR/VR and digital content creation. Existing approaches typically couple parametric body models with neural rendering or 3D Gaussian splatting and optimize all body regions jointly from short videos, which often degrades fidelity in the visible areas. To overcome this limitation, we introduce FlexiAvatar, a unified framework that explicitly optimizes only the visible body regions, effectively eliminating artifacts arising from unobserved limbs. Our method integrates occlusion-robust SMPL-X tracking with part-specific residual refinement to capture high-frequency geometric and appearance details. To complete entirely unseen regions (e.g., back views), we leverage a diffusion-based approach to generate texture consistent with the observed appearance. Experiments on full-body (NeuMan, ZJU-MoCap, WildAvatar), upper/half-body (talk-show clips), and head-only (INSTA) inputs show that FlexiAvatar delivers consistently higher reconstruction quality, outperforming state-of-the-art methods by an average PSNR improvement of approximately 3% across datasets. Finally, by restricting optimization to observed regions, our method reduces the effective number of Gaussians that must be optimized and rendered, leading to reduced runtime and memory overhead in partial-visibility scenarios.

Yihalem Yimolal Tiruneh, M. Ali, Uyoung Jeong et al. · 0 citations
Preprint Aug 2026

ViSculpt: Visual-Centric Agentic Geometry Editing

This work presents a training-free multi-agent system that edits existing 3D meshes directly in Blender by emulating the iterative workflow of human artists, and views this work as an exploratory step toward visual-centric agentic geometry editing in professional graphics software.

Bo Pang, Jiaqi Pan, Xiao-Chen Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.