Skip to content

Human4K: A Large-Scale 4K Multi-View Mocap Dataset for Whole-Body 3D Human Reconstruction

Jul 2026 · arXiv.org · Vol abs/2607.13646 · 0 citations · 50 references
Computer Science

TL;DR

Experimental results show that training with Human4K consistently improves whole-body reconstruction on standard benchmarks, with particularly large gains for hands, feet and depth-ambiguous limb configurations.

Abstract

Recent advances in 3D human reconstruction have improved overall performance, yet current models still fail in the most challenging real-world scenarios. They often produce unstable geometry, inaccurate limb articulation and unreliable predictions under depth ambiguity or self-occlusion. A key reason is that existing datasets still lack the combination of high-resolution images, high-precision annotations and diverse whole-body motions required to support robust reconstruction. To address this gap, we present Human4K, a large-scale 4K multi-view whole-body human reconstruction dataset with mocap-accurate SMPL-X annotations. Human4K contains over six million 4K images captured by an eight-view high-resolution camera system synchronized with a professional Vicon motion capture setup, covering 11 subjects performing complex, highly articulated and strongly self-occluded full-body motions. All sequences are processed by a Motion-Retargeting and Refinement Module (MRRM) to ensure precise alignment for the full body and extremities. Experimental results show that training with Human4K consistently improves whole-body reconstruction on standard benchmarks, with particularly large gains for hands, feet and depth-ambiguous limb configurations.

View source

Similar papers

Jul 2026

Multiview Multi-Person Human Mesh Recovery Under Large Scenes with Occlusions

Human mesh recovery (HMR) aims to recover 3D human meshes from images. Most existing HMR benchmarks and methods focus on either multi-person reconstruction from a single view or single-person reconstruction from multiple views, where the number of subjects and the scene scale are relatively limited. Such settings are insufficient for real-world applications with large scenes and severe inter-person occlusions. To address this limitation, we introduce a large-scale synthetic benchmark for multiview multi-person HMR, termed MVMP-HMR. The proposed dataset contains 15 complex scenes with up to 50 camera views and 30 interacting persons, featuring large spatial coverage and severe occlusions, which significantly increases the difficulty of human mesh recovery. Based on this benchmark, we further propose a multiview multi-person whole-body human mesh recovery model, referred to as MVMP-HMR model. The model first fuses multiview features into a scene-level 3D feature volume, and then leverages pelvis joints predicted by a 3D pose estimation network to extract person-specific queries from the 3D feature volume. These human queries are cross-attended with the 3D feature volume and integrated to decode each person's 3D mesh. Moreover, we introduce two novel losses--the orientation loss and the 3D joint density loss--to alleviate orientation and pose ambiguities under severe occlusions. Experiments demonstrate that existing state-of-the-art HMR methods struggle on the proposed MVMP-HMR benchmark, while our method consistently outperforms prior SOTAs in large-scale scenes with severe occlusions.

Qi Zhang, Tao Yu, Jiechao He et al. · 0 citations
Open access Aug 2026

HardMo++: A Large-Scale Hard-Case Dataset for Motion Capture

Recent years have witnessed rapid progress in monocular human mesh recovery. However, even strong benchmark models still exhibit systematic failures under unusual poses, especially hands, feet, and side-view configurations. Such failures are common in dance and martial arts but remain underrepresented in current datasets. We observe that these limitations mainly stem from insufficient pose diversity and inaccurate annotation of extreme joints rather than inherent architectural deficiencies. To address this issue, we introduce HardMo++, a scalable benchmark dataset built by an automated pipeline that combines large-scale data collection, re-annotation, and targeted optimization for systematic failure cases. HardMo++ contains approximately 830K images covering 15 dance categories, 14 martial arts categories, and diverse daily motions, with emphasis on wrist–hand and ankle–foot configurations. For the multi-view FreeMan-derived portion, cross-view refinement is additionally used to reduce side-view ambiguity. We also conduct a human verification study on 1000 sampled images as a descriptive check of annotation quality. Extensive experiments show that models trained with HardMo++ improve over the corresponding 4DHumans- and HardMo-trained baselines on the proposed hard-case benchmarks and retain competitive performance on standard benchmarks.

Junran Peng, Silei Shen, Zong-Xing Li et al. · 0 citations
Preprint Aug 2026

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

This work introduces HumanTracker, a preference-aligned metric trained on 12K motion pairs containing 24K motions that better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.

Dai-En Liu, Zekun Qi, Jiayu Zeng et al. · 0 citations
Jul 2026

FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility

Reconstructing animatable 3D human avatars from monocular video is a fundamental problem in computer vision with broad applications in AR/VR and digital content creation. Existing approaches typically couple parametric body models with neural rendering or 3D Gaussian splatting and optimize all body regions jointly from short videos, which often degrades fidelity in the visible areas. To overcome this limitation, we introduce FlexiAvatar, a unified framework that explicitly optimizes only the visible body regions, effectively eliminating artifacts arising from unobserved limbs. Our method integrates occlusion-robust SMPL-X tracking with part-specific residual refinement to capture high-frequency geometric and appearance details. To complete entirely unseen regions (e.g., back views), we leverage a diffusion-based approach to generate texture consistent with the observed appearance. Experiments on full-body (NeuMan, ZJU-MoCap, WildAvatar), upper/half-body (talk-show clips), and head-only (INSTA) inputs show that FlexiAvatar delivers consistently higher reconstruction quality, outperforming state-of-the-art methods by an average PSNR improvement of approximately 3% across datasets. Finally, by restricting optimization to observed regions, our method reduces the effective number of Gaussians that must be optimized and rendered, leading to reduced runtime and memory overhead in partial-visibility scenarios.

Yihalem Yimolal Tiruneh, M. Ali, Uyoung Jeong et al. · 0 citations
#artificial intelligence Preprint Aug 2026

HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction

This work presents HiPHI, a 600+ hour scale high-fidelity whole-body human motion dataset designed to systematically maximize coverage of the human motion and interaction manifold, and introduces a benchmark suite evaluating motion-space diversity, interaction grounding, object consistency, and physical AI applications.

Jiahao Ji, Ji Ma, Runhan Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.