HardMo++: A Large-Scale Hard-Case Dataset for Motion Capture
Abstract
Recent years have witnessed rapid progress in monocular human mesh recovery. However, even strong benchmark models still exhibit systematic failures under unusual poses, especially hands, feet, and side-view configurations. Such failures are common in dance and martial arts but remain underrepresented in current datasets. We observe that these limitations mainly stem from insufficient pose diversity and inaccurate annotation of extreme joints rather than inherent architectural deficiencies. To address this issue, we introduce HardMo++, a scalable benchmark dataset built by an automated pipeline that combines large-scale data collection, re-annotation, and targeted optimization for systematic failure cases. HardMo++ contains approximately 830K images covering 15 dance categories, 14 martial arts categories, and diverse daily motions, with emphasis on wrist–hand and ankle–foot configurations. For the multi-view FreeMan-derived portion, cross-view refinement is additionally used to reduce side-view ambiguity. We also conduct a human verification study on 1000 sampled images as a descriptive check of annotation quality. Extensive experiments show that models trained with HardMo++ improve over the corresponding 4DHumans- and HardMo-trained baselines on the proposed hard-case benchmarks and retain competitive performance on standard benchmarks.