Skip to content

Author

Cheng-Qun Yang

We have 5 of 6 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis

The capability to perceive and synthesize human-human interactions is fundamental to developing intelligent digital human systems. However, existing datasets and modeling approaches are fundamentally constrained by low-fidelity kinematics, the omission of dexterous hand gestures and a severe lack of rich multimodal annotations. Furthermore, fragmented interaction representations and inconsistent evaluation protocols also impede fair and rigorous benchmarking. To systematically address these bottlenecks, we present Inter-X++, a comprehensive and large-scale benchmark designed to empower versatile HHI analysis. Captured via a novel hybrid motion capture system, Inter-X++ provides 11,388 high-fidelity interaction sequences and over 8.1M frames, featuring precise whole-body movements and detailed finger articulations. Meanwhile, we enrich the data foundation with multifaceted annotations, including hierarchical fine-grained textual descriptions, interaction categories, causal interaction orders, the relationship and personality of the subjects, as well as vertex-level contact maps and physically regularized constraints. Leveraging these elaborate annotations, we formulate a unified testing ground comprising four categories of downstream tasks that symmetrically span both generative and perceptive paradigms. To eliminate benchmarking ambiguities, we systematically standardize the interaction representations and evaluation protocols. Finally, we go beyond dataset construction to propose OpenHHI, a single and unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding. Extensive experiments reveal that OpenHHI achieves state-of-the-art performance on both generation and perception tasks. This definitively proves that our unified representation successfully bridges interaction understanding and generation simultaneously.

Liang Xu, Chengqun Yang, Zili Lin et al. · 0 citations
Preprint Jul 2026

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control

This work presents Enfold, which transfers this computation that constructs a future into a representation predicted from the current visual context and language instruction, and recast a world generator as a source of predictive control representations if its internal structure can be enfolded into the present.

Wei-Li Zeng, Yi-Tong Xing, Fu-Long Liu et al. · 1 citation
Preprint Aug 2026

Spatiotemporally Decoupled Autoregressive Diffusion Model for Human Motion Generation

A unified spatiotemporally decoupled framework named DeMoDiff is proposed, which jointly redesigns representation and architecture and incorporates spatial-temporal masking and attention mechanisms into an autoregressive diffusion generator, achieving both generative capability and controllable editability.

Chengqun Yang, Liang Xu, Yanping Li et al. · 0 citations
Preprint Aug 2026

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval

This work proposes a lightweight granularity-aware model anchored at a frozen standard-caption-aligned retrieval model that improves mixed-granularity retrieval without compromising standard-caption performance, and believes that its MRBench provides a comprehensive testbed for advancing motion-language alignment evaluation.

Fulong Liu, Liang Xu, Chengqun Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.