Jul 2026· Computer graphics forum (Print)· 0 citations· 52 references
Computer Science
TL;DR
This paper introduces a novel generative model for close 3D human‐human interactions using a conditional variational autoencoder (cVAE), which generates poses for one human conditioned on the pose of another, allowing for controlled and diverse interaction synthesis.
Abstract
Creating realistic 3D human‐human interactions in virtual environments is challenging due to the high degrees of freedom in human body and the need for physically accurate poses that do not collide with each other. Traditional methods for humanhuman interaction are based on motion tracking or 3D body reconstruction, but lack generative capabilities. Recent generative methods enable the synthesis of individual or interacting motions via text or image input, but generally fall short in modeling close interactions. This paper introduces a novel generative model for close 3D human‐human interactions using a conditional variational autoencoder (cVAE), which generates poses for one human conditioned on the pose of another, allowing for controlled and diverse interaction synthesis. To train our model, we address two underlying long‐standing challenges in the field of human‐human interaction: data scarcity, for which we propose an automated supervised data augmentation strategy that generates synthetic yet realistic interaction poses; and collision awareness in generative approaches, for which we propose a self‐supervised loss based on a collision resolution technique using volumetric proxies to ensure physically correct interactions. We extensively evaluate the capabilities of our model, and demonstrate a wide variety of plausible and physically correct interactions, not possible to generate with current state‐of‐the‐art methods.
We present a real-time human-centric world model for upper-body interactive generation, aiming to synthesize coherent local world dynamics centered on a person, where coordinated body, hand, and facial motions evolve jointly with controllable human-object discrete interaction. To this end, we adopt a continuous-discrete joint control scheme with two complementary components: a continuous human state and a discrete interaction state. For continuous human-state control, we introduce a unified implicit representation based on multi-scale motion encoding, in which motion latents from the upper body, hands, and face are fused into a shared latent space. This multi-scale design improves expressiveness across different spatial scales, captures fine-grained human dynamics more effectively, and enables direct control without explicit retargeting. For discrete object interaction-state control, we represent object contact using a small set of language-encoded discrete interaction states, where text serves as an explicit interaction-state command, such as \emph{no contact} or \emph{grasp}, rather than an open-ended generation prompt, and we further construct a dedicated rendering pipeline for human-object interaction data to supervise such discrete interaction states. By combining continuous implicit human-state control with discrete interaction-state control, our model enables precise modeling of how a person moves and interacts with the local environment, including controllable changes to nearby scene states. Finally, we distill the model for efficient streaming real-time inference, achieving 25 FPS on two H100 GPUs. Experiments demonstrate improved fine-grained motion fidelity, more realistic hand-object coordination, and effective real-time interaction, establishing a practical step beyond motion reproduction toward real-time human-centric world modeling.
Chaonan Ji, Jinwei Qi, Peng Zhang et al.· 0 citations
Safe human-to-humanoid motion imitation is crucial for shared environments, where direct motion retargeting may induce self-collision or human–humanoid collision due to embodiment mismatch, kinematic limits, perception uncertainty, and human proximity. This paper presents an online vision-aided safe human-to-humanoid motion imitation framework that integrates skeleton-based upper-body pose estimation, joint-space retargeting, and capsule-based Control Barrier Function Quadratic Program (CBF-QP) safety filtering. Human skeletal observations are mapped to a reduced eight-degree-of-freedom (8-DoF) humanoid upper-body command, while the CBF-QP layer computes a safety-corrected target that minimally modifies the nominal imitation command subject to robot self-collision and human–humanoid collision constraints. The framework is evaluated via simulation and hardware experiments under representative self-collision and human–humanoid interaction episodes, complemented by a comparative benchmark against velocity damping and potential field baselines. Furthermore, this work introduces an evaluation protocol combining geometric safety, command deviation, and local-link similarity metrics to systematically characterize the safety–imitation trade-off governed by CBF parameters. The results demonstrate that, within the tested moderate-speed regime, the proposed framework substantially reduces geometric collision violations while balancing imitation fidelity with online computational feasibility, thereby providing a viable foundation for safe human-guided humanoid motion deployment.
Wenqi Cai, John Abanes, N. Evangeliou et al.· Big Data and Cognitive Compu...· 0 citations
Synthesizing realistic full-body human interactions with articulated objects is a fundamental challenge for embodied AI and graphics, with applications in robotics training and virtual agents. Existing models remain limited: some focus on simple activities with static objects, while others restrict attention to hand-only manipulation. This leaves open the problem of generating coordinated full-body motion that approaches, manipulates, and moves articulated objects in a realistic and generalizable way. The key difficulty lies in reasoning jointly about locomotion, fine-grained contact, and object articulation. Models must capture subtle hand-object correspondences that transfer across object geometries, while also producing seamless transitions from navigation to manipulation. At the same time, the scarcity of large-scale paired motion-scene data makes it difficult to generalize across diverse object positions and shapes. We introduce a text-conditioned diffusion model that addresses these challenges through three core ideas: an object-centric representation that unifies hand-object contact with object surfaces, a mixed-domain training strategy that balances locomotion and interaction, and a contact-based augmentation scheme that expands training diversity. Through experiments, our method demonstrated strong generalization to unseen object configurations, surpassing current state-of-the-art methods.
Xiaohan Zhang, Sebastian Starke, Alexander Winkler et al.· 0 citations
Pose-guided human image generation aims to synthesize an image of a target person based on a reference image and a target pose. Although diffusion-based methods have recently achieved significant progress in pose alignment and visual realism, generating high-fidelity images in regions involving complex pose interactions remains a challenge. Typical issues include structural confusion of limbs and texture distortion in occluded areas. To address this challenge, this paper proposes the HiPA-Gen framework, which is designed to automatically perceive complex interaction regions and generate high-fidelity content within them. Unlike existing methods that mainly rely on flat skeleton conditions or implicit attention responses, HiPA-Gen explicitly models hierarchical spatial relations and local appearance priors in complex interaction regions. Specifically, we design a Dual-Agent Hierarchical Reasoning Module (DHR) where a Prompt-Guided Reasoning Agent identifies interaction-related body parts and a Hierarchical Rendering Agent converts the inferred relations into a Hierarchical Correspondence Pose (HCP) map. The HCP map provides explicit front–back structural cues for limb overlap, hand occlusion, and body self-occlusion. To further reduce local texture ambiguity, we introduce a Part Prior Detail Alignment Module (PDA), which extracts Regional Reference Assets (RRAs) from the source image under HCP-guided part priors. These regional references are then injected into the Pose-Conditioned Interaction-Aware Detail Synthesis Network (PDS) through a Local Enhancement Branch (LEB), enabling more accurate local feature fusion during diffusion-based synthesis. Experiments on DeepFashion and Market-1501 show consistent numerical improvements under the reported evaluation protocols in structural similarity, perceptual quality, and distributional fidelity. Qualitative results further show that the proposed framework produces clearer limb boundaries, more reliable spatial ordering, and more consistent clothing textures in challenging pose interaction scenarios.
Juncheng Zhu, Haotian Yang, Mubai Li et al.· Electronics· 0 citations
This work presents HiPHI, a 600+ hour scale high-fidelity whole-body human motion dataset designed to systematically maximize coverage of the human motion and interaction manifold, and introduces a benchmark suite evaluating motion-space diversity, interaction grounding, object consistency, and physical AI applications.
Jiahao Ji, Ji Ma, Runhan Zhang et al.· 0 citations
Realistic animal motion for virtual production is typically obtained either through motion capture of highly trained performers who accurately mimic animal behavior, or by retargeting ordinary human motion using complex control setups. Both approaches are challenging and often fail to fully reproduce the nuances of natural animal motion, motivating data-driven alternatives. We present an automatic human-to-quadruped puppeteering framework that produces plausible and controllable quadruped motions from ordinary human motion data. Our approach employs a two-stage generative diffusion model trained purely on quadruped motion data. By introducing a structured conditioning and inpainting strategy, our method supports a wide range of actions, including walking, running, jumping, sitting, and lying. Furthermore, we enable fine-grained intuitive control of the quadruped motion such as head movement control and individual limb puppeteering. Experimental results demonstrate improved motion realism and controllability compared to existing retargeting approaches, highlighting the effectiveness of our framework as a tool for animation and virtual production applications.
Fatemeh Zargarbashi, Zehong Qiu, Dhruv Agrawal et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.