Skip to content
Preprint

Prior-First, Condition-Second: Scalable and Controllable Hand Motion Completion

Jul 2026 · 0 citations · 39 references
Computer Science

TL;DR

This work proposes a prior-first, condition-second framework for body-conditioned hand motion completion that improves kinematic plausibility, robustness, and controllability compared to end-to-end conditioned baselines, particularly in low-resource and cross-dataset settings.

Abstract

Synthesizing hand motion that matches the full body motion and the semantic labels is a difficult task due to their high degrees of freedom and the lack of semantic labels. To cope with this issue, we propose a prior-first, condition-second framework for body-conditioned hand motion completion. Our framework first learns a generic body-hand kinematic prior from large-scale unstructured and unlabeled motion data, capturing the intrinsic coordination between global body dynamics and hand articulation. Semantic control is then introduced through lightweight adaptation on top of the frozen prior, avoiding the need to relearn kinematic structure for each control interface. Our framework centers on a streaming, autoregressive body-hand prior that generates coherent, kinematically consistent hand motion from body dynamics in real time, using structured kinematic modeling to maintain mechanical body-hand coupling. To enable practical controllability under limited supervision, we introduce semantically-layered adapters that inject conditioning signals at appropriate kinematic levels, supporting both self-supervised attribute control and weakly supervised text-driven control with only a few hours of labeled data. Extensive evaluations demonstrate that our framework improves kinematic plausibility, robustness, and controllability compared to end-to-end conditioned baselines, particularly in low-resource and cross-dataset settings. We further showcase real-time inference and an interactive authoring workflow, highlighting the applicability to production animation pipelines. Homepage: https://AIGAnimation.github.io/HandPrior/

View source

Similar papers

Jul 2026

PAKE: Learning Whole-Body Loco-Manipulation with Partial Kinematic Embeddings

This paper introduces a new Kinematic Normalizing Flow (KNF) model, trained on a large-scale kinematic dataset, that generates diverse yet feasible partial reference motions that effectively addresses whole-body control challenges by decomposing the complex loco-manipulation problem into partial reference motion generation and low-level imitation control.

Zhengmao He, Moonkyu Jung, Hyeongjun Kim et al. · 0 citations
Open access Aug 2026

U-GRA: Uncertainty-Gated Residual Adaptation for Physically Robust Three-Finger Grasping

Robust three-finger grasping under physical-domain variation remains challenging because contact stability can change substantially with object mass, effective friction, and observation noise. This work develops U-GRA, a conservative offline-to-online residual adaptation framework for simulated three-finger grasping. U-GRA introduces a unified prior-preserving and critic-disagreement-regulated architecture that couples a frozen behavioral prior with a spectrally normalized and bounded residual stream, scalar Twin-Q reliability assessment, and critic-conditioned residual fusion. The framework first learns a nominal behavioral prior from successful demonstrations and then freezes it as a stable action anchor during online adaptation. Before execution, the twin critics evaluate a candidate action formed from the prior action and the bounded residual proposal, and their absolute scalar Q-value disagreement conditions a state-dependent gate that regulates residual-injection strength. Experiments are conducted in CoppeliaSim using an offline dataset of 40,000 successful demonstrations and online randomization of object mass, effective friction, and observation noise. Across three independent seeds, U-GRA achieves a mean success rate of 84.8±2.3%, a normalized return of 82.7±4.1, and a jitter value of 0.12±0.03. Relative to AWAC-Res, the strongest evaluated baseline, U-GRA improves mean success by 9.2 percentage points and reduces jitter by 57.1%. It also retains the highest mean success rate and normalized return over the unseen simulated high-mass–low-friction OOD region. These results provide simulation evidence that preserving a nominal behavioral prior while regulating bounded residual correction through critic disagreement improves three-finger grasping robustness under physical-domain variation.

Juncheng Zhu, Zhan Gao, Zhile Yang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

GigaBrain-WBC-0.5, the first Behavior World Model for humanoid whole-body control, is presented, which trains a causal Transformer to jointly predict its next action, next state, and the distribution over its next latent behavior command, so the network that acts also models how the environment shapes what it can do next.

Ziyang Cheng, Tianshu Tang, Jinxi Lan et al. · 0 citations
Preprint Aug 2026

DynamicWAM: Dual-Path Motion Conditioning for World-Action Models in Dynamic Manipulation

DynamicWAM introduces history-flow conditioning, encoding temporally aligned optical-flow frames alongside the current observation through a frozen pretrained video VAE to preserve spatial motion structure, while injecting kinematic descriptors of displacement, duration, velocity, and acceleration into the action expert to provide motion magnitude and timing.

Y. Lou, Hewen Gao, Xiyu Zhu et al. · 0 citations
Preprint Jul 2026

LAMP: Latent Motion Prior-Guided Real-World Learning for Dexterous Hand Manipulation

A latent motion prior module (\prior{}) is introduced that maps recent hand-action histories to a compact, history-conditioned latent prior and decodes continuous latent commands into executable high-dimensional hand targets and improves the policy with online residual RL in the same latent hand-action space.

Xinye Yang, Zhiyuan Ma, Hongze Yu et al. · 0 citations
Preprint Jul 2026

KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation

Learning manipulation from few demonstrations requires visual priors that capture not only where to interact, but also how the interaction should begin; static priors such as segmentation masks encode only the former. We present KAM-WM, a framework that extracts a coarse directional interaction cue from a frozen latent video world model without rollout or world-model fine-tuning. KAM-WM queries a Flow Matching image-to-video backbone once and interprets its single-step latent velocity as a Kinematic Affordance Map (KAM), which provides task-conditioned interaction regions and coarse motion structure. A lightweight Perceiver compresses KAM into tokens that condition a diffusion policy together with RGB observations and proprioception. Across LIBERO and RoboTwin2.0, KAM-WM reaches 90.6% average success on LIBERO and achieves 65.7% and 22.4% success rates in the Easy and Hard settings on RoboTwin2.0, respectively. Controlled comparisons against a zero-order mask prior suggest that part of the gains comes from directional information beyond spatial localization alone. These results indicate that, in the evaluated settings, a frozen video model can provide a useful first-order visual prior for control without the test-time cost of future rollout.

Xinyu Shao, Keru Zhou, Guowei Huang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.