Skip to content

Author

Dingyi Rong

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.

Dingyi Rong, Yue Shi, Chaofan Ma et al. · 1 citation
Jul 2026

EnerBridge-DPO: Energy-Aware Markov Bridge Inverse Folding for Protein Sequence Design

Designing protein sequences with favorable predicted energetic properties is an important challenge in protein inverse folding, because many existing deep learning methods are primarily trained by maximizing sequence recovery and do not explicitly incorporate energy-related preferences during generation. In this work, we propose EnerBridge-DPO, an energy-aware inverse folding framework that integrates Markov bridge sequence generation with preference optimization for protein complex design. The framework builds on the Markov bridge inverse-folding process to generate structure-compatible sequences from an informative prior sequence. It then introduces a Bridge-DPO objective that uses energy-related winner-loser preference pairs to bias the generator toward sequences favored by computational or experimental energy-related signals. In addition, we incorporate a quantitative energy-constrained loss based on mutation-induced binding free-energy changes to provide continuous ΔΔG supervision. Evaluations show that EnerBridge-DPO maintains competitive inverse-folding performance while obtaining lower predicted energy scores under selected computational scoring functions for protein complexes. On SKEMPI, EnerBridge-DPO achieves competitive ΔΔG prediction performance, with small numerical gains in several overall metrics that are not statistically conclusive under paired bootstrap analysis. These results suggest that incorporating energy-related preferences into Markov bridge inverse folding can improve computationally predicted energetic profiles, although experimental validation is required to confirm thermodynamic stability.

Dingyi Rong, Haotian Lu, Xupeng Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.