Skip to content

Author

Tongtong Cao

We have 5 of 16 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies

Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape without stating where...

Ding-Sheng Liu, Yang-Zheng Wu, Mahboubeh Asadi et al. · 0 citations
Preprint Aug 2026

Faster-WAM: Do World Action Models Need Deep Action Modules?

Faster-WAM, an instantiation of DoT for WAMs, which docks a single-layer action head onto a 30-layer video backbone, achieves competitive performance on LIBERO and RoboTwin 2.0 while demonstrating strong out-of-distribution generalization on LIBERO-Plus.

Liheng Ma, Rui Yang, Zhan-Guang Zhang et al. · 4 citations
Jul 2026

KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation

KAM-WM is presented, a framework that extracts a coarse directional interaction cue from a frozen latent video world model without rollout or world-model fine-tuning and indicates that a frozen video model can provide a useful first-order visual prior for control without the test-time cost of future rollout.

Xinyu Shao, Keru Zhou, Guowei Huang et al. · 1 citation
Jul 2026

RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning

RoboHarness is proposed, a unified framework that encapsulates independently developed robot control systems as reusable agentic skills as well as a general framework compatible with a broader range of robot policies, such as navigation policies, model predictive controllers, and world-action models.

Jinbang Huang, Yuan Hu, Zhiyuan Li et al. · 3 citations · ⚡1
Jul 2026

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation

Future-State-Conditioned VLN (FSC-VLN), a deployable model that augments a causal policy with a future-query token and uses training-only future-state supervision to distill information from future observations into the policy state, is proposed.

Lingfeng Zhang, Zhanguang Zhang, Liheng Ma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.