Understanding ongoing robot manipulation requires models to interpret visual observations in relation to interaction history and task progress. We introduce RoboChrono, a benchmark for streaming task understanding comprising 39 scenarios and 34,713 evaluation instances, constructed from real robot executions and comple...
Yu-Zhou Wu, Long-Teng Fan, Zi-Meng Li et al.· 0 citations
FedOGL preserves historical decision behavior through replay and task-start distillation, while protecting graph-propagation memory via projection onto a globally shared structure basis via projection onto a globally shared structure basis.
Ze-Kai Chen, Haodong Lu, Shi-Hao Li et al.· arXiv.org· 0 citations
This work proposes ReBind, a systematic framework that introduces semantic instructions with embedded reference tokens as the intermediate representation for multi-reference image-conditioned video editing and develops ReBind-Instruct, a specialized MLLM that learns to establish explicit bindings between visual attribu...
Xin-Yu Liu, Shi-Hao Li, Weihong Lin et al.· arXiv.org· 2 citations
A multimodal federated graph unlearning framework built around target-specific representation decoupling that effectively removes requested information, preserves retained graph utility, and achieves a speedup over full retraining.
Haodong Lu, Zekai Chen, Weiwei Ji et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.