Scaling robot data and model capacity has improved Vision-Language-Action (VLA) policies, but further progress is constrained by the high cost of robotic data. Verifier-guided test-time scaling offers an efficient alternative by sampling multiple action candidates and selecting the one most likely to lead to task succe...
Seongheon Park, Heecheol Kim, Shu-Lin Tian et al.· 0 citations
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding them requires reasoning...
Shu-Lin Tian, Junsu Kim, Shuai Liu et al.· 0 citations
Vision-language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single unsupported object, chart value, or intermediate inference can undermine an otherwise plausible response. We argue that this is a credit-assignment failure in multimodal post-training. Scalar outcome rewa...
Shu-Lin Tian, Ming-Lun Li, Yuhao Dong et al.· 0 citations
The Evaluation Agent framework is proposed, which employs human-like strategies for efficient, dynamic, multi-round evaluations, offering detailed, user-tailored analyses and is efficient, promptable, explainable, and scalable across models and tools.
Shu-Lin Tian, Zi-Qi Huang, Fan Zhang et al.· 2 citations
Apple-PI is introduced, the first benchmark that anchors video-model evaluation explicitly in physical laws, and is positioned as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.
Runmao Yao, Kairui Hu, Yukang Cao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.