Robot Question Answering (RQA), a modular evaluation framework and RoboVista, a benchmark curated from real robotic systems, research papers, and expert annotations are proposed and physical robot experiments suggest strong correlation between RoboVista performance and real-world task execution.
Abstract
Diverse applications for robotics, such as industry and agriculture, require robots to operate across various embodiments, changing visual conditions, and complex planning. Vision-Language Models (VLMs) offer a promising foundation for general-purpose and interpretable robotic reasoning. Aligning VLMs with diverse robot applications requires a modular understanding of the individual decision components that underlie robotic behavior. Capturing such structure is challenging for conventional robot benchmarks that are primarily based on teleoperated, end-to-end datasets. We propose Robot Question Answering (RQA), a modular evaluation framework and RoboVista, a benchmark curated from real robotic systems, research papers, and expert annotations. RoboVista contains 474 Visual Question Answering (VQA) instances with human annotated reasoning and covers 39 unique task types in agricultural, industrial, domestic, surgical robotics, autonomous driving, and open robot datasets. Experiments on RoboVista show that state-of-the-art VLMs exhibit substantial gaps. Physical robot experiments suggest strong correlation between RoboVista performance and real-world task execution.
RoboBRIDGE is presented, a modular framework that provides an orchestration layer over five coordinated modules, namely Monitor, Perceptor, Planner, Controller, and Robot Interface, to compose robust robotic agents from off-the-shelf components, including pretrained VLAs.
Experimental results in various task scenarios show that the proposed framework consistently improves overall task success rates compared with unimodal settings with different LLMs and achieves a higher success rate compared to using only visual or force data.
Evaluating robot manipulation policies is becoming increasingly important as generalist models, particularly vision-language-action (VLA) models, are deployed on physical robots. However, conventional real-world evaluation remains labor-intensive, unstable, and insufficiently informative. It requires repeated hardware trials, manual scene resets, and continuous operator monitoring, may produce different policy rankings across repeated evaluations, and primarily relies on success-rate metrics that provide limited information about execution quality. In contrast, humans assess robot performance by observing and comparing complete behaviors rather than relying solely on binary success outcomes. To this end, we propose R2S-Eval, an evaluation pipeline that combines real-to-sim calibration with vision-language model (VLM) preference evaluation. The real-to-sim component efficiently generates rollout videos in a simulator calibrated to the real-world evaluation setting, thereby reducing the need for repeated hardware trials. The VLM evaluator assesses the execution quality of rollout videos and produces pairwise preferences, which are subsequently aggregated into policy rankings. We further introduce a protocol to assess whether the proposed evaluation pipeline yields validated policy conclusions while mitigating the key challenges of conventional real-world evaluation. Experiments in both simulation and real-world settings demonstrate that R2S-Eval produces reliable and stable policy conclusions, achieves agreement with human preferences, substantially reduces repeated hardware-operation effort, and reveals behavior-quality differences that are not captured by binary success labels. In general, R2S-Eval advances robot evaluation from manual success counting toward automated, statistically stable, and quality-aware evaluation of robot behavior. Project page: https://r2s-eval.github.io.
Yi-Di Wang, Fei-Xiang Ruan, Ruo-Qu Chen et al.· 1 citation
OmniCAD is introduced, a large-scale benchmark for assembly-aware 3D spatial reasoning across diverse industrial systems, including robotic mechanisms, automotive components, aerospace structures, and agricultural machinery, and tool-augmented agentic reasoning.
Mingjia Wang, Taiting Lu, Ziwei Dong et al.· 0 citations
Generalist robots need to perform diverse tasks while operating in dynamic, uncertain, and unstructured environments, often around human beings. Vision-language-action (VLA) models have recently emerged as a promising and flexible framework for integrating perception, reasoning, robotic control, and action execution to develop generalist robotic policies. This systematic literature review (SLR) examines more than 140 VLA-related publications between 2020 and 2025 following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. To the best of our knowledge, it is the first PRISMA-compliant systematic review dedicated to VLA models, offering a structured discussion of robotic policies, VLA architectures, and inference optimization methods. The review also presents descriptive analyses of the included studies and a glossary defining the terminology commonly used in VLA and generalist robotic policy research. The findings reveal substantial diversity among VLA models in terms of their supported modalities, robotic embodiments, training strategies, and architectural designs. Despite the rapid growth of VLA research, several important areas remain underexplored, including the execution of complex, long-horizon tasks, effective integration of speech, and deployment on low-cost hardware, while ensuring robust, safe, and secure operation.
Umair Cheema, Y. Badr, T. Le et al.· Robotics· 0 citations
This work adapts Hugging Face's SmolVLA for Universal Robots lightweight robots, and releases the open-source repository ROS2SmolVLA that implements an interface for ROS 2 to SmolVLA, and makes it applicable for industrial-grade hardware.
Nils Mandischer, Noah Böckmann, Ludwig Holl et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.