Skip to content

GroundBench: A Factorized, Counterfactual Benchmark for Locating VLM Affordance Failures

Sep 2026 · 0 citations · 20 references
Computer Science

TL;DR

GroundBench is introduced, a diagnostic benchmark that separates explanations through six branch-and-merge conditions, each adding a controlled information bundle, and a counterfactual re-ask targeting a real alternate part visible in the same image.

Abstract

A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language models, with no model outperforming a constant baseline until the part was named. However, naming the part supplies information that a real system must infer, confounding visual grounding, mechanical reasoning, and category-to-action association. We introduce GroundBench, a diagnostic benchmark that separates these explanations through six branch-and-merge conditions, each adding a controlled information bundle, and a counterfactual re-ask targeting a real alternate part visible in the same image. Across three OpenAI models and 1,068 predictions, supplying the target region without its identity leaves action accuracy at or below the 0.53 majority baseline (0.26, 0.26, and 0.53), although the models largely reproduce the supplied region. Supplying identity without location instead yields 0.74, 0.68, and 0.68. Every above-baseline gain in this curated set occurs where the supplied part category itself determines the action. A no-vision control leaves GPT-5's scores unchanged or improved, providing evidence consistent with substantial category-to-action association. GPT-4o mini declines on one condition, so this interpretation is not universal. Adding joint type and motion axis does not improve accuracy across six model-stratum comparisons. On 74 counterfactual pairs from 32 objects, GPT-5 achieves 0.86 pair-weighted compliance with a 0.07 shortcut rate but fails all observed push-to-lift-vertical cases. GroundBench identifies which supplied information changes affordance behavior and tests whether apparently grounded performance can be reproduced through textual shortcuts.

View source

Similar papers

#machine learning Preprint Sep 2026

Part Grounding, Not Action Knowledge: Locating the Bottleneck in VLM Affordance Prediction

Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy score does not say which step fails. We separate two steps that affordance questions conflate: identifying which part of an object to act on, and knowing what action that part requires. Across 19 articulat...

Sarthak Sattigeri · 0 citations
Preprint Sep 2026

RoboIRGBench: Benchmarking Implicit Referential Grounding in Vision-Language-Action Models

Vision-Language-Action (VLA) models have shown strong capabilities in robotic manipulation, yet existing benchmarks typically assume that task-relevant information is explicitly specified in the instruction. In practice, however, humans frequently refer to objects, quantities, and relations implicitly, requiring robots...

Aernaer Akelijiang, Jian-Nan Li, Zhi-Neng Chen et al. · 0 citations
#artificial intelligence Preprint Aug 2026

LOCI: A Locator-Critic with Refinement Loop

Locator-Critic (LOCI) is proposed, a training-free framework that decouples visual search from evidence verification and improves accuracy for both open-weight models like Qwen3-VL and proprietary models like Gemini 2.5 Pro.

Walid Bousselham, Mathilde Caron, Arsha Nagrani et al. · 0 citations
#machine learning Preprint Sep 2026

From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models

Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterf...

Rong-Yu Xu, Prayag Tiwari, Shao-Lei Zhang · 0 citations

Related blog posts

Microsoft Research Blog Aug 11, 2026

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.