You Know What I Mean: A Benchmark for Agentic Conversational Reference Grounding
The results show that CoRG remains challenging for current agents, even the best agent reaches only 67.0% success rate, leaving one third of references unresolved, and position CoRG as a concrete benchmark for studying how agents search, inspect, and verify information in realistic multi-tool environments.