Maps are central to how humans make sense of the world, from navigation and environmental monitoring to military planning and historical interpretation. Yet despite rapid progress in large multimodal models (LMMs), these systems continue to struggle with interpreting maps – an essential skill for visual reasoning that...
Christian M. Arnold, Andrew Alini, Abdulrahman Alabdulkareem et al.· Proceedings of the 28th Inte...· 1 citation
To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a physically grounded hardware-software co-design challenge: engineering fully functional robots...
Kushagra Tiwary, David Mayo, Nikhil Behari et al.· 0 citations
Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute...
Vighnesh Subramaniam, B. Katz, Brian Cheung et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.