Vision Language Models (VLMs) should rely on visual evidence that directly determines the correct answer, but supervision for grounding visual reasoning is often expensive to obtain manually or tied to dataset-specific annotation primitives. We instead introduce model-causal visual evidence as an annotation target, def...
BOW is introduced, an RL framework that instead trains models to produce self-contained, neutral, and comprehensive descriptions of the plausible next-word space, and human evaluation shows that BOW-Reg produces broader next-word reasoning trajectories, while direct next-word-prediction evaluation shows that these traj...
This work proves a PAC-Bayes bound guaranteeing that a dictionary extracted from successful trajectories has bounded expected description length on future successful behavior, and introduces ReuseRL, which grounds agentic RL in the Minimum Description Length (MDL) principle.
Zhikun Xu, Yu Feng, Jacob Dineen et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.