Preprint
Aug 2026
Question-Guided Evidence Acquisition for Multimodal Visual Question Answering
Q-Guide is built, a small agent that reads a question, works out what evidence it is still missing, and calls targeted tool(s) to recover it---reading text where text is needed, zooming in where detail is needed, or grounding a region where position matters.
A. Popa
· 0 citations