Failing to See or Failing to Know? Attributing Errors in Vision-Language Models
A tree-structured framework is proposed that organizes failures in knowledge-intensive visual question answering into model-specific operational outcomes that support attribution-guided routing to targeted interventions, including image repair, entity support, question rewriting, and factual evidence.