Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization
Cross-Modal Visual Feedback (CMVF) incorporates a failure-conditioned visual diagnosis stage, in which a stronger optimizer VLM inspects each failed image without access to predictions or labels, and an error-aware aggregation stage that compresses these observations into reusable, task-level visual blind-spot patterns that drive the prompt rewrite.