Beyond Surface Features: Advancing Medical Vision-Language Alignment via Dynamic Evidence-Guided Preference Optimization
Dynamic Evidence-Guided Preference Optimization (DEPO) is proposed, a new framework that enables evidence-aware and adaptive preference learning for Med-LVLMs and introduces Multi-Modal Evidence Perturbation (MEP) to suppress non-causal textual and visual shortcuts and Dispre-ferred Evidence Resampling (DER) to continuously update dispreferred responses as hallucination patterns evolve.