Skip to content

Language-Bias-Resilient Visual Question Answering via Adaptive Multi-Margin Collaborative Debiasing

2025 · Neural Information Processing Systems · pp. 139410-139429 · 2 citations · 49 references
Computer Science

TL;DR

A novel Multi-Margin Collaborative Debiasing (MMCD) framework is proposed, which adaptively integrates frequency-aware, confidence-aware, and difficulty-aware angular margins with a dynamic, difficulty-aware contrastive learning mechanism to reshape decision boundaries under biased training conditions.

Abstract

Language bias in Visual Question Answering (VQA) arises when models exploit spurious statistical correlations between question templates and answers, particularly in out-of-distribution scenarios, thereby neglecting essential visual cues and compromising genuine multimodal reasoning. Despite numerous efforts to enhance the robustness of VQA models, a principled understanding of how such bias originates and influences model behavior remains underdeveloped. In this paper, we address this gap through a comprehensive empirical and theoretical analysis, revealing that modality-specific gradient imbalances, which originate from the inherent heterogeneity of multimodal data, lead to skewed feature fusion and biased classifier weights. To alleviate these issues, we propose a novel Multi-Margin Collaborative Debiasing (MMCD) framework 2 , which adaptively integrates frequency-aware, confidence-aware, and difficulty-aware angular margins with a dynamic, difficulty-aware contrastive learning mechanism to reshape decision boundaries under biased training conditions. Extensive experiments across multiple challenging VQA benchmarks confirm the consistent superiority of our proposed MMCD over state-of-the-art baselines in combating language bias.

View source

Similar papers

Conference Jul 2026

Improving Trustworthiness in Visual Question Answering Via Question-Conditioned Cross-Modal Verification

Visual Question Answering (VQA) is a challenging cross-disciplinary task that combine natural language processing and computer vision together. VQA answer natural language questions based on images. Current VQA systems prone to generate plausible but incorrect responses without indicating uncertainty. To address this limitation, we propose a training-free cross-modal answer verification framework. This framework based on question-conditioned CLIP scoring as a post-hoc reliability estimator for vision-language models. The technique improves the ability to distinguish between right and wrong short-form VQA answers by evaluating candidate solutions jointly with the original question using type-aware distractor pools. Experiments on the VQAv2 validation set evaluate raw accuracy, verified accuracy, coverage, and accuracy gain under threshold-based selective prediction. Results shows consistent reliability improvements across diverse model architectures with gains ranging from +1.13 to +32.98 percentage points depending on coverage and model characteristics. The proposed framework is model-agnostic, requires no retraining or architectural modification, and improves the trustworthiness of multimodal systems by selectively filtering unreliable predictions. These findings show that question-conditioned CLIP verification provides an effective and scalable reliability layer for VQA systems.

Prakhar Shukla, Ankit Kumar, Pulkit Singh et al. · 0 citations
Open access Jul 2026

Prompt-Driven Fuzzing Debiasing Framework for Robust Visual Question Answering

A unified prompt-driven debiasing framework that integrates generative prompt learning and a fuzzing-based bias correction mechanism is proposed, which significantly improves both in-distribution accuracy and out-of-distribution robustness, outperforming existing prompt-only or data-augmentation-only debiasing methods.

Ya-Li Fan, Gang-Yu Huang, Qiwen Lu et al. · 0 citations
2026

A Multi-Dimensional Evaluation of Explainability in Media Bias Detection

It is suggested that predictive performance, attribution plausibility, and mechanistic faithfulness characterize different aspects of model behavior and should be evaluated separately when studying explainability in media bias detection.

Tinghao Chen, Raina Zhang, Benjamin M. Ampel et al. · 0 citations
Preprint Aug 2026

Toward Fine-Grained Forgetting:Attribute Unlearning for Multimodal Large Language Models

This work proposes Causal Localization and Retain-Aware Projection (CLRP), a lightweight training-free framework that uses activation patching to identify the layer that causally mediates target-attribute disclosure, then applies a retain-aware projection that removes the target-attribute subspace while preserving same-identity evidence.

Junkai Lin, Junkai Chen, Siqi Hou et al. · 0 citations
Preprint Jul 2026

Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation

This work proposes Expert-Guided Mutual Distillation (EGMD), which learns what evidence to trust across the prediction pipeline, and constructs Weibo_Balanced, a domain-balanced benchmark that isolates the effect of imbalance on generalization.

Xuan Feng, Guihong Liu, Tianlong Gu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.