Interpreting Multimodal Gender Bias Scores Using Explainable AI
Abstract
Gender bias in artificial intelligence systems tends to take root in the data itself, yet the tools available for detecting and auditing it are still largely built around individual modalities with no unified way to assess the problem across all of them at once. This challenge is especially pronounced in multimodal settings because image, text, audio, and video datasets encode representational imbalance through different structural mechanisms. This paper presents the Multimodal Dataset Bias Interpretation Framework (MDBIF), a datasetcentered and model-independent framework for detecting and interpreting multimodal gender bias using explainable artificial intelligence (XAI). For each modality, MDBIF extracts interpretable features and computes a modality-specific bias score. In the image modality, object-context imbalance is combined with a Face Influence Score (FIS) derived from facial-region analysis, CLIP-based similarity, and depth cues. In the text modality, bias is quantified using Contextual Embedding Bias (CEB) and attention-weighted Weighted Feature Contribution (WFC) based on BERT representations. In the audio modality, a polynomial Acoustic Bias Score (ABS) captures gender-related acoustic imbalance using voiced pitch, energy, amplitude, and voice-activity statistics, including quadratic pitch interaction terms. In the video modality, person prominence, centering, screen time, motion, and temporal embeddings are integrated through a PCA-weighted bias score. To make the scores trustworthy, we built in a validation layer that runs three independent explanation methods: perturbation testing, feature ablation, and SHAP attribution, and checks how well they agree. We evaluated the framework across four gender-labeled datasets covering image, text, audio, and video. In each case, the bias patterns were interpretable, and the feature rankings produced by different methods lined up closely, with Kendall's $\tau$ between 0.76 and 0.89 and a Top 3 overlap of 83% or above throughout. This suggests MDBIF is a practical tool for auditing dataset-level gender bias before any model is trained on it.