Context-Aligned Latent Enhancement for Multimodal Hate Meme Detection
Abstract
Detecting hate memes on social media presents a formidable challenge due to their multimodal and often subtle nature. The hateful intent typically emerges from a nuanced interplay between visual and textual elements, which existing methods often fail to capture by inadequately modeling these cross-modal correlations. To address this gap, we propose semantic context alignment and latent enhancement network (SCALE), a novel framework designed to achieve a deep understanding of hateful content. Our approach operates in two core stages. First, for Semantic Context Alignment, we introduce an attention-enhanced multimodal alignment (AEMA) module. It dynamically aligns image and text features into a unified semantic space, enabling a fine-grained grasp of the visual-textual context. Second, to achieve Latent Enhancement, we employ a lightweight convolutional semantic enhancement network (CSEN). This network refines the aligned and fused representations, distilling high-level, abstract semantics that are crucial for deciphering implicit hateful meanings. Extensive experiments on three benchmark datasets validate the superiority of SCALE. Notably, it achieves a new state-of-the-art area under the receiver operating characteristic curve (AUROC) of 92.95% on the HarMeme dataset, significantly outperforming prior methods and demonstrating its effectiveness in tackling the complexities of hate meme detection.