Multi-agent framework for span-level detection of homophobic and transphobic comments in Hindi
Abstract
Social media platforms serve as spaces for both open expression and the spread of harmful narratives, often targeting marginalized communities. In the context of Hindi, automated detection of homophobic and transphobic comments at span level is still not studied widely. This work introduces a span-annotated dataset of Hindi social media comments, where identity-based toxic comments are labeled at the span level using character-based indices and encoded in BIO (Begin, Inside, Outside) format to support token-level classification. To address the scarcity of annotated resources, we adopt a human-in-the-loop framework that combines expert-labeled seed data with systematically generated synthetic data using GPT-4o prompting, followed by human validation to ensure cultural and contextual accuracy. We proposed novel augmentation strategies for annotation homophobic/transphopic at the span level which incorporates context-level, entity-level, combined context/entity, and noise-level strategies, enabling richer and more diverse training material. The dataset was built in three stages: a seed version, an intermediate augmented version, and a final version. We benched-marked our dataset with multiple prompting strategies, including zero-shot, few-shot, and augmentation-based approaches, across both vanilla and cooperative multi-agent framework. Our proposed multi-agent framework decomposes the detection process into specialized roles such as self-annotation, span-related feature extraction, demonstration scoring, and final prediction, resulting in a structured and interpretable pipeline for span-level homophobic/transphobic speech detection. Experiments are conducted on both the test split and the combined extended test set, with evaluation using strict token-level, hybrid, and span-overlap metrics to capture both exact and tolerant boundary matches. Our proposed multi-agent framework using augmented V3 data on the combined extended test set, achieved a Macro F1 score of 0.90 , outperforming the strongest single-agent baseline and demonstrating the effectiveness of combining augmentation, multi-agent reasoning, and refined evaluation. Disclaimer: This study includes examples of toxic language strictly for research and educational purposes. The content does not reflect the views of the authors. Reader discretion is advised. GitHub Link: https://github.com/Unit-for-Inclusive-AI/Multi-Agent-HT-Span-Hindi