This dataset contains 5,000 manually annotated Bangla-language text instances collected and curated for the detection of racism and body-shaming content on Bangla social media. It accompanies the paper "An Optimized Multi-Transformer Framework with Explainable AI for Racism and Body-Shaming Detection in Bangla Social Media." Files Bangla_Racism_BodyShaming_Dataset.csv — the full labeled dataset (UTF-8 encoded) Format CSV, UTF-8 (with BOM for Excel compatibility) Fields Column Description Sentence The target Bangla sentence being classified Sentiment Label: "Body Shaming" or "Racism" Story with Violence A short contextual narrative embedding the sentence, framed with a violent/harsh undertone Story without Violence A short contextual narrative embedding the same sentence, framed neutrally Class distribution Body Shaming: 2,573 samples Racism: 2,427 samples Total: 5,000 samples Language Bangla (Bengali script) Intended use Training and evaluation of transformer-based models (e.g., BanglaBERT, mBERT, MuRIL, XLM-RoBERTa) for hate speech / abusive language detection in low-resource Bangla NLP, and for explainable AI (XAI) research using techniques such as SHAP. Notes Samples were manually verified prior to inclusion. One fully empty row present in the original raw export was removed during cleaning. Please cite the accompanying paper if you use this dataset.
Shanjida Alam Hena, Antar Sarker, Tanveer Hasan et al.· Mendeley Data· 0 citations
This dataset contains 5,000 manually annotated Bangla-language text instances collected and curated for the detection of racism and body-shaming content on Bangla social media. It accompanies the paper "An Optimized Multi-Transformer Framework with Explainable AI for Racism and Body-Shaming Detection in Bangla Social Media." Files Bangla_Racism_BodyShaming_Dataset.csv — the full labeled dataset (UTF-8 encoded) Format CSV, UTF-8 (with BOM for Excel compatibility) Fields Column Description Sentence The target Bangla sentence being classified Sentiment Label: "Body Shaming" or "Racism" Story with Violence A short contextual narrative embedding the sentence, framed with a violent/harsh undertone Story without Violence A short contextual narrative embedding the same sentence, framed neutrally Class distribution Body Shaming: 2,573 samples Racism: 2,427 samples Total: 5,000 samples Language Bangla (Bengali script) Intended use Training and evaluation of transformer-based models (e.g., BanglaBERT, mBERT, MuRIL, XLM-RoBERTa) for hate speech / abusive language detection in low-resource Bangla NLP, and for explainable AI (XAI) research using techniques such as SHAP. Notes Samples were manually verified prior to inclusion. One fully empty row present in the original raw export was removed during cleaning. Please cite the accompanying paper if you use this dataset.
Shanjida Alam Hena, Antar Sarker, Tanveer Hasan et al.· Mendeley Data· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.