Skip to content

Author

Salim Rahman

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Dataset Open access Sep 2026

Bangla Mental Health Dataset V2

This dataset is a synthetically generated Bangla-language mental health dataset consisting of 10,000 structured conversational samples. It is designed to support research in natural language processing (NLP), particularly for low-resource languages such as Bangla, with applications in large language model (LLM) fine-tuning, mental health text classification, and dialogue system development. Each sample follows an instruction-based format (input–instruction–output), making the dataset directly suitable for supervised fine-tuning (SFT), Alpaca-style training, and parameter-efficient methods such as LoRA and QLoRA. The dataset captures a diverse range of approximately 40 mental health-related conditions, including stress, anxiety, overthinking, lack of emotional support, and self-confidence issues, expressed in natural Bangla conversational patterns. The dataset is fully synthetic and was generated using controlled text generation pipelines informed by mental health literature, psychological reports, media discussions, and publicly available educational content. No real user data or personally identifiable information (PII) is included. This dataset is intended strictly for research and educational purposes. It is not suitable for clinical use, diagnosis, or real-world mental health decision-making. The resource aims to facilitate safe and reproducible experimentation in Bangla NLP and conversational AI.

Esfer Sami, Salim Rahman · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.