Skip to content
Preprint

DocTalkBN: A Novel Dataset of Expert Telemedicine Conversations in Bengali

Aug 2026 · 0 citations · 34 references
Computer Science

TL;DR

A large-scale multimodal dataset of real-world expert telemedicine conversations in Bengali, collected from nationally broadcast telemedicine programs featuring board-certified physicians, showing that DocTalkBN is a practically useful resource, particularly for clinically grounded reasoning tasks.

Abstract

Reliable medical conversational AI requires authentic expert--patient interaction data, yet such datasets remain scarce, especially for low-resource languages such as Bengali. We present DocTalkBN, a large-scale multimodal dataset of real-world expert telemedicine conversations in Bengali, collected from nationally broadcast telemedicine programs featuring board-certified physicians. DocTalkBN contains 557.63 hours of paired audio and text, 1,515 multi-turn patient calls, 10,274 host--doctor question--answer exchanges, totaling 1.7M tokens, spanning 26 medical specialties. Unlike prior resources derived from medical forums, written health content, or synthetic data, our dataset preserves the spontaneity, contextual richness, and spoken characteristics of authentic medical interactions in a low-resource setting. To support benchmark-driven research, we further construct three downstream tasks from the corpus, medical triage classification, advice safety evaluation, and medical named entity recognition, and benchmark a diverse set of large language models and encoder-based baselines. Our results show that DocTalkBN is a practically useful resource, particularly for clinically grounded reasoning tasks. We release this resource to facilitate future research on reliable medical NLP and safer, more culturally grounded healthcare systems for low-resource languages. Our source codes and dataset are publicly available at https://anonymous.4open.science/r/doctalk.

View source

Similar papers

Open access Jul 2026

DawAI: AI Powered Medical Chatbot

DawAI is presented, a multimodal AI-powered virtual medical assistant designed to simulate real-time doctor– patient interactions and offers a scalable, accessible, and user-friendly solution for preliminary medical consultation, particularly benefiting users in remote and resource-constrained environments.

Dr. Abdul Khadeer, Mohammed Zubair Ahmed · 0 citations
Preprint Jul 2026

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

A large-scale benchmark for multimodal online medical consultation built from de-identified patient-doctor interactions collected from a nationwide Chinese internet hospital, MedRealMM offers a realistic and reproducible benchmark for evaluating multimodal medical reasoning in real-world online consultation.

Runhan Shi, Quan Zhou, Yuqian Xu et al. · 0 citations
Jul 2026

IndicMedQA: Multimodal Medical Query Analysis in Indian Languages

This work introduces IndicMedQA, a novel multimodal AI framework that integrates Indic large language models (LLMs) and visual encoders to analyze patient inquiries using both textual and visual cues, and creates a multilingual multimodal medical corpus spanning seven major Indian languages, translated using a semi-automated approach.

Akash Ghosh, Arkadeep Acharya, M. Muhsin et al. · 2 citations
Aug 2026

MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation

This work presents MTDiag, a large multi-turn diagnostic dialogue dataset constructed from three heterogeneous sources: DDXPlus, MIMIC-IV, and published case reports, and introduces and motivate clinical knowledge-grounded metrics for evaluating LLMs as diagnostic agents, beyond diagnostic accuracy, for the task of multi-turn differential diagnosis.

Pia Chouayfati, Alexander M. Fichtl, Miriam Anschütz et al. · 1 citation
Open access Jul 2026

HOSPIT-LLM: A Human-Centered Multimodal Dataset and Edge-Deployed LLM Pipeline for Emotion-Aware Hospitality Assistants

Large language models (LLMs) exhibit strong general conversational capabilities, yet their deployment in domain-specific service environments such as hospitality remains limited by the absence of emotionally grounded datasets and validated end-to-end system architectures. This paper presents HOSPIT-LLM, an EU-funded euROBIN Technology Exchange Program pilot, as a complete, integrated pilot pipeline for human-centered conversational AI in hotel reception scenarios. We deploy a multimodal hotel-terminal assistant in a real hotel reception, capturing synchronized dual-camera video and audio to collect authentic guest–staff interactions. Speech is transcribed using Whisper, and emotion is extracted from the corresponding video segments via DeepFace, producing 582 real Greek guest–receptionist exchange examples. The resulting data are classified into eight Standard Operating Procedure (SOP) categories. To address data scarcity, we augment the corpus with 1269 synthetic dialogues generated by eight diverse LLMs through the OpenRouter API, yielding a total of 1851 dialogue records with explicit emotion-token annotation. We fine-tune Qwen3.5-35B-A3B using Low-Rank Adaptation (LoRA) through a two-stage process: supervised fine-tuning (SFT) on an 888-example conversation pool and Simple Preference Optimization (SimPO) on a 1899-pair preference pool, each split 80/10/10 into training, validation, and test. The resulting model is integrated into an interactive hotel-terminal system combining YOLO-based person detection, face-recognition-driven guest personalization, Kokoro neural text-to-speech (TTS), and a multi-service orchestration layer connected to the hotel Property Management System (PMS). Evaluation combines standard text metrics, emotion-aware scoring, and a large-model judge. The results indicate targeted improvements in the rule-based contextual emotion-policy match and staff-emotion policy compliance compared to the base model, while general response-quality gains remain more modest. In particular, the rule-based contextual policy-match score improves from 0.614 to 0.901, while forbidden staff-emotion outputs decrease from 0.142 to 0.018. The deployed pilot demonstrates the practical integration of a personalized, emotion-aware LLM assistant in an interactive hotel-terminal setting; end-to-end latency and fully hotel-side edge deployment were not evaluated and are left for future work. HOSPIT-LLM provides a reproducible framework for multimodal dataset creation, preference-based fine-tuning, and deployment of human-centered AI systems. A mobile robotic embodiment is planned as future work.

Homer Papadopoulos, Antonis Korakis, G. Balaskas · 0 citations
Preprint Jul 2026

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge

This paper describes TalTech's submissions to the Beyond Transcription Challenge (BeTraC), which requires generating SOAP notes directly from long doctor-patient conversation recordings, without intermediate transcription. After screening open-weight speech LLMs for long-audio robustness, we adapted Voxtral Mini (lightweight track) and Voxtral Small (heavyweight track) with LoRA supervised fine-tuning followed by DAPO reinforcement learning that uses the challenge metric, Open Medical Concept F1, as its reward. Our systems ranked first in both tracks, and an independent LLM-as-a-judge evaluation showed the lowest hallucination rate among all submissions, indicating that reinforcement learning against a concept-matching metric need not compromise factual reliability. We also find that fine-tuning on text transcripts transfers well to speech input and appears to improve robustness on out-of-domain real recordings.

Aivo Olev, Tanel Alumäe · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.