This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents, and concludes with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures.
Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti et al.· 2 citations
An overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation is presented, covering spoken visual question answering and image-grounded hallucination detection in English and Modern Standard Arabic, and CRAI-Bench, evaluating the cultural accuracy of text-to-image generation.
This work evaluates four omni LLMs in a zero-shot setting and shows that fine-tuning consistently outperforms zero-shot inference, and explores synthetic data augmentation by using an LLM to generate culturally grounded Tunisian Derja utterances, followed by voice cloning to generate synthetic speech.
Tajwaar Shafiq, Hunzalah Hassan Bhatti, S. Chowdhury et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.