We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed YouTube speech span...
Bashar Talafha, Samar M. Magdy, Aisha Alansari et al.· 0 citations
This paper presents the Mawqif-XT, consisting of 996 manually annotated Arabic tweets collected from three public targets: Women Driving, E-Cars, and Trimester System, which provides a benchmark for evaluating cross-target generalization in Arabic stance detection.
BUL, a multi-dialect Arabic ASR dataset collected from 275 speakers in 11 Arab countries, includes structured dialect and sub-dialect coverage, as well as recordings of classical Arabic and modern standard Arabic spoken by participants in their native dialectal accents to support accent-aware modeling.
Ahmed Ashraf, Aisha Alansari, Fadel Al Abbas et al.· 1 citation
CrossHallu is presented, the first study to evaluate the cross-lingual and cross-domain generalization of hallucination detection using internal representations from six LLMs on the generative question-answering task.
Aisha Alansari, Malak Alkhorasani, H. Luqman· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.