The shared task is a shared task for evaluating hallucination detection and factual verification in Arabic question answering under challenging generalization settings, based on two Arabic datasets: HalluScore and HalluTruthQA.
Aisha Alansari, Abdessalam Bouchekif, A. Hasanaath et al.· 0 citations
HalluTruthQA-4K provides a reusable resource for hallucination detection, span-level error localization, explanation generation, factual verification, and the broader evaluation of factual reliability in Arabic language models.
S. E. Bekhouche, Abdessalam Bouchekif, H. Telli et al.· 0 citations
Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selec...
Abdessalam Bouchekif, M. Zighem, S. E. Bekhouche et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.