The shared task is a shared task for evaluating hallucination detection and factual verification in Arabic question answering under challenging generalization settings, based on two Arabic datasets: HalluScore and HalluTruthQA.
Aisha Alansari, Abdessalam Bouchekif, A. Hasanaath et al.· 0 citations
This work builds a retrieval test collection for Arabic fiqh and uses it to evaluate dense, lexical, hybrid, fine-tuned, and madhhab-aware retrieval strategies, and presents an error analysis showing that the main challenge is distinguishing answer-bearing passages from topically similar passages that do not contain th...
Somaya Eltanbouly, Heba Sbahi, Samer Rashwani et al.· 0 citations
HalluTruthQA-4K provides a reusable resource for hallucination detection, span-level error localization, explanation generation, factual verification, and the broader evaluation of factual reliability in Arabic language models.
S. E. Bekhouche, Abdessalam Bouchekif, H. Telli et al.· 0 citations
Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selec...
Abdessalam Bouchekif, M. Zighem, S. E. Bekhouche et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.