Skip to content
Book Open access

Does LLM Relevance Labelling Work for Arabic?

Jul 2026 · Annual International ACM SIGIR Conference on Research and Development in Information Retrieval · pp. 3568-3573 · 0 citations · 29 references
Computer Science

TL;DR

It is shown that agreement on binary labels and system ordering is generally high, with no cases of significant opposite conclusions; however, there are cases of false or missed improvements and noticeable limitations when labelling manual, highly performing systems.

Abstract

Large Language Models (LLMs) are increasingly used in Information Retrieval, both within retrieval pipelines and for constructing evaluation resources. Existing studies on using LLMs for IR evaluation, however, focus almost exclusively on English, leaving their applicability to other languages, where evaluation resources are often limited and highly needed, unexplored. We examine the use of LLMs to generate relevance labels for an Arabic test collection (ArTest). Using about 10K relevance labels, generated by three LLMs, and used to order eight automated systems and nine simulated manual systems, we show that agreement on binary labels and system ordering is generally high, with no cases of significant opposite conclusions; however, there are cases of false or missed improvements and noticeable limitations when labelling manual, highly performing systems. These findings align with results reported for English and indicate that LLMs could, with some caveats, offer a viable approach to supporting IR evaluation in languages with limited evaluation resources.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu...

Farah Adeeba, A. Khan, Rajesh Bhatt et al. · 0 citations
Jul 2026

Challenges in annotations by humans and LLMs: A case study of evaluative language

In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the...

Mirela Imamović, Aenne Knierim, Khushi Pitroda et al. · 0 citations
Sep 2026

Benchmark of linguistic variation in LLM‑generated texts

This study investigates the register variation in texts written by humans and comparable texts produced by large language models (LLMs). Multidimensional analysis (MDA) is applied to a sample of human-written texts and AI-generated counterpart texts to find the dimensions of variation in which LLMs differ most si...

Jiří Milička, Anna Marklová, V. Cvrček · 0 citations
Jul 2026

Benchmarking and AI-assisted human-like evaluation of retrieval-augmented generation for Arabic and English documents

An important practical reproducible framework for multilingual RAG benchmarking and insights for optimizing performance on resource-constrained devices are contributed and support the latent language hypothesis by suggesting an internal model bias toward high-resource languages.

B. J. Mohd, Khalil M. Ahmad Yousef, Salah G. Abu Ghalyon · 0 citations
Aug 2026

STAR: instruction tuning for Arabic across tasks, datasets, and models

An in-depth evaluation of instruction tuning for Arabic NLP tasks using three prominent LLMs: LLaMA 3.1-8B, AceGPT-v2-8B, and Qwen3-8B shows that instruction tuning consistently improves performance across most tasks, with notable variations in effectiveness across different tasks and prompts.

Maged Saeed Al-shaibani, Zaid Alyafeai, Irfan Ahmad · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.