Jul 2026· Annual International ACM SIGIR Conference on Research and Development in Information Retrieval· pp. 3568-3573· 0 citations· 29 references
Computer Science
TL;DR
It is shown that agreement on binary labels and system ordering is generally high, with no cases of significant opposite conclusions; however, there are cases of false or missed improvements and noticeable limitations when labelling manual, highly performing systems.
Abstract
Large Language Models (LLMs) are increasingly used in Information Retrieval, both within retrieval pipelines and for constructing evaluation resources. Existing studies on using LLMs for IR evaluation, however, focus almost exclusively on English, leaving their applicability to other languages, where evaluation resources are often limited and highly needed, unexplored. We examine the use of LLMs to generate relevance labels for an Arabic test collection (ArTest). Using about 10K relevance labels, generated by three LLMs, and used to order eight automated systems and nine simulated manual systems, we show that agreement on binary labels and system ordering is generally high, with no cases of significant opposite conclusions; however, there are cases of false or missed improvements and noticeable limitations when labelling manual, highly performing systems. These findings align with results reported for English and indicate that LLMs could, with some caveats, offer a viable approach to supporting IR evaluation in languages with limited evaluation resources.
Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu...
Farah Adeeba, A. Khan, Rajesh Bhatt et al.· 0 citations
In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the...
This study investigates the register variation in texts written by humans and comparable texts produced by large language
models (LLMs). Multidimensional analysis (MDA) is applied to a sample of human-written texts and AI-generated counterpart texts to find the
dimensions of variation in which LLMs differ most si...
Jiří Milička, Anna Marklová, V. Cvrček· International Journal of Cor...· 0 citations
An important practical reproducible framework for multilingual RAG benchmarking and insights for optimizing performance on resource-constrained devices are contributed and support the latent language hypothesis by suggesting an internal model bias toward high-resource languages.
B. J. Mohd, Khalil M. Ahmad Yousef, Salah G. Abu Ghalyon· Language Resources and Evalu...· 0 citations
An in-depth evaluation of instruction tuning for Arabic NLP tasks using three prominent LLMs: LLaMA 3.1-8B, AceGPT-v2-8B, and Qwen3-8B shows that instruction tuning consistently improves performance across most tasks, with notable variations in effectiveness across different tasks and prompts.
Maged Saeed Al-shaibani, Zaid Alyafeai, Irfan Ahmad· Language Resources and Evalu...· 0 citations