Skip to content
Book Open access

Towards Building a Standard Benchmark for Low-Resource Nepali Information Retrieval

Jul 2026 · Annual International ACM SIGIR Conference on Research and Development in Information Retrieval · 0 citations · 28 references
Computer Science

Abstract

Despite being spoken by over 26 million people, Nepali lacks a standardized Information Retrieval (IR) benchmark, preventing systematic evaluation of retrieval models in this low-resource setting. While progress has been made in several Nepali NLP tasks, no publicly available test collection with queries and relevance judgments exists for ad-hoc retrieval task. We propose the creation of a first standardized Nepali IR benchmark supporting monolingual (Nepali queries \rightarrow Nepali documents), cross-lingual (English queries \rightarrow Nepali documents), and code-mixed (Nepali-English queries \rightarrow Nepali documents) retrieval. We argue that Nepali provides a compelling testbed for studying morphology-aware retrieval, code-mixed query processing, and multilingual embedding transfer.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.