Towards Building a Standard Benchmark for Low-Resource Nepali Information Retrieval
Despite being spoken by over 26 million people, Nepali lacks a standardized Information Retrieval (IR) benchmark, preventing systematic evaluation of retrieval models in this low-resource setting. While progress has been made in several Nepali NLP tasks, no publicly available test collection with queries and relevance judgments exists for ad-hoc retrieval task. We propose the creation of a first standardized Nepali IR benchmark supporting monolingual (Nepali queries \rightarrow Nepali documents), cross-lingual (English queries \rightarrow Nepali documents), and code-mixed (Nepali-English queries \rightarrow Nepali documents) retrieval. We argue that Nepali provides a compelling testbed for studying morphology-aware retrieval, code-mixed query processing, and multilingual embedding transfer.