Semantic takhrij for cross-collection Hadith retrieval: a leakage-controlled evaluation framework
Abstract
Hadith scholarship requires tracing semantically related reports across collections, variants, and transmission settings, yet existing computational work concentrates on lexical search, narrow question answering, or heterogeneous evaluation setups that limit scholarly usefulness and weaken cross-study comparability. This article specifies semantic takhrij as a distinct cross-collection retrieval task in which a query Hadith is used to recover canonically grounded, semantically related evidence across collections, and constructs a leakage-controlled evaluation framework suited to the structural properties of Hadith corpora. The framework defines the query unit, candidate evidence space, graded relevance schema, provenance-aware output record, annotation logic, family-aware split philosophy, stable canonical identifiers, transparent normalization rules, and documentation requirements needed for auditable and reproducible evaluation. A four-level relevance model treats exact matches, near-parallel reports, semantically supportive evidence, and nonrelevant results as distinct outcomes rather than collapsing them into a single similarity class, while a family-aware partitioning strategy prevents evaluation-breaking leakage inherent in cross-collection Hadith settings. This is the first work to formalize semantic takhrij as a provenance-aware, graded, and leakage-controlled retrieval task, integrating task definition, annotation logic, and split discipline as a single coherent methodological unit, thereby advancing digital humanities scholarship by providing operational design vocabulary for evidence-oriented retrieval in Arabic religious corpora and directly supporting reproducible and source-grounded computational study of Islamic textual traditions.