Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

A two-stage hybrid intelligence framework for subject indexing via semantic embedding and LLM collaborative optimization

Approaches to subject indexing need to consider a trade-off between operational efficiency and the maintenance of indexing quality. Given that existing automated approaches struggle with domain adaptability, this study proposes the recall–rank–rerank (R3) framework—a two-stage hybrid intelligence method for automated subject indexing. By simulating human indexing cognition through case-based reasoning and human–machine collaboration trained on the TIBKAT data set, the results demonstrate that R3 significantly improves efficiency and semantic precision within the standard test collection evaluation framework in information retrieval research. R3 operates without model fine-tuning: its first stage uses cross-lingual semantic embedding to recall and rank candidate subjects from similar documents, while the second employs a large language model (LLM) as a simulated expert for refinement—calibrating semantics, enriching implicit concepts, and reranking relevance. Results show R3 outperforms supervised fine-tuning and semantic recall (R2) methods. On the validation set, it lifts mean average precision (MAP) from 41.19% to 45.24% (a 4% gain) and increases top-5 precision (P@5) by 2.49%. In the LLMs4Subjects task, R3 achieves top recall rates—65.68% for core subjects and 58.56% for all subjects—leading the SemEval’25 shared evaluation. With its lightweight design and strong performance, R3 offers adaptability and scalability, delivering a possible solution for large-scale indexing in practice. Fusing semantic embedding with LLM reasoning, the proposed framework can boost accuracy and provide new paths for hybrid intelligence in artificial intelligence–assisted subject indexing. Implications for the design and evaluation of automated subject indexing systems are discussed.

T. Xia, Xin Yang, Wenjing Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.