Skip to content
Open access

MPNet-Based Semantic Retrieval for Legal Evidence Selection: A Set Cover and Subset Selection Perspective

Oct 2026 · Engineering, Technology & Applied Science Research · 0 citations · 13 references

Abstract

Legal decision support often depends on multiple conditions distributed across policy and regulatory documents. Retrieval-only systems can identify individually relevant passages, but may return redundant or collectively insufficient evidence. This study proposes a Legal NLP framework that combines MPNet-based semantic retrieval with controlled-cardinality evidence selection. Retrieved clauses are organized into compact subsets by jointly optimizing condition coverage, semantic relevance, redundancy, and subset size. The evidence-selection stage is formulated as a constrained set-cover and subset-selection problem. Experiments were conducted on LexIR-PolicyEvidence, a clause-level corpus comprising 850 policy and regulatory documents, approximately 14,000 clauses, and 1,800 query or case-description instances. The evaluation compared BM25, MPNet top-k retrieval, MPNet with greedy set cover, and the proposed method. MPNet improved Recall@5 from 0.68 to 0.79, MRR from 0.61 to 0.73, and nDCG@10 from 0.66 to 0.78. The proposed selection method achieved coverage of 0.86, sufficiency of 0.84, and redundancy of 0.24. The results demonstrate that retrieval quality alone does not ensure decision-ready legal support and that explicit evidence construction is required to produce compact and interpretable clause bundles.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.