Engineering a Legal Corpus: A Lightweight and Transferable Pipeline for Building and Refining a Case-Law Corpus
Legal research traditionally relies on qualitative1 analysis of legal texts, including case law. Yet, the growing number of judicial decisions makes systematic manual review increasingly difficult. This paper demonstrates how lightweight computational techniques, accessible to researchers with limited technical expertise, can support doctrinal legal research through the systematic collection, enrichment, and delimitation of legal corpus. Using asylum case law of the Swiss Federal Administrative Court (SFAC) as a case study, we design a transferable pipeline producing a structured database covering most decisions since 2007. Focusing on sexual orientation and gender identity (SOGI) asylum claims, the pipeline first identifies a SOGI-related corpus, representing around 1.25% of the collected database, which is then refined through citation network reconstruction. The pipeline enables the identification of structurally relevant decisions for qualitative analysis while producing data suitable for quantitative and computational research.