This paper introduces the task of identifying and segmenting legal conditions (Tatbestand) and legal consequences (Rechtsfolge) within German statutory texts and presents ANNOTARES (Annotations of Tatbestand-Rechtsfolge Sequences), a novel dataset comprising German law texts with span-level annotations.
Abstract
The automatic structural analysis of legal texts is a cornerstone of legal technology, yet the extraction of their logical components remains a significant challenge. In this paper, we introduce the task of identifying and segmenting legal conditions (Tatbestand) and legal consequences (Rechtsfolge) within German statutory texts. To support this task, we present ANNOTARES (Annotations of Tatbestand-Rechtsfolge Sequences), a novel dataset comprising German law texts with span-level annotations. Spanning three distinct legal codes, the dataset is designed to evaluate both domain-specific performance and cross-statute generalizability. We benchmark diverse architectural approaches: a rule-based baseline, CRFs, BiLSTMs, BiLSTM-CRF, and modern Transformer-based models, including BERT variants and LLM-based methods. Our results demonstrate that BERT and LLM-based models achieve superior performance in capturing the complex syntactic structures of legal language. We release our dataset to facilitate further research in automated legal reasoning.
This paper outlines a unique method of legal text processing using Natural Language Processing (NLP) technology to extract the information from the legal texts meaningfully and naturally. The proposed system is designed in a data pipeline architecture by integrating the NLP functionalities such as tokenization, part-of-speech tagging, named entity recognition (NER), and dependency parsing to systematize the processing of typologies of legal text hubs, including legal briefs, statutes, and case law. The methodology presented concerns the importance of pre-processing legal texts that address domain-specific challenges. The texts may contain ambiguities, while the legal language itself is a very intricate kind of language. The system uses advanced methods like syntactic parsing and semantic role labeling to parse and find relevant entities, relationships, and context, ensuring the automation of the large amount of raw legal data for review and analysis. Besides that, the first is leveraging machine learning models to optimize the data extraction process and to ensure high efficiency and scalability. This methodology guarantees that accurate and reliable information is extracted and reduces the time and costs that conventionally come with manual legal analysis. The focal point of the offered system is overcoming legal workflow issues and bringing model texts to widespread use. Therefore, the proposed system aims to facilitate decision-making processes in legal practice and even the accuracy of the proposed model.
S. A. Gade, Sivaram Ponnusamy· Journal of Intelligent Decis...· 0 citations
This work presents PROSLEX (PRediction Of Statutes and LEgal eXplanation), a comprehensive dataset comprising 1,623 expert-annotated legal documents from the Indian context, positioning PROSLEX as a benchmark for developing explainable AI systems that can support legal practitioners while advancing research in interpretable legal NLP.
Subinay Adhikary, Upal Bhattacharya, Vivek K. Singh et al.· 0 citations
Legal research traditionally relies on qualitative1 analysis of legal texts, including case law. Yet, the growing number of judicial decisions makes systematic manual review increasingly difficult. This paper demonstrates how lightweight computational techniques, accessible to researchers with limited technical expertise, can support doctrinal legal research through the systematic collection, enrichment, and delimitation of legal corpus. Using asylum case law of the Swiss Federal Administrative Court (SFAC) as a case study, we design a transferable pipeline producing a structured database covering most decisions since 2007. Focusing on sexual orientation and gender identity (SOGI) asylum claims, the pipeline first identifies a SOGI-related corpus, representing around 1.25% of the collected database, which is then refined through citation network reconstruction. The pipeline enables the identification of structurally relevant decisions for qualitative analysis while producing data suitable for quantitative and computational research.
Mélanie Braillard, Hugo Hueber· Proceedings of the 2026 ACM...· 0 citations
The proposed LeDA system, a system for Legal Data Annotation for Legal Data Annotation, offers the generic functionality of annotating and adjudicating entities or concepts within documents via a web-based interface and allows to dynamic create new tags for annotation.
Legislative knowledge evolves as an intricate hypertext in which documents are interconnected through complex, often implicit relationships. In this paper, we introduce ReSB2, a framework for retrieving and linking similar legislative bills that supports human–machine collaboration and helps reduce redundancy in the lawmaking process. The framework fine-tunes two ModernBERT-based language models on authentic legislative data, incorporating domain-specific formatting and procedural constraints derived from real workflows in a Brazilian state-level legislative assembly. To ensure transparency, ReSB2 integrates an explainability module based on Integrated Gradients, enabling analysts to inspect which textual elements most influence model decisions. Evaluated on a large corpus of official bills, the framework outperforms both general-purpose and domain-specific baselines in identifying semantically similar documents, achieving recall values of approximately 0.9. Human-centric evaluation with domain experts further demonstrates that ReSB2 serves as an effective human-centered augmentation tool, supporting the consistency and governance of legislative knowledge.
Lucas G. L. Costa, Átila Souza, Elves Rodrigues et al.· Proceedings of the 37th ACM...· 0 citations
Named entity recognition (NER) is a core information extraction (IE) task dependent on high-quality annotated data that are expensive and time-intensive to produce. Large language models (LLMs) offer a promising alternative through LLM-generated pseudo-annotations, yet their reliability in domain-specific legal settings remains insufficiently studied. This study investigates the use of LLM-generated annotations to expand the training set for supervised NER models applied to sentences from Dutch administrative decisions as a low-resource domain and language. To this end, LLM-based annotations of predefined legal entities are created using a schema-driven few-shot prompt, which are evaluated against a human-annotated dataset. Next, two different NER architectures are trained—a token-level NER model and a span-based NER–RE model (joint NER and relationship extraction (RE))—under three training settings: (1) human annotations only, (2) LLM-generated annotations only, and (3) models trained on LLM-generated annotations further fine-tuned on human annotations. The results indicate that LLMs can accurately generate annotations for legal entities that are explicitly defined in legislation but generate less reliable annotations for other legal entities that require deeper contextual understanding beyond what is explicitly stated in the text or prompt. Furthermore, fine-tuning a NER model trained on these LLM-generated annotations with human annotations turns out to slightly outperform models trained on human-annotated data only. Our findings highlight the potential of hybrid supervision strategies to scale low-resource legal NER tasks while maintaining human-level accuracy.
H. Nan, Samaneh Khoshrou, Johan Wolswinkel· Journal of Computational Law...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.