The proposed LeDA system, a system for Legal Data Annotation for Legal Data Annotation, offers the generic functionality of annotating and adjudicating entities or concepts within documents via a web-based interface and allows to dynamic create new tags for annotation.
Abstract
In this paper, we mainly concentrate on finding concepts or topics from the legal case proceedings, since adopting a structured representation for legal documents, as opposed to a mere bag-of-words flat text representation, can significantly enhance processing capabilities. To achieve this objective, we put forward a set of diverse concepts for legal case proceedings. With this motivation, we propose LeDA, a system for Legal Data Annotation. The system offers the generic functionality of annotating and adjudicating entities or concepts within documents via a web-based interface. A novel feature of our system is that it allows to dynamic create new tags for annotation, which is a particularly useful provision for situations where there exists no pre-defined ontology for the entities (concepts) that need to be annotated - these being rather discovered by annotators as they continue examining more documents. The system that we demonstrate is currently in use to annotate a set of concepts from legal documents to construct semantic representations of documents as bags of concepts that can then be used for several downstream tasks, such as prior case retrieval, judgment prediction, and so on. Along with the system features in general, we also describe how LeDA was used by 3 assessors to annotate and adjudicate legal concept names from Indian Supreme Court case proceedings.
This paper introduces the task of identifying and segmenting legal conditions (Tatbestand) and legal consequences (Rechtsfolge) within German statutory texts and presents ANNOTARES (Annotations of Tatbestand-Rechtsfolge Sequences), a novel dataset comprising German law texts with span-level annotations.
The preservation of rare documents in the form of image collections presents significant challenges regarding access to their documentary content. To enable this accessibility for software agents, this article proposes a formal representation of this type of document through a semantic description layer. This layer includes a set of descriptive metadata attached to the document, alongside the minimal and strictly necessary vocabulary required to formalize the explicit textual and visual knowledge of its documentary content. To achieve this, we present a construction methodology based on a Semantic Model of Document (SMD), where a document is treated as a core documentary resource containing a set of information resources. The semantic description of these resources, aligned with RDF framework logic, produces an Ontological Core of Document (OCD) that formally describes the document's logical structure and captures its underlying semantics. Finally, we demonstrate the practical utility of these Ontological Cores through three distinct use cases—each targeting a specific dataset level (structural, administrative, and semantic)—showing how they allow software applications to move beyond simple collection searching toward intelligent, precise information extraction directly from the documentary content.
M. El Ouaazizi· International Journal of Edu...· 0 citations
Legislative knowledge evolves as an intricate hypertext in which documents are interconnected through complex, often implicit relationships. In this paper, we introduce ReSB2, a framework for retrieving and linking similar legislative bills that supports human–machine collaboration and helps reduce redundancy in the lawmaking process. The framework fine-tunes two ModernBERT-based language models on authentic legislative data, incorporating domain-specific formatting and procedural constraints derived from real workflows in a Brazilian state-level legislative assembly. To ensure transparency, ReSB2 integrates an explainability module based on Integrated Gradients, enabling analysts to inspect which textual elements most influence model decisions. Evaluated on a large corpus of official bills, the framework outperforms both general-purpose and domain-specific baselines in identifying semantically similar documents, achieving recall values of approximately 0.9. Human-centric evaluation with domain experts further demonstrates that ReSB2 serves as an effective human-centered augmentation tool, supporting the consistency and governance of legislative knowledge.
Lucas G. L. Costa, Átila Souza, Elves Rodrigues et al.· Proceedings of the 37th ACM...· 0 citations
This paper outlines a unique method of legal text processing using Natural Language Processing (NLP) technology to extract the information from the legal texts meaningfully and naturally. The proposed system is designed in a data pipeline architecture by integrating the NLP functionalities such as tokenization, part-of-speech tagging, named entity recognition (NER), and dependency parsing to systematize the processing of typologies of legal text hubs, including legal briefs, statutes, and case law. The methodology presented concerns the importance of pre-processing legal texts that address domain-specific challenges. The texts may contain ambiguities, while the legal language itself is a very intricate kind of language. The system uses advanced methods like syntactic parsing and semantic role labeling to parse and find relevant entities, relationships, and context, ensuring the automation of the large amount of raw legal data for review and analysis. Besides that, the first is leveraging machine learning models to optimize the data extraction process and to ensure high efficiency and scalability. This methodology guarantees that accurate and reliable information is extracted and reduces the time and costs that conventionally come with manual legal analysis. The focal point of the offered system is overcoming legal workflow issues and bringing model texts to widespread use. Therefore, the proposed system aims to facilitate decision-making processes in legal practice and even the accuracy of the proposed model.
S. A. Gade, Sivaram Ponnusamy· Journal of Intelligent Decis...· 0 citations
Intelligent Target Locator (ITL), a domain-agnostic and language-portable methodology that estimates the affinity between the textual units of a target document and the concepts defined in a structured reference document, is presented.
R. Giráldez, Dayrelis Mena, Jesús S. Aguilar-Ruiz· 0 citations
Software systems must comply with legal regulations, which is a resource-intensive task, particularly for small organizations and startups lacking dedicated legal expertise. Extracting metadata from regulations to elicit legal requirements for software is a critical step to ensure compliance. However, it is a cumbersome task due to the length and complex nature of legal text. Although prior work has pursued automated methods for extracting structural and semantic metadata from legal text, they do not consider the interplay and interrelationships among attributes associated with these metadata types, and they rely on manual labeling or heuristic-driven machine learning, which does not always generalize to new documents. In this paper, we introduce a decomposition-based in-context learning method for automatically generating a canonical representation of legal text encoded as executable Python code. Our representation is instantiated from a manually designed Python class structure that serves as a domain-specific metamodel, capturing both structural and semantic legal metadata and their interrelationships. Our corpus contains 13 US state data breach notification laws (332 paragraphs), of which six unseen laws (182 paragraphs) form the held-out test set. On this test set, our proposed method using GPT−5.1 achieves 90.5% semantic test accuracy with a precision of 79.4% and a recall of 81.9%. We also assess the generalizability of the method to the Children’s Online Privacy Protection Act (COPPA), a US federal law. The results demonstrate that, once a domain metamodel and expert-authored examples are available, few-shot code generation can extract legal metadata relationships without training a task-specific supervised model and can be adapted to unseen legislation.
Anmol Singhal, Travis D. Breaux· Requirements Engineering· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.