Intelligent Target Locator (ITL), a domain-agnostic and language-portable methodology that estimates the affinity between the textual units of a target document and the concepts defined in a structured reference document, is presented.
Abstract
Measuring alignment between documents and structured reference frameworks requires identifying conceptual evidence distributed throughout the text and reporting it through measures that are quantitative, interpretable, and traceable. Many commonly used retrieval and classification approaches return either pairwise similarity scores or one or more class labels, whereas fewer methods provide concept-level scores that are directly traceable to the terminological evidence supporting them. We present \emph{Intelligent Target Locator} (ITL), a domain-agnostic and language-portable methodology that estimates the affinity between the textual units of a target document and the concepts defined in a \emph{Structured Reference Document} ($SRD$). From the $SRD$, ITL induces concept-specific terminological profiles built from independent terms, bigrams, trigrams, and co-occurrences. Each term is assigned an importance weight that combines concept membership, term-type specificity and inter-concept discriminability. The output is a textual-unit--concept affinity matrix that can be aggregated at different levels of granularity. We conduct an internal consistency assessment using the 17 Sustainable Development Goals (SDGs), evaluating each official goal statement against the $SRD$ induced from the same set of descriptors. Every statement reached its highest affinity with the corresponding concept, and the mean affinity across the remaining concepts stayed marginal relative to the mean reference affinity. This separation indicates that ITL distinguishes the conceptual profiles of the framework. ITL thus offers a general basis for quantifying document alignment with structured frameworks while keeping each result traceable to the terminological evidence that supports it.
Long-document question-answering experiments show that human-verified TOC hierarchies and contextual relationships improve reasoning, with their combination providing complementary benefits.
Yuefeng Zou, Yichen Lu, Jingxiao Yang et al.· 0 citations
Research teams and organizations often explore unfamiliar free-text collections, from survey comments and reviews to reports and domain documents, before labels, queries or coding schemes exist. At this stage, the first thematic map shapes what users notice, prioritize and carry into downstream analysis, so it should be trusted only insofar as it can be verified. Existing options force a trade-off between scale and verifiability. Qualitative coding preserves evidence but is slow. Search presupposes a query. Clustering and topic models scale but produce labels users must interpret. One-shot large language model (LLM) summaries are fluent yet difficult to reproduce or audit. We present EviMap, an interactive system providing researchers and practitioners with an auditable thematic overview of such corpora. Guided by model-generated context describing the corpus and hypothesized stakeholder concerns, EviMap extracts within-document evidence phrases and organizes them, rather than whole documents, into a three-level map of aspects, groups and fine-grained topics. Embedding-based clustering narrows the search space for finer semantic judgments by the LLM. Each node traces back to supporting phrase spans, so documents link to topics through evidence they contain and users can audit labels against the original text. Users can start from a top-level corpus map, drill into topics, inspect highlighted evidence in original documents, and combine two topics to find documents discussing both. We demonstrate this workflow across six heterogeneous corpora spanning 2,108 to 101,699 documents, with a comparison against flat and hierarchical LLM baselines. By grounding every label in verbatim source spans, EviMap makes a topic map not just readable, but verifiable. Code, demo video, and interactive dashboard are available at https://github.com/zhiyintan/EviMap.
Legislative knowledge evolves as an intricate hypertext in which documents are interconnected through complex, often implicit relationships. In this paper, we introduce ReSB2, a framework for retrieving and linking similar legislative bills that supports human–machine collaboration and helps reduce redundancy in the lawmaking process. The framework fine-tunes two ModernBERT-based language models on authentic legislative data, incorporating domain-specific formatting and procedural constraints derived from real workflows in a Brazilian state-level legislative assembly. To ensure transparency, ReSB2 integrates an explainability module based on Integrated Gradients, enabling analysts to inspect which textual elements most influence model decisions. Evaluated on a large corpus of official bills, the framework outperforms both general-purpose and domain-specific baselines in identifying semantically similar documents, achieving recall values of approximately 0.9. Human-centric evaluation with domain experts further demonstrates that ReSB2 serves as an effective human-centered augmentation tool, supporting the consistency and governance of legislative knowledge.
Lucas G. L. Costa, Átila Souza, Elves Rodrigues et al.· Proceedings of the 37th ACM...· 0 citations
Large Language Models are deployed in financial applications such as research synthesis and risk analysis, yet their effectiveness is constrained by the limitations of conventional retrieval methods. Existing approaches rely primarily on semantic similarity or token-level matching, which fails in structured domains like finance where relevance depends on precise alignment across entity, temporal and document-type dimensions. This paper proposes the Financial Knowledge Integration Framework (FKIF), a hybrid metadata-aware retrieval system that integrates dense semantic similarity, sparse lexical matching, and structured metadata signals into a unified ranking function. Unlike conventional hybrid retrieval approaches, FKIF treats metadata as a first-class relevance signal rather than an auxiliary feature. Tested on a held-out set of 22 queries drawn from a corpus of 57 SEC filings, FKIF achieves an MRR@10 of 0.98 against TF-IDF’s 0.35, BM25’s 0.17, and dense retrieval’s 0.13. The results demonstrate that metadata-aware retrieval significantly enhances retrieval accuracy and provides a foundation for reliable financial Retrieval-Augmented Generation (RAG) systems.
S. Jambula, Srihari Kumar Pendyala, Rajesh Kumar Butteddi et al.· 2026 International Conferenc...· 0 citations
Scientific papers require models to reason jointly over text, equations, figures, tables, code, and datasets while preserving the provenance of supporting evidence. Existing benchmarks typically evaluate these capabilities in isolation, leaving unclear whether multimodal models can support realistic scientific-reading workflows. We introduce SciDocBench, a workflow-centered benchmark for scientific document understanding. It contains 124 expert-authored and difficulty-screened questions organized into seven research-assistant capability groups and 19 subtasks across five scientific domains. Each question is instantiated under four matched conditions combining English or Chinese questions with all-images-first or interleaved document representations, yielding 496 evaluation instances for controlled analysis. The strongest evaluated system achieves only 62.6/100, with pronounced weaknesses in document perception, evidence grounding, verification, and cross-document reasoning. To translate these diagnostics into scalable training signals, we introduce SciDocIR, a typed evidence-graph representation that preserves scientific document objects, layout and cross-reference relations, and provenance. Building on SciDocIR, we construct SciDocDataset, comprising approximately 15K supervised fine-tuning samples and 8K reinforcement-learning samples across 14 verifiable subtasks. Together, SciDocBench, SciDocIR, and SciDocDataset form an evaluation-to-training framework for diagnosing and improving scientific-document assistants. The project page is available at https://github.com/InternLM/SciDocBench.
Shenxi Wu, Yu-Hong Liu, Haosong Zhang et al.· 0 citations
LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost, and is the first to score value accuracy, record completeness at scale, grounding, and measured cost together.
Boyang Zhang, Adrian Lyjak, Elizabeth Stewart et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.