OmniExtract is an automatic data extraction tool with user-friendly configuration files which can adapt to various data extraction tasks and it can support a comprehensive data extraction including text and tables.
Abstract
Extracting structured information from documents or scientific papers is crucial for data sharing and retrieval. Recently, Large Language Model (LLM) has shown its impressive ability in text understanding and several tools based on LLM has been developed. However, it’s still difficult to find a universal and user-friendly tool for various practical extraction tasks. To address this challenge, we propose OmniExtract, an automatic data extraction tool with user-friendly configuration files which can adapt to various data extraction tasks. OmniExtract uses a prompt optimized engineering to improve prompt and obtain high performance, and it can support a comprehensive data extraction including text and tables. Evaluation results show that OmniExtract obtains a high accuracy over 80% for 3 datasets. Furthermore, two additional data extraction applications using OmniExtract have been provided, achieving an accuracy of 92.21% and an average F1 score of 0.83 respectively. The data reliability performance shows that OmniExtract is a valuable tool for database updating.
Experiments on three real-world datasets demonstrate that AnnoIndex consistently outperforms state-of-the-art baselines, achieving the highest average F1 score while maintaining robust performance on complex multi-hop join and progressive reasoning queries.
Policy and procedural documentation are essential for the effective operation of government agencies. However, the vast volume of these documents often obscures critical information within hundreds of pages of unrelated content. Recent advancements in large language models (LLMs) enhance text search and response genera...
The increasing importance of Information Retrieval (IR) in managing large datasets has highlighted significant limitations in traditional keyword-based search systems. Context-aware chat-based search methods, such as Retrieval Augmented Generation (RAG), have recently emerged, but their evaluation compared to keyword-b...
Mohamed Ben Salha, Fiete Lüer, M. Betka et al.· 0 citations
The goal is to optimize the normalized discounted cumulative gain (NDCG) metric, which measures the ranking quality of the retrieved documents, by integrating the fine-tuned LLM model with the chosen information retrieval system by incorporating the model’s outputs into ranking algorithms.
C. Vaidya, Amudhavel Jayavel, Pradeep Kumar Mishra et al.· Journal of Nonlinear, Comple...· 0 citations
This study designs a unique, open-ended SIE task of extracting mutations in a given virus that modify its interaction with the host, and develops a new, multi-step retrieval augmented generation (RAG) framework called VILLA for SIE.
Blessy Antony, Amartya Dutta, Sneha Aggarwal et al.· Proceedings of the 32nd ACM...· 0 citations
Design documents contain essential design knowledge such as designers’ intent, decision-making criteria, and constraints, and are widely used to support accurate and consistent product development. Most design documents are extensive and composed of unstructured natural language-based text, which makes it difficult f...
Junho Kim, Sangwook Park, Seungeun Lim et al.· Journal of Computational Des...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.