Skip to content
Open access

OmniExtract: an automatic data extraction tool based on large language model and prompt engineering

Sep 2025 · bioRxiv · Vol 27 · 0 citations · 29 references
Biology Medicine

TL;DR

OmniExtract is an automatic data extraction tool with user-friendly configuration files which can adapt to various data extraction tasks and it can support a comprehensive data extraction including text and tables.

Abstract

Extracting structured information from documents or scientific papers is crucial for data sharing and retrieval. Recently, Large Language Model (LLM) has shown its impressive ability in text understanding and several tools based on LLM has been developed. However, it’s still difficult to find a universal and user-friendly tool for various practical extraction tasks. To address this challenge, we propose OmniExtract, an automatic data extraction tool with user-friendly configuration files which can adapt to various data extraction tasks. OmniExtract uses a prompt optimized engineering to improve prompt and obtain high performance, and it can support a comprehensive data extraction including text and tables. Evaluation results show that OmniExtract obtains a high accuracy over 80% for 3 datasets. Furthermore, two additional data extraction applications using OmniExtract have been provided, achieving an accuracy of 92.21% and an average F1 score of 0.83 respectively. The data reliability performance shows that OmniExtract is a valuable tool for database updating.

Read PDF

Similar papers

Preprint Aug 2026

Structure then Query: Enabling Precise Analytical Queries over Unstructured Documents

Experiments on three real-world datasets demonstrate that AnnoIndex consistently outperforms state-of-the-art baselines, achieving the highest average F1 score while maintaining robust performance on complex multi-hop join and progressive reasoning queries.

Teng Lin, Yu-Yu Luo, Nan Tang · 2 citations
Conference

Exploring Large Language Models as Decision Support Tools: A Proof of Concept for Procedure-Related Information Retrieval

Policy and procedural documentation are essential for the effective operation of government agencies. However, the vast volume of these documents often obscures critical information within hundreds of pages of unrelated content. Recent advancements in large language models (LLMs) enhance text search and response genera...

Nathaniel Shepherd, Thomas Berg, Xue-Ping Li · 0 citations
Preprint Sep 2026

Towards Semi-Automatically Comparing Keyword-Based and Semantic Search Accuracy

The increasing importance of Information Retrieval (IR) in managing large datasets has highlighted significant limitations in traditional keyword-based search systems. Context-aware chat-based search methods, such as Retrieval Augmented Generation (RAG), have recently emerged, but their evaluation compared to keyword-b...

Mohamed Ben Salha, Fiete Lüer, M. Betka et al. · 0 citations
Aug 2026

Optimizing information retrieval tasks with large language model for data enhancement

The goal is to optimize the normalized discounted cumulative gain (NDCG) metric, which measures the ranking quality of the retrieved documents, by integrating the fine-tuned LLM model with the chosen information retrieval system by incorporating the model’s outputs into ranking algorithms.

C. Vaidya, Amudhavel Jayavel, Pradeep Kumar Mishra et al. · 0 citations
Book Open access Mar 2026

VILLA: Versatile Information Retrieval from Scientific Literature Using Large Language Models

This study designs a unique, open-ended SIE task of extracting mutations in a given virus that modify its interaction with the host, and develops a new, multi-step retrieval augmented generation (RAG) framework called VILLA for SIE.

Blessy Antony, Amartya Dutta, Sneha Aggarwal et al. · 0 citations
Open access Aug 2026

LLM-based design document understanding for semantic knowledge extraction and design verification

Design documents contain essential design knowledge such as designers’ intent, decision-making criteria, and constraints, and are widely used to support accurate and consistent product development. Most design documents are extensive and composed of unstructured natural language-based text, which makes it difficult f...

Junho Kim, Sangwook Park, Seungeun Lim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.