May 2026· International Convention on Information and Communication Technology, Electronics and Microelectronics· pp. 1001-1005· 1 citation· 14 references
TL;DR
The results indicate that the proposed approach reduces the effort required for manual dataset construction while preserving data quality through mandatory human validation, and highlights the effectiveness of hybrid automation workflows in accelerating fine-tuning dataset preparation without compromising reliability.
Abstract
Preparing fine-tuning datasets for large language models (LLMs) commonly involves substantial manual effort, particularly in extracting, structuring, and validating data from unstructured sources. This study proposes a semi-automated, human-in-the-loop approach for generating fine-tuning question–answer (QA) pairs from PDF documents. The research investigates how unstructured textual content can be systematically transformed into validated QA data suitable for fine-tuning, while mitigating the risks associated with hallucinated or low-quality model outputs.The proposed system consists of a web-based architecture combining a React frontend with a Flask backend interfacing with the OpenAI API. Users provide a PDF document and a target page range, after which the system extracts text and generates candidate QA pairs. These candidates are presented for manual inspection, filtering, and refinement, prior to export in a structured JSON format compatible with fine-tuning pipelines.The results indicate that the proposed approach reduces the effort required for manual dataset construction while preserving data quality through mandatory human validation. The study highlights the effectiveness of hybrid automation workflows in accelerating fine-tuning dataset preparation without compromising reliability, and contributes design insights for human-centered tools supporting LLM customization.
It is asserted that the present contribution consists of an interpretable domain palette, a constructed benchmark of diverse tabular datasets, and reproducible code and data to enable further research on domain discovery and domain-aware tooling for tabular data.
This work proposes a two-step validation method that integrates a PLM block into the generative IE pipeline and thereby leverages LLMs' correction capability, discovering that such a validation task enhances LLM performance, particularly on the extraction of weakly expressed, low-salience entities that appear sparsely throughout the text.
Yi-Sheng Hsu, Nermeen Abou Baker, U. Handmann· arXiv.org· 0 citations
This study explores a semantic variation methodology to augment training data by generating question-answer pairs with explicit control over semantic similarity, and shows that semantically controlled augmentation improves domain-specific knowledge acquisition while preserving consistency.
Alexander Chen, Caroline Tang, Jennifer Sleeman· TEXT2KG/BiKE@ESWC· 0 citations
A multimodal core-opinion extraction framework in which visual evidence serves as a contextual anchor for textual judgment is proposed, providing a case-level value signal for downstream STI screening.
Sheng Hong, Xuan-Qi Wang, Jiachen Wang et al.· 0 citations
Evidence of cross-lingual efficacy of code-based LLMs for Chinese QA tasks, further enhanced through Code Llama-M's expanded Chinese vocabulary is found, and successful application of the fine-tuned LLM in a live assistant system, enhancing user experience is demonstrated.
Jiajun Yu, Linghan Zheng, Hui Liu et al.· Annual International ACM SIG...· 0 citations
A preliminary version of a framework that improves the most widespread DLA datasets quality by assessing and correcting layout coherence in scholarly documents and provides a more reliable ground truth with improved structural and semantic coherence for training and evaluating document segmentation and understanding models.
L. Massai, S. Marinai· 1 citation
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 18, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, human-editable specifications. The post Flint: A visualization language for the AI era appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduJun 3, 2026
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.