This paper presents an end-to-end pipeline for acquiring, classifying, and extracting structured data from selectable text documents and shows that the proposed extraction from each property document is both feasible and reliable at this scale.
Abstract
Real estate property listings expose structured metadata through the API. Still, the richest property-level information (i.e., legal status, structural condition, utility supplies, heating systems) sits in attached questionnaire documents that no automated system currently processes at scale. These documents are heterogeneous. Some are digitally generated with selectable text, others are scanned physical forms. There are even more complex layouts that contain checkbox annotations that defeat conventional text extraction. In this paper, we present an end-to-end pipeline for acquiring, classifying, and extracting structured data from selectable text documents. The pipeline was applied to 3965 questionnaire documents collected from a live property platform via reverse-engineered REST APIs. First, we classified each document into one of three structural categories (text_only, scanned, and special_char), then extracted 35 predefined property attributes from eligible documents using DeepSeek R1 as the Large Language Model, prompted to return a structured JSON object. All 2781 submitted documents were processed successfully, producing a final dataset of 2766 unique property records. Downstream validation confirmed the data quality. Cosine similarity matching achieves a Jaccard consistency score of 0.82, and K-Means clustering produces interpretable market segments with a silhouette score of 0.2088. Results show that the proposed extraction from each property document is both feasible and reliable at this scale.
A comprehensive natural language processing (NLP) pipeline for extracting key information and identifying industries, which outperforms traditional machine learning baselines and single deep learning models, offering more reliable recognition for minority classes.
Xin-Yi Xu· Applied and Computational En...· 0 citations
Experiments on three real-world datasets demonstrate that AnnoIndex consistently outperforms state-of-the-art baselines, achieving the highest average F1 score while maintaining robust performance on complex multi-hop join and progressive reasoning queries.
Unstructured text data such as crime reports and witness statements often contain essential connections between crime entities such as suspects, weapons, locations, and crime types. However, it can be difficult to extract and analyze these linkages because they are often inserted within unstructured narratives. This pa...
Sukhvinder Kaur Walia, S. Masih, U. Suman· International Journal of Lat...· 0 citations
Table extraction from texts is an important task for information systems, and recent approaches that prompt large language models (LLMs) with instructions have drawn great attention for their strong performance. Existing works have assumed the input texts to be table descriptions or specialized documents. However, thes...
Tong Li, Shu-Ye Ding, Jia-Chuan Wang et al.· 0 citations
The design realization and evaluation of an Automated Summarization Tool (AST) is presented which is a document intelligence platform based on google gemini 2.5 flash that outperforms the strongest fine-tuned transformer baselines (PEGASUS, BART) by ~14 points and is clearly ahead of BERTSUM-ext (a strong transformer b...
K. Kumar, A. Amandeep, Dharmender Kumar et al.· International Journal of Inn...· 0 citations
The raw input data exhibited severe structural heterogeneity, manifesting as seven distinct structural formats (Formats A–G) characterized by inconsistent header placement, nested merged cells, variable string representations, and missing values. To address these challenges without relying on static rule-based scripts,...
Arya Pratama Tarigan, M. A. Budiman, Ade Candra· Green Intelligent Systems an...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.