Skip to content

Structured Data Extraction from Real Estate Documents using Clustering, Classification, and Large Language Models

Jul 2026 · arXiv.org · Vol abs/2607.06012 · 0 citations · 20 references
Computer Science

TL;DR

This paper presents an end-to-end pipeline for acquiring, classifying, and extracting structured data from selectable text documents and shows that the proposed extraction from each property document is both feasible and reliable at this scale.

Abstract

Real estate property listings expose structured metadata through the API. Still, the richest property-level information (i.e., legal status, structural condition, utility supplies, heating systems) sits in attached questionnaire documents that no automated system currently processes at scale. These documents are heterogeneous. Some are digitally generated with selectable text, others are scanned physical forms. There are even more complex layouts that contain checkbox annotations that defeat conventional text extraction. In this paper, we present an end-to-end pipeline for acquiring, classifying, and extracting structured data from selectable text documents. The pipeline was applied to 3965 questionnaire documents collected from a live property platform via reverse-engineered REST APIs. First, we classified each document into one of three structural categories (text_only, scanned, and special_char), then extracted 35 predefined property attributes from eligible documents using DeepSeek R1 as the Large Language Model, prompted to return a structured JSON object. All 2781 submitted documents were processed successfully, producing a final dataset of 2766 unique property records. Downstream validation confirmed the data quality. Cosine similarity matching achieves a Jaccard consistency score of 0.82, and K-Means clustering produces interpretable market segments with a silhouette score of 0.2088. Results show that the proposed extraction from each property document is both feasible and reliable at this scale.

View source

Similar papers

#large language models Open access Sep 2026

Research on Text Information Extraction and Imbalanced Classification Methods for Enterprise Profiling

A comprehensive natural language processing (NLP) pipeline for extracting key information and identifying industries, which outperforms traditional machine learning baselines and single deep learning models, offering more reliable recognition for minority classes.

Xin-Yi Xu · 0 citations
Preprint Aug 2026

Structure then Query: Enabling Precise Analytical Queries over Unstructured Documents

Experiments on three real-world datasets demonstrate that AnnoIndex consistently outperforms state-of-the-art baselines, achieving the highest average F1 score while maintaining robust performance on complex multi-hop join and progressive reasoning queries.

Teng Lin, Yuyu Luo, Nan Tang · 1 citation
2026

A Hybrid LLM Based Relationship Extraction Algorithm for Unstructured Data

Unstructured text data such as crime reports and witness statements often contain essential connections between crime entities such as suspects, weapons, locations, and crime types. However, it can be difficult to extract and analyze these linkages because they are often inserted within unstructured narratives. This pa...

Sukhvinder Kaur Walia, S. Masih, U. Suman · 0 citations
#artificial intelligence Preprint Sep 2026

TEAR: Table Extraction with Attribute Recommendation from Texts via Large Language Models

Table extraction from texts is an important task for information systems, and recent approaches that prompt large language models (LLMs) with instructions have drawn great attention for their strong performance. Existing works have assumed the input texts to be table descriptions or specialized documents. However, thes...

Tong Li, Shu-Ye Ding, Jia-Chuan Wang et al. · 0 citations
Open access Jul 2026

Automated Summarization Tool

The design realization and evaluation of an Automated Summarization Tool (AST) is presented which is a document intelligence platform based on google gemini 2.5 flash that outperforms the strongest fine-tuned transformer baselines (PEGASUS, BART) by ~14 points and is clearly ahead of BERTSUM-ext (a strong transformer b...

K. Kumar, A. Amandeep, Dharmender Kumar et al. · 0 citations
Open access Jul 2026

Preprocessing and Adaptive Parsing of Unstructured Voter Turnout Data with Spatial Feature Enrichment

The raw input data exhibited severe structural heterogeneity, manifesting as seven distinct structural formats (Formats A–G) characterized by inconsistent header placement, nested merged cells, variable string representations, and missing values. To address these challenges without relying on static rule-based scripts,...

Arya Pratama Tarigan, M. A. Budiman, Ade Candra · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.