Skip to content
Open access

Design of an Iterative Model for Contextual Structuring of Unstructured NoSQL Data Using Hybrid Topological, Probabilistic, and Relational Learning Models

Aug 2026 · Journal of Intelligent Decision Making and Information Science · 0 citations · 24 references

TL;DR

The entire process of structuring unstructured NoSQL data forms a fully automated method, semantically grounded and performance-aware, towards compatibility with SQL-ready, graph analytics, and downstream workflows in machine learning process.

Abstract

Unstructured exponential growth of NoSQL data is causing an effective challenge for data processing in general, semantic querying, and integration into the traditional analytics pipeline for efficient data processing. Current solutions depend at least in part on rule-based static transformation or partial machine learning models, often lacking any preservation of semantic integrity, schema heterogeneity, or alignment with real-world usage patterns. These constraints severely cripple downstream applications like query optimization, relational mapping, and ML pipeline integration in process. This work proposes a thorough Contextual Structuring Pipeline that automatically converts unstructured NoSQL data into structured, schema-consistent, and optimized relational representations in the queries. The pipeline includes five new learning models designed to contribute to a specific subtask in the transformation. First, CADA-Net (Context-Aware Document Attention Network) uses transformer-based hierarchical encoding to segment raw NoSQL records into semantically meaningful key Value structures. Second, TopoGraph-X constructs an entity-type graph with multi-level topologies from latent document hierarchies through topological learning. Third, HarmoField employs domain-specific BERT embeddings along with the Gaussian mixture model to normalize field representation across heterogeneous sources. Fourth, MetaRel-Frame constructs meta-relational abstractions on logical entity-role patterns to discover relational table blueprints. Finally, QueryStruct transforms the blueprint into a use-optimized dataset through restructuring, indexing, reordering, and compressing its components, taking into account historic query logs. The structural accuracy achieved by this framework is 92.3%, schema normalized is 91.2%, and query latency is decreased by 39.4%. The entire process of structuring unstructured NoSQL data forms a fully automated method, semantically grounded and performance-aware, towards compatibility with SQL-ready, graph analytics, and downstream workflows in machine learning process.

Read PDF

Similar papers

Preprint Aug 2026

Structure then Query: Enabling Precise Analytical Queries over Unstructured Documents

Experiments on three real-world datasets demonstrate that AnnoIndex consistently outperforms state-of-the-art baselines, achieving the highest average F1 score while maintaining robust performance on complex multi-hop join and progressive reasoning queries.

Teng Lin, Yuyu Luo, Nan Tang · 1 citation
Conference Open access Aug 2026

AI-Driven Knowledge Externalisation: From Unstructured Documents to Structured Data Models

The findings suggest that AI-based structured extraction may redefine how organisations formalise expertise, shifting from document-centric storage toward schema-driven knowledge architectures.

Dilyan Georgiev, E. Gourova · 0 citations

ExpeSQL: An Efficient, Experience-Guided Decompositional Search Framework for Text-to-SQL

This work introduces ExpeSQL, a zero-shot, open-source–compatible, and efficient framework that combines divide-and-conquer reasoning, Best-of-N candidate selection, and self-critique with experience-guided refinement that establishes a new paradigm for deployable, self-improving Text-to-SQL systems in dynamic, real-wo...

Unknown authors · 0 citations
Open access Jul 2026

Schema-Aware Query Translation and Tabular Reasoning for Enterprise Databases

Schema-Aware Query Translation and Tabular Reasoning for Enterprise Databases aka Inference-from-RDBMS is presented, an open-source framework designed for schema-aware query translation, dynamic context pruning, and execution-guided tabular inference over complex RDBMS structures.

Harshil Lodhiya · 0 citations
#natural language process... Preprint Aug 2026

RENSA: Rich Environment Metadata to Navigate Shared and Distributed Endpoints for Automated Federated SPARQL Query Generation

This work proposes RENSA, a federated SPARQL query generation framework that leverages an extension of SPARQL Builder Metadata (SBM), and demonstrates that RENSA infers class and authority constraints for query variables, enabling the identification of data sources even across heterogeneous endpoints.

Victor Eiti Yamamoto, Hideaki Takeda, Yasunori Yamamoto · 0 citations
Preprint Aug 2026

Guided Table Retrieval for Structured Data Search

guided table retrieval is presented, a four-phase pipeline that combines deterministic grounding via hash-based predictors, structural exploration of join-graph reachability, LLM-powered disambiguation of sources and targets, and algorithmic merging into minimal, topologically ordered join trees.

Alekh Jindal, J. Pandey, C. Pavlopoulou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.