Skip to content
Review Open access

Graph-Based Data Engineering Models for Large-Scale Knowledge Discovery

2025 · International Journal of Data Engineering and Intelligent Computing · Vol 8, pp. 01-16 · 0 citations

TL;DR

Experimental evaluation demonstrates improved relationship discovery, query performance, and knowledge extraction compared with conventional relational approaches, making the proposed framework suitable for intelligent applications in healthcare, cybersecurity, finance, smart manufacturing, and enterprise knowledge management.

Abstract

Graph-based data engineering has become a powerful approach for managing and analyzing highly interconnected data across enterprise systems, IoT, social media, healthcare, finance, and scientific domains. Unlike traditional relational databases, graph-based models represent data as interconnected nodes and edges, enabling efficient relationship analysis, semantic understanding, and knowledge discovery. This paper surveys recent advances in graph databases, knowledge graphs, graph neural networks (GNNs), and distributed graph analytics, and proposes an integrated framework for scalable graph construction, semantic enrichment, graph analytics, and AI-driven knowledge extraction. The framework emphasizes scalability, semantic consistency, explainable AI, and continuous graph evolution. Experimental evaluation demonstrates improved relationship discovery, query performance, and knowledge extraction compared with conventional relational approaches, making the proposed framework suitable for intelligent applications in healthcare, cybersecurity, finance, smart manufacturing, and enterprise knowledge management.

Read PDF

Similar papers

Open access Aug 2026

Knowledge Graph–Driven Enterprise Data Integration for Autonomous Decision Intelligence

Modern enterprises generate vast amounts of data from diverse sources, including business applications, cloud platforms, IoT devices, social networks, and transactional systems. Integrating and analyzing this heterogeneous data efficiently remains a significant challenge due to data silos, semantic inconsistencies, and complex relationships among entities. Knowledge Graphs (KGs) have emerged as a powerful technology for representing interconnected enterprise data through semantic relationships, enabling enhanced data integration, contextual understanding, and intelligent knowledge discovery. This paper presents a Knowledge Graph–Driven Enterprise Data Integration Framework for Autonomous Decision Intelligence that unifies heterogeneous data sources into a semantically enriched knowledge ecosystem. The proposed framework employs ontology modeling, entity resolution, semantic mapping, graph construction, and intelligent reasoning mechanisms to establish meaningful relationships among enterprise data assets. Advanced graph analytics and machine learning techniques are integrated to support autonomous decision-making by generating contextual insights, identifying hidden patterns, and providing real-time recommendations. The framework further incorporates automated data governance, metadata management, and explainable reasoning capabilities to ensure data quality, transparency, and regulatory compliance. Experimental evaluation demonstrates that the proposed approach significantly improves data integration accuracy, knowledge discovery efficiency, and decision intelligence performance compared with traditional data integration systems. By leveraging knowledge graphs and intelligent reasoning engines, the framework enables organizations to transform fragmented enterprise data into actionable knowledge, thereby enhancing operational efficiency, strategic planning, and autonomous business decision-making. The proposed solution provides a scalable and intelligent foundation for next-generation enterprise analytics and AIdriven decision support systems.

Shashank Akinapalli · 0 citations
Open access 2024

AI-Based Knowledge Graphs for Intelligent Decision Support

Experimental results show that AI-driven knowledge graphs significantly enhance decision accuracy, reduce ambiguity, and improve interpretability, achieving up to 85–92% higher decision efficiency compared to traditional methods.

Venkatesh Iyer, Nandhini Ravi · 0 citations
Open access 2020

Knowledge Graphs for Intelligent Data Integration Systems

The increasing presence of heterogeneous data sources in modern information systems has intensified the need for intelligent data integration processes capable of handling semantic complexity, structural diversity, and dynamic changes. Traditional data integration methods, primarily based on relational schemas and syntactic mappings, struggle to address semantic heterogeneity in large-scale distributed environments. Knowledge graphs have emerged as a powerful paradigm, enabling semantically rich, flexible, and scalable integration by representing data as interconnected entities with metadata, ontologies, and inference capabilities. Using technologies such as RDF and OWL, knowledge graphs support interoperability, contextual reasoning, and unified data views across systems. This paper examines knowledge graph-based intelligent data integration systems, focusing on their architecture, methodology, and practical applications. It highlights their advantages in schema alignment, entity resolution, and semantic enrichment over traditional ETL approaches. The integration of machine learning techniques further enhances automation in data mapping, anomaly detection, and knowledge discovery. A systematic framework is proposed, covering ontology design, data ingestion, graph construction, and query optimization. A conceptual case study demonstrates improved integration accuracy, scalability, and query performance. Evaluation results indicate enhanced data quality, interoperability, and reasoning capabilities, along with reduced integration latency. Overall, knowledge graphs serve as a key enabler for next-generation intelligent data integration, supporting complex relationships and data-driven decision-making. Future work includes improving scalability, real-time processing, and integration with deep learning models.

Muhammad Al-Azar · 0 citations
Open access 2022

Graph Database Pipeline Integration for Dynamic Data Analytics

The integration of graph databases into dynamic data analytics pipelines presents a promising solution for efficiently processing and analyzing complex, interconnected data. As organizations generate and consume increasing amounts of data, traditional data models struggle to keep up with the demands for real-time insights and dynamic data processing. Graph databases, with their ability to represent relationships between entities in a highly flexible structure, offer significant advantages in such scenarios. This paper explores the design, implementation, and challenges of integrating graph databases into modern data pipeline architectures. It discusses how graph models can be dynamically updated, analyzed, and visualized in real-time to unlock advanced analytics capabilities. We also explore use cases from industries such as social networks, fraud detection, and IoT, demonstrating the value of graph databases in handling dynamic and evolving datasets. Finally, we identify key challenges and future research opportunities in the field, emphasizing the need for scalability, performance optimization, and real-time analytics.

Michael Rabin, Amir Pnueli · 0 citations
Open access Aug 2026

A Comparative Analysis of Spark GraphX and GraphFrames for Healthcare Data Analytics

Healthcare data analytics is essential for identifying disease patterns, understanding patient relationships, and supporting data-driven decision-making in modern healthcare systems. As healthcare datasets continue to grow in size and complexity, scalable graph-processing frameworks have become increasingly important for analyzing interconnected patient data. This paper compares two Apache Spark graph-processing frameworks, GraphX and GraphFrames, using the Centers for Disease Control and Prevention Diabetes Health Indicators dataset. A patient similarity graph was constructed by representing 70,692 patients as vertices and connecting highly similar patients through cosine–similarity relationships, resulting in 16,009,101 graph edges. The two frameworks were evaluated using the same graph structure and execution environment with respect to graph construction time, execution performance, memory consumption, application programming interface usability, and scalability. In addition to the framework comparison, graph-based features, including in-degree, out-degree, total degree, and PageRank, were extracted to examine the structure of the patient similarity network. The experimental results showed that GraphX completed graph construction, PageRank computation, and degree calculations considerably faster than GraphFrames. However, GraphFrames offered a higher-level programming interface, simpler integration with Spark SQL, and easier implementation of graph analytics workflows. The extracted graph measures also highlighted highly connected and structurally important patients within the network, demonstrating the usefulness of graph-based feature extraction for healthcare data analysis. GraphX completed graph construction approximately 33 times faster (92.5 s vs. 3037 s), PageRank computation approximately 780 times faster (9.1 s vs. 7132 s), and degree computation over 10,000 times faster (1.0 s vs. 10202 s) than GraphFrames, whereas GraphFrames used substantially less memory; these results indicate that GraphX is more suitable for performance-oriented graph workloads, whereas GraphFrames is advantageous when development flexibility and DataFrame integration are primary considerations.

Soz Raouf Hama, A. Jumaa · 0 citations
Open access 2025

Graph Foundation Models for Cross-Domain Knowledge Integration and Analytics

The proposed framework provides a scalable foundation for graph-based artificial intelligence and has applications in biomedical knowledge discovery, financial fraud detection, industrial digital twins, recommendation systems, cybersecurity intelligence, scientific literature mining, and smart governance.

Seppo Linnainmaa, A. Salomaa · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.