Skip to content
Open access

Efficient Storage and Query Optimization for Large-Scale Data Sets in Distributed Architectures

Unknown authors
Aug 2026 · Advanced Electromagnetics · 0 citations

Abstract

With the rapid development of the big data industry, data volume across various industries has exploded, and large-scale datasets at PB and EB levels have become mainstream objects for data processing. Relying on core theories of distributed storage and query, this paper constructs an integrated collaborative optimization system covering storage, query and caching. Storage efficiency is improved by designing an adaptive dynamic sharding strategy and intelligent multi-replica placement policy, building a tiered storage architecture for hot and cold data, and optimizing storage encoding and compression mechanisms. Query overhead is reduced by reconstructing query execution plans based on cost models, pushing down operators, and optimizing cross-node transmission. A collaborative global-local indexing framework and a multi-level caching linkage mechanism are established to further boost query response performance. A standardized distributed experimental cluster is deployed, and multi-dimensional comparative experiments are conducted to verify the performance of the proposed optimization scheme. Experimental results demonstrate that the proposed optimization solution can effectively cut storage redundancy overhead, improve cluster resource utilization, drastically reduce query latency for large-scale data, and raise concurrent throughput. It can provide theoretical support and engineering practice references for distributed system optimization in industrial big data, internet massive data processing and other scenarios.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.