Skip to content
Conference

Design of a Distributed Semantic Analysis System for Scalable Processing of Heterogeneous Text Data

Jul 2026 · International Conference on Big Data Computing Service and Applications · pp. 248-252 · 0 citations · 21 references

Abstract

This study addresses the challenge of achieving consistent and scalable quantification of semantic divergence in large-scale heterogeneous textual data by proposing an integrated solution that combines methodological design and system architecture. Existing approaches primarily rely on lexical statistics or vector-based distances, which are insufficient for capturing the global structure of high-dimensional semantic spaces and lack consistency across different text granularities. To overcome these limitations, we propose the Core Semantic Variance Index (CSVI), which leverages semantic embeddings together with Principal Component Analysis (PCA) and the Participation Ratio (PR) to characterize the distribution of semantic variance. In addition, a distributed semantic analysis service platform is developed to support unified analysis and efficient computation across diverse text structures. Comparative experiments against traditional lexical-based and existing embedding-based methods show that CSVI achieves accuracies of 86.35%, 93.05%, and 88.32% on Multilingual-STSB, Python Dialogue, and 20 Newsgroups, respectively, demonstrating its effectiveness and applicability within distributed service environments.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.