Skip to content
Conference Open access

A Review of Distributed Computing Technologies for Financial Data Preprocessing

Sep 2026 · Exploring Science Academic Conference Series · 0 citations · 10 references

Abstract

Financial data preprocessing is a key link in financial analysis and modeling. With the exponential growth of data scale, the single-machine architecture is facing severe bottlenecks. Distributed computing provides a feasible path to break through performa nce limitations through multi-node collaborative processing. This article systematically sorts out the research status and technical progress of distributed computing in the field of financial data preprocessing. First of all, analyze the common characteri stics and pre-processing task genealogy of financial data, and combine Tang Yao’s stock linkage effect research to show the typical process of single-machine preprocessing pipelines and its scale bottlenecks; Then review the mainstream frameworks such as Hadoop, Spark and Flink from the perspective of technological evolution, and summ arize the three core empowerment mechanisms of data parallelism, computing parallelism and stream processing; On this basis, classify and summarize the application research of distributed computing in data cleaning, feature engineering, real-time processing and other links, and take Ma Chiyu’s financial news sentiment analysis based on SparkR as a case to verify the performance advantages of distributed preprocessing in actual tasks; Finally, we will comment on the limitations of existing research and look forward to the future direction of ad aptive preprocessing, explainable attribution, privacy protection calculation, etc. This article aims to provide a systematic reference for distributed computing applications in the field of financial data preprocessing.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.