Skip to content
Open access

Enhancing Stock Price Prediction through Artificial Intelligence-Driven Integration of Market Data and Sentiment Analysis

Jul 2026 · International Journal of Economics and Management Studies · 0 citations · 29 references

Abstract

Forecasting equity prices is hard, and the literature has been telling itself this for fifty years. Markets are volatile, nonlinear, and shaped by investor psychology in ways that Autoregressive Integrated Moving Average (ARIMA) and Generalised Autoregressive Conditional Heteroskedasticity (GARCH) were never designed to capture. The deep-learning literature has been pushing back on this for the last decade, but the typical study tests one or two networks against a single baseline, draws sentiment from a single platform, and reports on a handful of stocks within one exchange. A head-to-head benchmark that holds preprocessing, validation, and sentiment integration constant across many models and many markets is conspicuously absent. This paper supplies one. Ten forecasters were trained and tested on five years of daily prices for 25 technology equities listed across NYSE, NASDAQ, the Mexican Stock Exchange, Shanghai, and Korea: ARIMA, Exponential Smoothing State Space (ETS), GARCH, Random Forest Regression (RFR), Extreme Gradient Boosting (XGBoost), Support Vector Machine (SVM), Prophet, Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), and an ARIMA-LSTM hybrid. A parallel sentiment pipeline was built around 1,849 news articles drawn from NewsAPI.org, with two vocabulary sizes (5,000 and 3,000 tokens) feeding both LSTM and Random-Forest classifiers. Dropout (rate 0.2), early stopping (patience 10), and walk-forward validation were applied to the recurrent models so that accuracy gains could be attributed to the architecture rather than to a forgiving data split. LSTM achieved the lowest mean Root Mean Square Error (RMSE) at 77.29 and Mean Absolute Percentage Error (MAPE) of 2.58%; GRU was a close second at RMSE 66.03 and MAPE 2.49%. ARIMA produced a mean MAPE of 53.01% and Prophet 33.87%. Diebold–Mariano testing confirms the LSTM advantage over both at the 1% level. Negative-sentiment recall, however, remained weak — a finding with direct consequences for any production system that needs news polarity to anticipate downside risk.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.