Jul 2026· NLP & Big Data· pp. 15-24· 0 citations· 13 references
TL;DR
It is suggested that LLMs interpret financial context rather than merely counting sentiment-bearing words, offering a meaningful advance for early-warning and financial risk monitoring applications.
Abstract
This paper evaluates whether large language model (LLM)-based sentiment analysis can detect financial distress more accurately than traditional dictionary-based methods. Using the 2023 U.S. bank failures as a natural experiment, Silicon Valley Bank (SVB), Signature Bank, and First Republic Bank each failed during March–May 2023, we construct monthly sentiment indices for five banks using an LLM alongside VADER, TextBlob, and FinBERT under an identical weighting framework. The LLM index consistently declines ahead of and during the failure period for the three distressed institutions while remaining stable for the two control banks (Bank of America, JPMorgan Chase). VADER, TextBlob, and FinBERT fail to detect the distress, remaining strongly positive throughout. Cohen’s Kappa coefficients near zero (0.01– 0.22) confirm that the methods capture fundamentally different signals. The LLM index is constructed using a severity-weighted aggregation scheme incorporating source credibility, model confidence, and recency, normalised via a tanh transformation. These findings suggest that LLMs interpret financial context rather than merely counting sentiment-bearing words, offering a meaningful advance for early-warning and financial risk monitoring applications.
This paper presents an empirical comparison of lexicon-based and Large Language Model (LLM)-based sentiment analysis for extracting market-relevant signals from social media discourse in highly volatile equity markets. Using Reddit data from r/WallStreetBets and focusing on meme stocks (GME, AMC, NOK), we construct time-aligned sentiment indicators and evaluate their relationship with market returns, with particular attention to extreme positive return events in the upper tail of the return distribution. The LLM-based approach generates multidimensional sentiment representations capturing emotional polarity, bullishness, sarcasm likelihood, and topical relevance, whereas the baseline relies on the VADER lexicon-based model. We evaluate both approaches using lead/lag correlation analysis, OLS regression, ROC-AUC-based directional classification, and a quantile-based early-warning framework. The results indicate that LLM-derived indicators provide a richer multidimensional representation and exhibit stronger asset-specific statistical structure than the lexicon-based baseline. However, their relationship with market movements remains heterogeneous across assets, suggesting that increased linguistic expressiveness does not necessarily translate into stable forecasting performance in retail-driven volatility regimes.
An applied natural-language-processing framework for measuring the tone of Brazilian Monetary Policy Committee (Copom) statements, explicitly inspired by iSent, Ita\'u's Central Bank sentiment classifier, that separates rhetorical tone from policy guidance and uncertainty.
Drawing on more than 3.77 million (3,778,954) U.S. Consumer Financial Protection Bureau (CFPB) complaints filed between 2015 and 2026, this study examines the dynamic linkage between the Chicago Board Options Exchange Market Volatility Index (VIX), a metric for expected market volatility, and the structural composition of consumer financial complaints. To address severe class imbalance and substantial textual noise present in the CFPB corpus, we aggregate complaints into five categories via a class-balanced linear support-vector machine (SVM) that applies per-class thresholding within an interpretable pipeline combining Singular Value Decomposition (SVD) semantic embedding and isotonic calibration; the classifier achieves 85.35% accuracy and a macro-F1 of 0.8207, and although it is trained only on data spanning 2019-2022, it generalizes across time, meaning the COVID-19 period does not appear to act as a meaningful confounder. We then model weekly category shares using a Bayesian Dirichlet-multinomial regression driven by weekly VIX peak values. A rise in the VIX significantly expands the credit-reporting share while reducing the debt-collection and mortgages-and-loans shares, whereas credit card and retail banking categories exhibit no significant shifts in their shares; this association reflects persistent level co-movement instead of a genuine multi-week causal lag. A lag-augmented Toda-Yamamoto test further identifies no reverse Granger feedback running from complaint structure to the VIX. Combined with the macro-exogenous property of the VIX, this finding supports a unidirectional relationship in which market panic shapes complaint composition rather than the opposite. The framework therefore provides regulators with a forward-looking, high-confidence early-warning signal to predict and mitigate consumer-complaint pressure.
Measuring sentiment from financial news is a central task in economics and finance, yet most existing indicators rely on dictionary-based approaches that infer sentiment from word counts and only partially capture context, negation, and semantic structure. This paper proposes a framework for constructing daily news mood indices using transformer-based language models and evaluates whether they better represent sentiment than dictionary-based alternatives. Using 143,755 financial news articles from Factiva, we classify sentiment at the sentence level with FinBERT and aggregate these predictions into article-level and daily sentiment measures through alternative normalization schemes. We compare the resulting indices with benchmark measures based on Shapiro et al., 2022 and Barbaglia et al., 2025. A central contribution is the validation of alternative sentiment measures against human judgments. We conducted an incentivized annotation exercise in which 444 participants evaluated a validation subsample of 588 financial news articles. Consensus ratings from independent human evaluations serve as an external benchmark for assessing the quality of automated sentiment measures. Across correlation, regression, and classification exercises, transformer-based measures show stronger agreement with human judgments than vocabulary-based alternatives and perform substantially better in distinguishing positive, neutral, and negative articles. Overall, the results suggest that incorporating contextual information through transformer-based language models produces sentiment measures that more closely reflect human assessments of financial news.
Maria Saveria Mavillonio, Stefano Borgioli, C. Giannetti et al.· 0 citations
The findings show that QLoRA is effective for financial sentiment adaptation, while also documenting a clear gap between classification accuracy and tradable cross-sectional signals.
Central-bank language often conveys policy direction through qualifying clauses rather than isolated sentiment terms, so evidence-calibrated decision cards are therefore most appropriate for conservative analyst triage rather than autonomous policy interpretation.
Hugo Yan· Journal of Artificial Intell...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.