Skip to content
Preprint

Reading Copom's Tone: A Weighted LLM Framework for Hawkish-Dovish Sentiment, Forward Guidance, and Uncertainty

Aug 2026 · 0 citations · 8 references
Economics Computer Science

TL;DR

An applied natural-language-processing framework for measuring the tone of Brazilian Monetary Policy Committee (Copom) statements, explicitly inspired by iSent, Ita\'u's Central Bank sentiment classifier, that separates rhetorical tone from policy guidance and uncertainty.

Abstract

This paper documents an applied natural-language-processing framework for measuring the tone of Brazilian Monetary Policy Committee (Copom) statements. The project is explicitly inspired by iSent, Ita\'u's Central Bank sentiment classifier, particularly its sentence-level division of official communication into hawkish, dovish, neutral, and out-of-context classes. The implementation extends that idea in three directions. First, an LLM identifies short hawkish and dovish expressions and assigns each a 0-to-1 intensity weight. Second, the document index combines sentence counts with document-specific average signal intensities, producing a bounded score from -1 to 1. Third, a separate full-document layer measures forward-guidance direction, guidance explicitness, uncertainty level, and change in uncertainty. The empirical sample is restricted to communications dated August 2016 or later and contains 80 statements and 1,498 classified sentences from August 31, 2016 through August 5, 2026. Across this sample, 33.3% of sentences are hawkish, 18.0% dovish, 42.1% neutral, and 6.5% out of context. The average document score is +0.107, while the most hawkish reading is +0.570 in August 2021. The latest statement, dated August 5, 2026, scores +0.232, with eight hawkish, two dovish, and nine neutral sentences. Its structural overlay is more nuanced: guidance is directionally ambiguous but partly explicit, while uncertainty is classified as central and higher than at the prior meeting. Tone and the guidance-direction score have a contemporaneous Pearson correlation of 0.719. These are descriptive outputs, not a validated forecast of Selic decisions or DI returns. The main contribution is therefore methodological: a transparent, incremental, auditable system that separates rhetorical tone from policy guidance and uncertainty.

View source

Similar papers

Open access Jul 2026

Evidence-Calibrated Financial Language Models for Macro-Policy Stance Classification and Decision Cards

Central-bank language often conveys policy direction through qualifying clauses rather than isolated sentiment terms, so evidence-calibrated decision cards are therefore most appropriate for conservative analyst triage rather than autonomous policy interpretation.

Hugo Yan · 0 citations
Review Open access 2026

CogDeBias: An LLM-Based Multilingual Framework for Cognitive Bias Detection and Mitigation in Corporate Decision-Making Texts

Cognitive biases embedded in corporate strategic communications pose significant risks to investment decisions and governance quality, yet existing detection approaches rely on manual analysis that lacks scalability. This study presents CogDeBias, a multilingual framework integrating large language models (LLMs) and machine learning (ML) for automated detection and mitigation of cognitive biases in enterprise decision-making texts. A bilingual annotated corpus of 1,200 corporate annual reports (600 English, 600 Chinese) was constructed, encompassing 30,416 sentences, of which 8,082 are bias-positive (26.6%) and carry approximately 9,230 bias-category instances across six categories: confirmation bias, sunk cost fallacy, overconfidence, anchoring effect, bandwagon effect, and recency bias. The hybrid architecture combines XLM-RoBERTa multilingual encoding, Llama-3-70B prompt-based classification, and XGBoost ensemble learning, achieving a weighted F1-score of 0.82 on the test set, surpassing rule-based (0.59), BERT-based (0.72), zero-shot GPT-4 (0.76), and few-shot Llama-3 (0.785) baselines, with the 3.5-percentage-point margin over the strongest baseline statistically significant (p = 0.008, McNemar’s test). On a fully expert-annotated subset labelled without any model assistance, the framework retained a weighted F1 of 0.81, indicating that the reported performance is not an artefact of model-to-model label agreement. Cross-lingual evaluation revealed English F1 of 0.84 versus Chinese F1 of 0.80, with zero-shot transfer from English to Chinese yielding F1 of 0.68, improving to 0.78 with minimal target-language fine-tuning. The Mixtral-8x7B-powered mitigation generator produced actionable suggestions rated 4.1/5 by 30 financial analysts (Fleiss’ kappa = 0.71). These suggestions are decision-support outputs for human review, not automated corrections to regulated filings. Processing efficiency of 0.42 seconds for core model inference (3.0 seconds end-to-end per document) supports batch-oriented use as decision support in investor due diligence, corporate audit, and regulatory monitoring workflows, subject to domain-specific calibration. This work establishes a computational paradigm for decision science applications, demonstrating that linguistic patterns reliably surface systematic reasoning distortions across languages and corporate contexts.

Yutong Shen, Wang Yang, Yue Shen · 0 citations
Jul 2026

LLM-Assisted Sentiment Analysis for Indonesia's Coretax Policy: A Knowledge Distillation and Human-in-the-Loop Approach

The lack of high-quality labeled datasets remains a major challenge for sentiment analysis in low-resource languages such as Indonesian, particularly in specialized domains like fiscal policy. This study investigates the effectiveness of Large Language Models (LLMs) as automated annotators within a teacher-student knowledge distillation framework. Using social media data from X related to Indonesia's Coretax system, three training scenarios were evaluated: AI-labeled data, human-labeled data, and a hybrid approach. The results show that GPT-4o achieves substantial agreement with human annotators, with a Cohen's Kappa score of 0.61. Furthermore, the student model IndoBERT trained on the combined dataset outperforms other configurations, achieving a Macro F1-score of 0.64 and a Macro ROC-AUC of 0.84. These findings indicate that while LLMs cannot fully replace human judgment, they significantly enhance scalability and enable near real-time policy evaluation in low-resource settings through effective human-AI collaboration.

Novialdi Ashari, Ulfah Oktarida Sihaloho, Novi Aulia Sari · 0 citations
Preprint Jul 2026

Measuring Sentiment News with Transformer-Based Language Models

Measuring sentiment from financial news is a central task in economics and finance, yet most existing indicators rely on dictionary-based approaches that infer sentiment from word counts and only partially capture context, negation, and semantic structure. This paper proposes a framework for constructing daily news mood indices using transformer-based language models and evaluates whether they better represent sentiment than dictionary-based alternatives. Using 143,755 financial news articles from Factiva, we classify sentiment at the sentence level with FinBERT and aggregate these predictions into article-level and daily sentiment measures through alternative normalization schemes. We compare the resulting indices with benchmark measures based on Shapiro et al., 2022 and Barbaglia et al., 2025. A central contribution is the validation of alternative sentiment measures against human judgments. We conducted an incentivized annotation exercise in which 444 participants evaluated a validation subsample of 588 financial news articles. Consensus ratings from independent human evaluations serve as an external benchmark for assessing the quality of automated sentiment measures. Across correlation, regression, and classification exercises, transformer-based measures show stronger agreement with human judgments than vocabulary-based alternatives and perform substantially better in distinguishing positive, neutral, and negative articles. Overall, the results suggest that incorporating contextual information through transformer-based language models produces sentiment measures that more closely reflect human assessments of financial news.

Maria Saveria Mavillonio, Stefano Borgioli, C. Giannetti et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.