Skip to content
Review Open access

CogDeBias: An LLM-Based Multilingual Framework for Cognitive Bias Detection and Mitigation in Corporate Decision-Making Texts

2026 · IEEE Access · Vol 14, pp. 123144-123164 · 0 citations · 39 references

Abstract

Cognitive biases embedded in corporate strategic communications pose significant risks to investment decisions and governance quality, yet existing detection approaches rely on manual analysis that lacks scalability. This study presents CogDeBias, a multilingual framework integrating large language models (LLMs) and machine learning (ML) for automated detection and mitigation of cognitive biases in enterprise decision-making texts. A bilingual annotated corpus of 1,200 corporate annual reports (600 English, 600 Chinese) was constructed, encompassing 30,416 sentences, of which 8,082 are bias-positive (26.6%) and carry approximately 9,230 bias-category instances across six categories: confirmation bias, sunk cost fallacy, overconfidence, anchoring effect, bandwagon effect, and recency bias. The hybrid architecture combines XLM-RoBERTa multilingual encoding, Llama-3-70B prompt-based classification, and XGBoost ensemble learning, achieving a weighted F1-score of 0.82 on the test set, surpassing rule-based (0.59), BERT-based (0.72), zero-shot GPT-4 (0.76), and few-shot Llama-3 (0.785) baselines, with the 3.5-percentage-point margin over the strongest baseline statistically significant (p = 0.008, McNemar’s test). On a fully expert-annotated subset labelled without any model assistance, the framework retained a weighted F1 of 0.81, indicating that the reported performance is not an artefact of model-to-model label agreement. Cross-lingual evaluation revealed English F1 of 0.84 versus Chinese F1 of 0.80, with zero-shot transfer from English to Chinese yielding F1 of 0.68, improving to 0.78 with minimal target-language fine-tuning. The Mixtral-8x7B-powered mitigation generator produced actionable suggestions rated 4.1/5 by 30 financial analysts (Fleiss’ kappa = 0.71). These suggestions are decision-support outputs for human review, not automated corrections to regulated filings. Processing efficiency of 0.42 seconds for core model inference (3.0 seconds end-to-end per document) supports batch-oriented use as decision support in investor due diligence, corporate audit, and regulatory monitoring workflows, subject to domain-specific calibration. This work establishes a computational paradigm for decision science applications, demonstrating that linguistic patterns reliably surface systematic reasoning distortions across languages and corporate contexts.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.