Skip to content

Cross-lingual transfer learning for fake news detection: leveraging high-resource language for low-resource adaptation

Aug 2026 · Neural computing & applications (Print) · Vol 38 · 0 citations · 40 references

TL;DR

A transfer-cum-ensemble learning framework that integrates a task-specific pretrained model (XLM-RoBERTa) with a lightweight Indian-language model (IndicBERT) with a weighted-average attention mechanism is used to combine these sophisticated language models for further enhancing performance.

View source

Similar papers

Open access Jul 2026

FakeDiverse a curated multi-source news corpus for context-aware fake news detection using BERT and DeBERTa

This study examines the effectiveness of two transformer-based architectures—BERT and DeBERTa—for identifying fake news using only textual information from headlines and article bodies and achieves strong performance on FakeDiverse corpus, demonstrating the need for enhanced generalization strategies as well as domain adaptation.

A. Kumar, A. S, Akshara G. Bhat et al. · 0 citations
Open access Aug 2026

Multilingual Fake News Detection Using Machine Learning with Contextual-Based Feature Extraction

The proposed approach provides a simple and efficient solution for multilingual fake news detection in data-scarce environments with ensemble-based classifiers such as Random Forest and Gradient Boosting achieving reliable performance across both languages.

Nikita Garg, Pritam Singh Negi · 0 citations
Open access Jul 2026

Fake News Identification Using Hybrid Transformer Ensemble Approach

A hybrid transformer-based ensemble model for automated fake news identification using the FakeNewsNet dataset is proposed and Experimental results show that the ensemble model achieves an accuracy of approximately 93%, outperforming the individual constituent models.

E. C. Babu, G. Sukanya · 0 citations
Open access Aug 2026

A Three-Stage Cross-Lingual Knowledge Transfer Approach Based on the XLM-RoBERTa Model for Detecting Fake News in Ukrainian

In recent years, there has been an increase in the amount of fake news in the media, which is why fact-checking systems are gaining popularity, particularly those that use natural language processing (NLP) to quickly identify and flag fake news. One of the main limitations in the development of such systems is the limited number of datasets containing verified information, which are necessary for the effective training of models. The situation is particularly critical for non-English datasets, specifically those in the Ukrainian language. This article proposes a three-stage algorithm for training a model to recognize fake news in the Ukrainian language. At the core of the proposed approach lies the multilingual transformer model XLM-RoBERTa, which solves this problem by utilizing cross-lingual knowledge transfer from English to Ukrainian. This approach means there is no need to search for a large, high-quality dataset in Ukrainian; instead, a significantly smaller dataset in Ukrainian can be used for the final calibration of the model. The model developed as a result of the experiment proved effective in extreme low-resource scenarios, achieving 90.7% accuracy on just 500 training records and outperforming the baseline model by 9.7%.

Volodymyr Smahliuk, Ya. Kovivchak, Yu. Kynash · 0 citations
Open access Aug 2026

Fake News Detection Using Machine Learning and LLM Embeddings: A Comparative Study of TF-IDF and BERT Representations on the Welfake Dataset

The proposed framework highlights the potential of integrating transformer-based language models with classical machine learning algorithms to build robust and scalable fake news detection systems.

Umme Noor Us Saqa, Sreenivasa B. R. · 0 citations
Open access Jul 2026

AZFAKENEWS: A BENCHMARK DATASET FOR FAKE NEWS DETECTION IN THE AZERBAIJANI LANGUAGE

Disinformation on digital platforms is a very important problem for public trust and political debate. Even so, most research on automated fake news detection has stayed on a small number of high-resource languages. This paper presents AzFakeNews, the first large-scale benchmark dataset for fake news detection in Azerbaijani. Azerbaijani is a low-resource Turkic language with more than 30 million speakers. We built the dataset from two sources: scraping of news articles from five major Azerbaijani media websites, and generation of synthetic fake articles via Meta's LLaMA language model. The final dataset has 3,782 articles (3,000 authentic and 782 synthetic) across 34 topic categories. For the baseline we fine-tuned the Azerbaijani aLLMA model on this corpus. It reached 91.03% accuracy and 91.11% macro-F1 on the test split, because of which we consider it a strong baseline compared to mBERT, XLM-RoBERTa and the base LLaMA on the same data. We hope the dataset helps close a gap in Azerbaijani NLP and supports cross-language research on disinformation.

Jalal Mehdiyev, Vusal Shahbazov · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.