Disinformation on digital platforms is a very important problem for public trust and
political debate. Even so, most research on automated fake news detection has stayed
on a small number of high-resource languages. This paper presents AzFakeNews, the
first large-scale benchmark dataset for fake news detection in Azerbaijani. Azerbaijani
is a low-resource Turkic language with more than 30 million speakers. We built the
dataset from two sources: scraping of news articles from five major Azerbaijani media
websites, and generation of synthetic fake articles via Meta's LLaMA language model.
The final dataset has 3,782 articles (3,000 authentic and 782 synthetic) across 34 topic
categories. For the baseline we fine-tuned the Azerbaijani aLLMA model on this corpus.
It reached 91.03% accuracy and 91.11% macro-F1 on the test split, because of which we
consider it a strong baseline compared to mBERT, XLM-RoBERTa and the base LLaMA
on the same data. We hope the dataset helps close a gap in Azerbaijani NLP and
supports cross-language research on disinformation.
Jalal Mehdiyev, Vusal Shahbazov· Problems of Information Tech...· 0 citations
Despite the advancements made by researchers, spam emails remain one of the biggest challenges in the field of cybersecurity. Spam emails can serve as phishing emails or carry viruses that compromise the security of an organization's system. Current detection techniques depend on supervised learning or rely on cloud-based services, which can compromise user data privacy and affect implementation flexibility. This paper evaluates the capability of five large language models (LLMs) in zero-shot spam email classification. The models used in this study include llama3.1:8b, deepseek-r1:8b, gemma3:4b, falcon3:7b, and mistral:7b. In addition to predicting whether the email is spam or not, the LLM was also asked to generate an explanation of its prediction in natural language form. The experiments were conducted on two benchmark datasets: the Ling and TREC2007 datasets. In terms of performance, llama3.1:8b outperformed other LLMs when evaluated on the TREC2007 dataset (98.78% accuracy) and deepseek-r1:8b had the best performance on the Ling dataset (98.79%). The results show that open-weight LLMs can achieve competitive spam detection performance in a local, privacy-preserving environment without any fine-tuning.
Vusal Shahbazov· 2026 7th International Confe...· 0 citations
Examination of sentiment analysis methods applied to social media text data, covering lexicon-based, machine learning, and deep learning approaches, including transformerbased architectures, as well as widely used datasets, shows how sentiment analysis can be applied to the detection of threats that exploit human emotions.
Vusal Shahbazov· “Kibertəhlükəsizlik və rəqəm...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.