Feature Based Survey on Fake News Detection: Statistical and Semantic
Abstract
The digital news portals and social media are rapidly expanding, which has significantly increased the spread of fake news, which affects public opinion, social harmony, and trust in information sources. Detection of fake news at an early stage is a critical research challenge. In recent years, researchers have applied multiple methods for the classification of news articles as fake using Natural Language Processing (NLP), Machine Learning (ML), and Deep Learning (DL) techniques. This paper discusses a feature-based analysis of text-oriented fake news detection methods published from 2017 to 2025. The analysis demonstrates how different textual features, such as linguistic, stylistic, psychological, statistical, semantic, and syntactic features, are used for the identification of fake or real news. A comparative analysis of existing studies shows that most research primarily depends on statistical and semantic representations like N-grams, TF-IDF, and word embeddings, whereas linguistic, stylistic, and psychological cues are comparatively less explored. In addition, syntactic features have gained very limited attention despite their potential to enhance detection performance. The review emphasizes integrating multiple feature types to develop more reliable and interpretable detection systems. It also identifies research gaps and suggests future directions for developing comprehensive feature-based frameworks for fake news detection