It is suggested that predictive performance, attribution plausibility, and mechanistic faithfulness characterize different aspects of model behavior and should be evaluated separately when studying explainability in media bias detection.
This study deploys a scalable machine learning pipeline: combining a transformer-based classifier applied to 2.01 million English-language AI-related news headlines (July 2022–July 2024) with large-language-model and human-annotator validation (three annotators, Fleiss’ κ=0.80) on stratified subsamples, to extract six interpretable, bias-linked discourse indicators computed at the AI-domain level: evaluative orientation (valence), loss salience, narrative drift, exposure-adjusted sentiment, cross-source divergence, and novelty-phase framing. Each operationalizes an established cognitive-psychology construct as a computable property of the information environment associated with biased risk–benefit reasoning. Results show systematic variation across domains: technical and methodological areas such as deep learning and natural language processing exhibit gain-salient framing, while safety-critical topics such as deepfakes (loss-to-gain headline ratio = 3.17) and facial recognition show strongly loss-salient profiles. Cross-model validation using an LLM on a stratified sample of 1000 headlines confirms that domain-level indicator rankings are robust to classifier choice (Spearman ρ=0.83; p<0.001), establishing the rank stability of pipeline outputs independently of the specific classification architecture. As a contextual application, domain-level profiles are mapped to European Union AI governance instruments, documenting parallels between discourse patterns and regulatory risk tiers. The framework provides a scalable, reproducible methodology for monitoring evaluative conditions in technology news across domains, sources, and time.
O. Topal, Inna Novalija, Joao Pita Costa et al.· Applied Informatics· 0 citations
Media bias detection relies on definitions and examples that specify what counts as bias, yet these specifications often vary across datasets or remain implicit, even when given the same name. Such variation makes it unclear whether models trained for the same bias category learn the same construct or different phenomena, a problem largely overlooked in prior work. We examine how definition choice affects bias annotation in a between-subjects experiment with 354 participants and a parallel evaluation with four LLMs. Participants and models rate six news articles across four bias categories using definitions that vary in conceptual framing and elaboration. Across 8,496 human and 28,800 LLM ratings, we find that the conceptual target of a definition drives annotation divergence, while construct-preserving elaboration does not: conceptual framing significantly shifts annotations for humans and does so even more strongly for LLMs. We discuss implications for construct specification in annotation protocols and prompt-based measurement, and consider how definitional sensitivity may propagate to downstream classification beyond media bias. We also release MUDD, the Multi-Definition Bias Detection Dataset.
Martin Wessel, Timo Spinde, Jürgen Pfeffer et al.· 0 citations
HEF-XFND is proposed, a hybrid explainable feature-fusion framework that combines sparse lexical evidence, contextual transformer representations, source-level credibility indicators, and calibrated ensemble learning that addresses three recurring limitations in fake-news research.
Raju M, Subalakshmi Kannan, P. P.· International journal of res...· 0 citations
The Israel-Palestine conflict is one on which public opinion is greatly influenced by media bias. Detecting and understanding such bias in news reporting is necessary to promote transparency and accountability in journalism. The issue of media bias detection using advanced deep learning techniques is addressed in this research, particularly transformer-based architectures such as Distilled Bidirectional Encoder Representations from Transformers (DistilBERT), Bidirectional Encoder Representations from Transformers Mini (BERT-Mini), Compact Bidirectional Encoder Representations from Transformers (TinyBERT). A dataset of 9,000 news articles is used, labeled with sentiment using the TextBlob library, with sentiment serving as a measurable proxy for the emotional framing dimension of media bias. We evaluate the performance of these transformer models against more traditional deep learning architectures consisting of Gated Recurrent Unit (GRU), Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM). We find that transformer-based models greatly outperform their sequential counterparts. The best performance was achieved by DistilBERT with 94% accuracy, 89% precision, 86% recall and F1 score of 87%, verifying that DistilBERT possesses the best capability to capture subtle contextual nuances which specify media bias. In this study, we demonstrate the potential of lightweight transformer-based language models to enable bias detection in digital journalism with a scalable and consistent approach. Finally, our findings are situated in the growing field of automated media bias profiling, and we call for further work to extend the use of machine learning for promoting fair, transparent news coverage. Future work should combine multilingual and multimodal data to get a more complete sense of bias over different media landscapes.
Saba Saddique, Usman Ahmad· Journal of Intelligent Syste...· 0 citations
Purpose: Automated fake news detection can support verification, but predictions are less useful when textual and contextual evidence cannot be inspected. The study examined whether combining claim text with speaker metadata could improve veracity classification while retaining explanations for journalistic review.
Methods: The LIAR train, validation, and test partitions were used, with six veracity labels remapped into Fake and Real classes. A BERT-base text classifier and Extreme Gradient Boosting metadata classifier produced class probabilities combined by an L1-regularized logistic regression meta-learner with disagreement and interaction features. SHapley Additive exPlanations, Local Interpretable Model-agnostic Explanations, and Integrated Gradients inspected metadata, ensemble, and token-level behavior.
Findings: The stacking model reached 0.7514 accuracy, 0.7483 macro-F1, 0.8288 area under the receiver operating characteristic curve, and 0.4970 Matthews correlation coefficient. It exceeded the BERT text branch (macro-F1=0.6203) and metadata branch (macro-F1=0.7066). Metadata provided the stronger signal, with credibility score and speaker-history variables receiving the largest SHAP importance. High text-metadata disagreement occurred in 30.4% of test cases, and the ensemble reached 0.8234 accuracy within this subset.
Originality: The study extends explainable fake-news detection by conceptualizing disagreement between evidence sources as a dimension of interpretability, separating semantic claim evidence, contextual speaker-history evidence, and their interaction at the fusion stage.
K. M. T. Bin Parves, Prashanta Kumar Shill· Jurnal the Messenger· 0 citations
In many settings, studying causal questions based on text data requires adjusting for confounding information within texts. Yet there is a tradeoff in constructing text representations for adjustment: they must be sufficiently large and/or dense to preserve the confounding variables necessary for unbiased effect estimation, but sufficiently small and/or sparse to satisfy finite-sample overlap and yield low-variance estimates. To address this tradeoff, we turn to sparse autoencoders (SAEs), and propose a novel causal adjustment pipeline that iteratively selects a minimal set of SAE features via conditional independence tests. We find that SAE representations achieve better adjustments (lower bias and and higher coverage) than alternative representations in standard semi-synthetic evaluations with binary confounders, and their interpretability offers opportunities for falsification. We also introduce a more realistic semi-synthetic evaluation that uses multi-label data as the unobserved confounders and find off-the-shelf adjustment methods require increased investigation for these more complex settings. Code: https://github.com/mianzg/sae-text-confounder
Mian Zhong, Katherine A. Keith, Anjalie Field· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.