Skip to content
Open access

Calibrating LLM-Derived Trust Scores for News Outlets When Public Factuality Scorecards Disappear

Unknown authors
Sep 2026 · Information · 0 citations · 10 references

Abstract

Third-party news-source factuality scorecards are valuable but increasingly fragile. Web pages change, access conditions shift and underlying datasets may disappear. The challenge is therefore not only benchmark imperfection but also benchmark sustainability as credibility datasets, search interfaces and platform reputation signals become harder to access reproducibly. This study investigates whether a fixed large language model (LLM) scoring procedure can generate durable, replayable outlet-level trust scores that align with a frozen external factuality benchmark rather than objective ground truth. Fifty-two English-language news outlets were assessed across nine predefined trust dimensions and compared with a frozen Media Bias Fact Check (MBFC) factuality snapshot. Raw LLM scores were rank-aware but compressed (Pearson’s r=0.801, Spearman’s ρ=0.843, full-cohort mean GAP =0.221). An affine calibration fitted on 42 training outlets increased full-cohort Pearson alignment to r=0.828 and reduced mean GAP to 0.090; on the fixed ten-outlet validation fold, mean GAP fell from 0.162 to 0.063. Across 1000 additional stratified 42/10 splits, median validation GAP was 0.078 (central 95% split range 0.048–0.110). Wikipedia lead and source-weighted web enrichment did not outperform the calibrated archival path in the retained data. The Step 4 unweighted web-search meter improved on the Wikipedia-lead meter (Pearson’s r=0.697, Spearman’s ρ=0.576, full-cohort mean GAP =0.200; fixed-validation GAP =0.146) but remained below the calibrated archival path. RSS monitoring is reported separately as an asymmetric, bounded adverse-event signal rather than a second factuality benchmark. These findings support calibrated LLM trust vectors as a potentially useful archival proxy while highlighting benchmark dependence, sampling constraints, model sensitivity and the importance of reproducible data provenance.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.