A blockchain-based commit-reveal protocol using Autonomous Economic Agents on an Ethereum-compatible ledger creates a tamper-evident audit trail that separates blind evaluation from post-hoc claims and reduces the verification burden on independent researchers and leaderboard operators.
Abstract
LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims that DeepSeek R1 outperformed OpenAI's o1 contributed to market panic on January 27, 2025, when Nvidia lost USD589 billion in market value. Yet vendor benchmarks often depend on an honor system. Academic reassessments and independent leaderboards have found undisclosed changes to proprietary models, contaminated training data, and selective reporting. LLM-as-a-judge methods scale evaluation by reducing human review. Studies, however, suggest that judges may show identity-aware bias, scoring an answer according to its source model rather than its quality. This bias has not been fully measured or corrected across politically sensitive, reasoning-intensive, and preference-based tasks. We examine this problem using seven verifier models: GPT-OSS 120B, Llama 3.3 70B, GLM 5.1, Qwen3 32B, DeepSeek V4 Pro, Mistral Large3, and Sarvam M. They score anonymous and identity-disclosed responses from three primary models on 58 factual, reasoning, political, and preference-based questions. Identity disclosure slightly raises scores for factual questions, moderately affects stress-reasoning tasks, and causes large changes for geopolitically sensitive topics. Notable results include GLM5.1 (+7.00 points, p = 0.0249) and Llama 3.3 70B (+1.56 points, p = 0.00). We also introduce a blockchain-based commit-reveal protocol using Autonomous Economic Agents on an Ethereum-compatible ledger. In Phase 1, each judge records a one-way hash of its score and a secret salt before candidate identities are revealed. In Phase 2, the identity and raw score are disclosed and verified on-chain. This creates a tamper-evident audit trail that separates blind evaluation from post-hoc claims and reduces the verification burden on independent researchers and leaderboard operators.
This work argues for a hybrid public-private approach: keeping the benchmark content private to maintain integrity while making the benchmark design public to build trust, which offers a path toward trustworthy measurement of ever-growing data-intensive AI systems.
Sang T. Truong, S. Wang, Angelina Wang et al.· 0 citations
Findings indicate that state-of-the-art LLMs can achieve a high degree of demographic neutrality; fundamental artefacts such as positional bias can nonetheless produce severely discriminatory outcomes; and bias auditing must extend beyond demographic parity to interaction artefacts and ecosystem structure.
A. Camargo, Rafaela Silva Figueiredo Camargo· Revista de Geopolítica· 0 citations
Third-party news-source factuality scorecards are valuable but increasingly fragile. Web pages change, access conditions shift and underlying datasets may disappear. The challenge is therefore not only benchmark imperfection but also benchmark sustainability as credibility datasets, search interfaces and platform reputation signals become harder to access reproducibly. This study investigates whether a fixed large language model (LLM) scoring procedure can generate durable, replayable outlet-level trust scores that align with a frozen external factuality benchmark rather than objective ground truth. Fifty-two English-language news outlets were assessed across nine predefined trust dimensions and compared with a frozen Media Bias Fact Check (MBFC) factuality snapshot. Raw LLM scores were rank-aware but compressed (Pearson’s r=0.801, Spearman’s ρ=0.843, full-cohort mean GAP =0.221). An affine calibration fitted on 42 training outlets increased full-cohort Pearson alignment to r=0.828 and reduced mean GAP to 0.090; on the fixed ten-outlet validation fold, mean GAP fell from 0.162 to 0.063. Across 1000 additional stratified 42/10 splits, median validation GAP was 0.078 (central 95% split range 0.048–0.110). Wikipedia lead and source-weighted web enrichment did not outperform the calibrated archival path in the retained data. The Step 4 unweighted web-search meter improved on the Wikipedia-lead meter (Pearson’s r=0.697, Spearman’s ρ=0.576, full-cohort mean GAP =0.200; fixed-validation GAP =0.146) but remained below the calibrated archival path. RSS monitoring is reported separately as an asymmetric, bounded adverse-event signal rather than a second factuality benchmark. These findings support calibrated LLM trust vectors as a potentially useful archival proxy while highlighting benchmark dependence, sampling constraints, model sensitivity and the importance of reproducible data provenance.
FairFund-Bench is introduced, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised, indicating that current LLMs robustly reproduce human deservingness evaluations.
This work benchmarks six estimators across five datasets and two known-effect sweeps, and validate the mechanisms against a non-simulated paired reference, finding that weak overlap is governed by logger-target action alignment, not by logging sharpness alone.
The Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable.
Ziyue Wang, Aomufei Yuan, Yi-Ran Yao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.