Skip to content

InfoOps Bench: A live information operations safety benchmark

Jul 2026 · arXiv.org · Vol abs/2607.28503 · 0 citations · 48 references
Computer Science

TL;DR

An active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted for use by authoritarian state information operations, and shows the potential for contemporary information operations to be substantially aided by frontier AI models.

Abstract

In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted for use by authoritarian state"information operations": intentional, coordinated activities by one state to influence public opinion and information ecosystems in another state. These information operations are a well documented, persistent threat against contemporary democracy. Our benchmark is based on real examples from over 2,100 information operations drawn from a live monitoring pipeline which tracks online information assets with links to authoritarian regimes. Alongside this paper, we also release a companion website that updates the benchmark weekly with new claims. The dynamic nature of this public facing benchmark makes it resistant to saturation. In the benchmark, we test 17 models from 8 providers across four prompt framings. We find that most models can be co-opted for information operations at least some of the time. Integrity scores, defined as the share of judged responses in which the model neither preserved nor amplified the claim, range from 9.3% to 91%, an 81.7-percentage-point spread not explained by model size. Models approach participation in information operations in a variety of ways. Some models fabricate details and produce output more harmful than the original input claim; others make claims less harmful even while complying and producing some output. Fact-checking rates vary from 3.2% to 80.8%. Integrity against information operations is at least partly related to refusal to produce content even for benign claims, illustrating the challenge of balancing model usability with safety. Overall, our results show the potential for contemporary information operations to be substantially aided by frontier AI models.

View source

Similar papers

Review Jul 2026

Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications

Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monitoring, and human escalation. Yet evaluation often remains model-centric: benchmark scores, task accuracy, or one-off qualitative reviews are treated as evidence of readiness. In financial settings, this is insufficient. We take the position that financial LLM systems should not be approved for production based on benchmark performance alone. They require system-level validation evidence across the application stack: data, model design, retrieval and generation performance, agent behavior, governance, and implementation. Drawing on industry experience validating GenAI applications in financial institutions, we outline a multi-layer validation view and explain why hybrid evaluation is necessary. We discuss where LLM-as-a-judge methods are useful and why they require controls such as multiple judges, rubrics, agreement, and auditability checks. We also highlight failure modes poorly captured by static benchmarks, including retrieval failures, unfaithful generation, tool misuse, escalation errors, and operational instability. Our position is that financial LLM validation should be an ongoing system discipline rather than a one-time model scoring exercise. Validation should produce decision-ready evidence, not only scores. We conclude with a research agenda for system-aware benchmarks, agent trace validation, judge alignment protocols, and lifecycle validation standards.

Burak Payzun, Irem Demirtas, Simona Scala et al. · 0 citations
Aug 2026

What is Missing? An HFACS Analysis of the VERIS Community Database

Results suggest that current public reporting can identify the type of error that occurred, but is limited in explaining why it occurred, and suggest that current public reporting can identify the type of error that occurred, but is limited in explaining why it occurred.

Saroja Roy Grandhi, Jeremiah D. Still · 0 citations
Open access Aug 2026

LAADS: Design and Implementation of a Sector-Aware Security Advisory Platform for Academic Institutions

Academic institutions increasingly hold data and infrastructure worth targeting, yet most lack a dedicated Security Operations Center capable of turning raw threat intelligence into advisories that are actually distributed to and acted on by their community. Existing Security Information and Event Management stacks handle log aggregation, and threat-intelligence platforms handle indicator storage, but neither produces institution-tailored, human-readable bulletins or delivers them across the channels an institutional audience actually uses. This paper presents the design and implementation of LAADS (LLM-Augmented Advisory Authoring and Dissemination System), a sector-aware security advisory platform built for this gap: a five-layer web system that authors advisories through three entry paths (manual authoring, AI-assisted extraction from a submitted article, and escalation of an external feed item), enforces sector-scoped multi-tenancy across five operational sectors at four independent layers, correlates advisories against an institutional asset inventory through a three-tier matching engine, and disseminates published advisories over six channels spanning real-time push, e-mail, chat webhooks, PDF export and share links. We describe the architecture, data model and enforcement mechanisms in detail and report functional verification of each subsystem. Quantitative evaluation of enrichment performance under production load is left as future work.

Rama Murthy Naidu Yerramsetty Chinna, Krishna T Siva Rama · 0 citations

DelistBench: Evaluating Search-Enabled LLMs for Auditable Corporate-Event Database Completion

Financial institutions need an independent way to detect missing, stale, and misclassified corporate-event records in vendor databases. We introduce Search-to-Record, a database-assurance task in which search-enabled large language models reconstruct institution-defined event records from public sources for a known security universe and historical cutoff, and DelistBench, a 1,200-record benchmark for security-level delisting announcements. We evaluate five models in paired closed-book and web-enabled conditions. Web access raises announcement-date accuracy within seven days by 34.0 to 48.0 percentage points and event-status accuracy by approximately 2.8 to 21.7 points; the best system achieves 81.5% overall joint accuracy within seven days. Economy web systems achieve 75.9-78.3% overall joint accuracy within seven days at 4.5-6.6% of the API cost of the most expensive web system. Risk-based triage identifies low-error subsets, although the highest-coverage operating point still sends 27.3% of the balanced test set to review. The evaluation identifies web retrieval as the main source of timing gains and shows that low-cost systems can approach the best system's accuracy. Together, Search-to-Record, DelistBench, and the evaluation provide concrete deployment guidance: calibrate triage to local event prevalence and market mix, preserve positive-event recall, and route positive and ambiguous cases to targeted review.

Yao Xuan, Shuping Li, Yang Dai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.