Skip to content
Preprint

AI Security Leaderboard: Methodology, Results and Minimal Standard

Aug 2026 · 1 citation · 21 references
Computer Science

TL;DR

This report tested four leading models for universal jailbreaks in the context of the FAR$AI Minimal Standard for Safeguards, and found more than a hundredfold difference in security.

Abstract

The AI Security Leaderboard is an independent benchmark that ranks the safeguards of frontier AI models from least to most secure. It tests models against the FAR$.$AI Minimal Standard for Safeguards, which represents a minimum bar for security: meeting it does not guarantee a secure model, but failing to meet it guarantees a lack of state-of-the-art security. Version 1.0 covers severe misuse requests across chemical, biological, radiological, nuclear, and explosive (CBRNE) threats and offensive cybersecurity. In this report, we tested four leading models for universal jailbreaks in the context of this minimal standard, and found more than a hundredfold difference in security. Claude Fable 5 and GPT-5.6 Sol held against every attack we ran, with no universal jailbreak found; we estimate they would likely cost more than \$14,200 to jailbreak, if it is possible with this methodology at all. Meanwhile, we found hundreds of universal jailbreaks for Grok 4.5 and Gemini 3.1 Pro; each broke for under \$300, with universal jailbreaks in Grok's weakest domain, cybersecurity, accessible for as little as \$24. The gap is fixable: every weakness we found belongs to a known class of attack that already has a defense deployed in production models. The leaderboard will be updated on a rolling basis as new models are released, and the evaluation methodology and Minimal Standard will be periodically revised to take into account the latest capabilities and the state-of-the-art in safeguards. The leaderboard is available at leaderboard.far.ai.

View source

Similar papers

#cybersecurity Preprint Aug 2026

The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark

SRE-Bench is introduced, the first realistic, contamination-free RE benchmark, and results indicate that strong source-code security capabilities do not yet transfer to binary analysis, highlighting RE as an important frontier for agentic cybersecurity and SRE-Bench as a rigorous testbed to measure progress.

J. Spence, Nicholas Assaderaghi, Feng Xiao et al. · 1 citation
#artificial intelligence Review Sep 2026

Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges

This survey systematically analyze threat surfaces across intra-execution loops and interaction planes, formulate a multi-layered zero-trust defense-in-depth architecture integrating Dual-LLM isolation, Capability-Based Access Control, kernel eBPF probes, and sandboxed runtimes, and map technical controls to internatio...

Seyedakbar Mostafavi · 0 citations

AIOracles: Modeling Agentic AI Security *

A formal framework for security in agentic AI is developed, built around an ideal predicate ϕ that decides, for a given ( prompt, context, result ) triple, whether the agent’s behaviour is acceptable.

F. Durak, Tadayoshi Kohno, Franziska Roesner · 0 citations
#artificial intelligence Preprint Oct 2026

Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks

Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark's measur...

Neeraj Karamchandani, Piyush Nagasubramaniam, Xin-Hong Xie et al. · 0 citations
Case report Open access Aug 2026

Open by Design, Exposed by Default: Data Security and the Strategic Challenge of PLA Intelligentised Warfare

Key Takeaways: - What counts as relevant is being redefined by intelligentisation, where data quality determines AI effectiveness, making peacetime civilian data acquisition a form of military preparation. - The establishment of ISF in April 2024 institutionalised the doctrinal shift of information gathering into int...

M. Descamps · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.