This report tested four leading models for universal jailbreaks in the context of the FAR$AI Minimal Standard for Safeguards, and found more than a hundredfold difference in security.
Abstract
The AI Security Leaderboard is an independent benchmark that ranks the safeguards of frontier AI models from least to most secure. It tests models against the FAR$.$AI Minimal Standard for Safeguards, which represents a minimum bar for security: meeting it does not guarantee a secure model, but failing to meet it guarantees a lack of state-of-the-art security. Version 1.0 covers severe misuse requests across chemical, biological, radiological, nuclear, and explosive (CBRNE) threats and offensive cybersecurity. In this report, we tested four leading models for universal jailbreaks in the context of this minimal standard, and found more than a hundredfold difference in security. Claude Fable 5 and GPT-5.6 Sol held against every attack we ran, with no universal jailbreak found; we estimate they would likely cost more than \$14,200 to jailbreak, if it is possible with this methodology at all. Meanwhile, we found hundreds of universal jailbreaks for Grok 4.5 and Gemini 3.1 Pro; each broke for under \$300, with universal jailbreaks in Grok's weakest domain, cybersecurity, accessible for as little as \$24. The gap is fixable: every weakness we found belongs to a known class of attack that already has a defense deployed in production models. The leaderboard will be updated on a rolling basis as new models are released, and the evaluation methodology and Minimal Standard will be periodically revised to take into account the latest capabilities and the state-of-the-art in safeguards. The leaderboard is available at leaderboard.far.ai.
SRE-Bench is introduced, the first realistic, contamination-free RE benchmark, and results indicate that strong source-code security capabilities do not yet transfer to binary analysis, highlighting RE as an important frontier for agentic cybersecurity and SRE-Bench as a rigorous testbed to measure progress.
J. Spence, Nicholas Assaderaghi, Feng Xiao et al.· 1 citation
A formal framework for security in agentic AI is developed, built around an ideal predicate ϕ that decides, for a given ( prompt, context, result ) triple, whether the agent’s behaviour is acceptable.
F. Durak, Tadayoshi Kohno, Franziska Roesner· 0 citations
Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark's measur...
Neeraj Karamchandani, Piyush Nagasubramaniam, Xin-Hong Xie et al.· 0 citations
Key Takeaways:
- What counts as relevant is being redefined by intelligentisation, where data quality determines AI effectiveness, making peacetime civilian data acquisition a form of military preparation.
- The establishment of ISF in April 2024 institutionalised the doctrinal shift of information gathering into int...
M. Descamps· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.