Two Frontier Models Cracked Wide Open. Two Held. Now We Have Numbers.

Two Frontier Models Cracked Wide Open. Two Held. Now We Have Numbers.

The AI safety layer your enterprise deal depends on just got independently stress-tested — and the results are not a tie. For the first time, a public benchmark puts a dollar figure on how much it costs an attacker to break each major frontier model. Some of those numbers are embarrassingly small.

What happened

FAR.AI published the first public, quantified jailbreak leaderboard for frontier models, evaluating Claude Fable 5, GPT-5.6 Sol, Gemini 3.1 Pro, and Grok 4.5 against 360 attacker goals across CBRNE (chemical, biological, radiological/nuclear, explosive) threats and offensive cyber. They built a taxonomy of 67 static jailbreak techniques and used a three-stage funnel to hunt for “universal jailbreaks” — single prompt templates that succeed on over 75% of goals within a domain. The cost-to-jailbreak metric is the killer finding: random search found 63 universal jailbreaks against Grok 4.5 at roughly $58 per jailbreak, and 18 against Gemini 3.1 Pro at roughly $278 per jailbreak. Expert-guided composition raised those numbers to 385 and 231 jailbreaks respectively. Claude Fable 5 and GPT-5.6 Sol yielded zero universal jailbreaks under either strategy.

Cold read

“No universal jailbreak found” is a lower bound, not a certificate of safety — the paper explicitly right-censors these results, meaning Claude and GPT-5.6 Sol may have failure modes that a larger compute budget or smarter attack strategy would surface. The benchmark covers 360 attacker goals, which sounds thorough but is a sample of a very large attack space; coverage gaps are baked in by design. The evaluation is point-in-time: models are updated continuously, and a patch or a new fine-tune can shift these numbers overnight in either direction. The “operationally compliant” response standard used to score success is itself a judgment call — LLM-as-judge pipelines carry their own error rates, and the paper doesn’t detail inter-rater reliability on that scoring. Finally, the 67-technique taxonomy reflects publicly known attacks; nation-state or well-resourced attackers working with novel, unpublished techniques are outside scope entirely.

What it means for you

  • Signal maturity: 4/5 — Methodology is unusually rigorous for a security benchmark; dollar-cost framing is concrete and auditable
  • Who gets hurt: Any SaaS or API product that routes user input through Grok 4.5 or Gemini 3.1 Pro for high-stakes tasks (compliance, medical, defense-adjacent) and has been treating “the model has guardrails” as a line item on the risk register
  • What breaks if this is true: Enterprise buyers demanding model-specific security attestations before signing; the “we use a frontier model so we’re covered” compliance argument collapses for two of the four major providers
  • Why it might not land: Grok and Google will patch — the paper itself says these gaps are “closable with current techniques” — so by the time procurement cycles complete, the specific vulnerabilities may be remediated, muting urgency
  • Watch for: Whether enterprise AI procurement RFPs start requiring published jailbreak-resistance scores or third-party red-team attestations; that’s the market signal this benchmark found a foothold

Forecast as of 2026-08-06

By Q2 2027, at least two of the four evaluated providers will publish formal responses to the FAR.AI leaderboard methodology — either contesting the scoring criteria or announcing updated model versions with measurable improvement on the cost-to-jailbreak metric as independently verified. If neither Grok nor Gemini close the gap to “no universal jailbreak found” under the same protocol within 12 months, that will be a material procurement liability in regulated industries.


Source: AI Security Leaderboard: Methodology, Results and Minimal Standard — Jasper Timm, Lukas Struppek, Ziwei Xu, Grace Cheong, Oscar Mata, Dan Zhao, Mick Yang, Isadora De Andrade, Xiaojun Jia, Yiming Li, Samuel Bauer, Heather McIntyre, Adam Gleave, Edward Yee, Kellin Pelrine. https://arxiv.org/abs/2608.03070v1

Similar Posts