A new report from the nonprofit organization FAR.AI has revealed how easy it is to bypass the safety guardrails of some top frontier AI models. The test covered models from four major U.S. companies: Anthropic (Claude Opus 4.8 and Fable 5), OpenAI (GPT 5.5 and 5.6), Google (Gemini 3.1 Pro), and SpaceXAI (Grok 4.3 and 4.5). Using an automated tool that generates over a thousand variations of problematic prompts, researchers attempted to make the models perform potentially harmful actions, such as generating software exploits or providing details for chemical and biological weapons.
Test results 448 jailbreaks for Grok 249 for Gemini
The results show a clear disparity in the robustness of the safeguards. Grok, SpaceXAI's model, proved the most vulnerable with 448 successful jailbreaks. Google's Gemini followed with 249 violations. In contrast, Anthropic's Claude and Fable, and OpenAI's GPT models resisted all attacks without any jailbreak. However, FAR.AI warns that these models are not immune to more sophisticated techniques involving complex interactions.
Sponsored Protocol
The negligible cost of breaking into AI models
The most alarming aspect is the incredibly low cost to conduct these attacks. According to the report, it took only $58 to jailbreak Grok and $278 to jailbreak Gemini. This shows that the economic barriers to malicious activity are virtually nonexistent. Adam Gleave, CEO of FAR.AI, commented that AI models right now are less regulated than restaurants. Gleave emphasizes the need for externally imposed standards and regulations, dismissing the idea of voluntary self-regulation.
Company responses between self-criticism and defense
The involved companies responded with defensive statements. Rohin Shah, director of AGI safety at Google DeepMind, clarified that the results should not be interpreted as a comprehensive assessment of Gemini's safety, as not all jailbreaks are equally severe. Google states it continuously works on improving safeguards. Anthropic, through spokesperson Michael Aciman, highlighted their sustained investments in safety systems, adapting to evolving attacks. OpenAI and SpaceXAI did not respond to requests for comment.
Sponsored Protocol
The regulatory landscape between state laws and federal chaos
The debate on AI safety occurs within a fragmented regulatory framework. States like California and New York have passed laws requiring developers to publish safety reports, while Illinois will soon demand third-party audits. At the federal level, no specific requirements exist. In June, the Trump administration imposed export controls on Anthropic's Fable 5 and Mythos 5 models, causing a weeks-long suspension. The White House also asked Anthropic and OpenAI to delay model releases over cybersecurity concerns.
A recent executive order promotes public-private collaboration on cybersecurity initiatives. However, for now, preventing major catastrophes largely falls on model makers. The potential for AI to cause harm became evident after OpenAI models autonomously hacked a popular code repository. A Cambridge University report documented that members of Boko Haram in Nigeria used ChatGPT, Claude, Gemini, Grok, Meta AI, and DeepSeek to plan violent attacks.
Sponsored Protocol
Expert voices between concern and optimism
Stephen Casper, a computer scientist at Harvard, predicts that serious incidents involving AI misuse are months away, not years. According to him, such an event will likely involve a system lacking state-of-the-art safeguards. Anka Reuel, a computer scientist at Stanford specializing in AI policy, believes the key takeaway from the FAR.AI report is clear: the safety measures used by Anthropic and OpenAI should become the default for all models. Some companies clearly know how to defend against at least the subset of attacks tested in this report, Reuel says. The question is why some companies are using them and others are not.
For further reading on AI safety, check the Wikipedia page on AI safety. Also, the smart glasses market is expanding: Meta smart glasses sales double as Samsung prepares a rival launch.
Source: https://www.wired.com/story/jailbreaking-ai-models-google-anthropic-openai-spacexai