New tool easily jailbreaks safeguards of frontier AI models, raising alarm for crypto and tech sectors

1 hour ago 22

The safety guardrails on the world’s most powerful AI models are, to put it gently, not holding up. A Nature study published in July 2026 found that four large reasoning models, deployed as adversarial attackers, achieved a 97.14% overall jailbreak success rate against nine frontier AI models from companies including OpenAI, Anthropic, and Google.

How the guardrails crumbled

The research tested four large reasoning models as autonomous adversaries: DeepSeek-R1, Gemini 2.5 Flash, Grok 3 Mini, and Qwen3 235B. These weren’t sophisticated nation-state tools. They used simple prompts and strategic persuasion techniques in multi-turn conversations to bypass safety filters.

DeepSeek-R1 was the standout performer, if you can call it that. It recorded a 100% attack success rate on HarmBench prompts in Cisco-linked testing. Every single prompt got through. The targets, models from OpenAI, Google, and Anthropic, showed only partial resistance at best.

Separate research focusing specifically on Claude models found that even advanced jailbreak techniques caused only a 7.7% performance degradation in the highest-performing variants.

Security firm HiddenLayer has documented universal bypass techniques that work across multiple leading large language models, including GPT-4, Claude, and Gemini.

The commercialization problem

A Russian-speaking threat actor operating under the name “Trim” began commercializing jailbreak techniques between March and June 2026, packaging them into a paid offensive security platform. This isn’t academic research being responsibly disclosed. It’s an exploit-as-a-service business model.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article