A study published in Nature in July 2026 revealed that four large reasoning models, deployed as adversarial attackers, achieved a 97.14% success rate in bypassing safety guardrails on nine frontier AI models from OpenAI, Anthropic, and Google. DeepSeek-R1 recorded a 100% success rate on HarmBench prompts in Cisco-linked testing. The targeted models, including Claude, showed only a 7.7% performance degradation against the most advanced jailbreak techniques. A Russian-speaking threat actor operating under the name « Trim » has been commercializing these techniques as a paid offensive security platform since March 2026, turning these vulnerabilities into an exploit-as-a-service business.
Source: Read the original article

