Open-Weight AI Models Narrow the Gap — But the Safety Gap Keeps Growing

Share

Open-weight AI models are rapidly closing the gap with frontier systems, but the safety divide is widening, according to multiple reports released this summer.

🔑 Key Takeaways

  • GLM-5.2 from Z.ai refused zero offensive cybersecurity or biology tasks
  • Safety guardrails become unenforceable once open weights are downloaded
  • Hundreds of universal jailbreaks identified in closed models
  • Discussions underway at the White House to restrict open-weight models

GLM-5.2: Comparable Performance, Zero Refusals

GLM-5.2, the open-weight model developed by Chinese company Z.ai, is dangerously close to OpenAI and Anthropic systems in terms of cybersecurity and biological capabilities. The evaluation conducted by SaferAI, a nonprofit organization specializing in AI safety, reveals that GLM-5.2 refused none of the offensive tasks submitted through its public API. By comparison, Anthropic’s Claude Opus 4.7 « refused so consistently that SaferAI could not complete CyberGym on this model », the report states. CyberGym is a benchmark evaluating cybersecurity capabilities of AI systems.

« The capability frontier is not the risk frontier, and so we must take into account the state of mitigation measures to properly assess risks. »

Henry Papadatos, Executive Director of SaferAI

The Core Problem: Guardrails Become Unenforceable

Once downloaded and run on proprietary hardware, open-weight models escape all control. Whereas companies like OpenAI and Anthropic rely on classifiers, refusal training, and API-level controls, these protections become totally inapplicable when weights are run locally. This fundamental limitation represents a major challenge for the governance of these systems.

SaferAI notes that Z.ai has published no security framework, no pre-deployment testing commitments, and no risk assessment for GLM-5.2. This lack of transparency complicates the assessment of risks by users and regulators alike.

Closed Models Are Not Infallible Either

Far.ai, an AI safety organization, identified hundreds of universal jailbreaks in closed models such as xAI’s Grok 4.5 and Google DeepMind’s Gemini 3.1 Pro. These vulnerabilities exploit a combination of techniques: role-playing, authority impersonation, falsified conversation history, and complementary prompts. This discovery nuances the picture: security issues are not exclusive to open models.

Washington Mulls Over Restrictions

The debate is shifting to the regulatory arena. According to Nathan Lambert, an AI policy specialist, discussions are underway at the White House regarding a new executive order that could restrict open-weight models, particularly those of Chinese origin and government uses.

« The most likely action would be to ban or indefinitely postpone any open-weight model significantly exceeding the capability level of GPT 5.5, Claude Opus 4.8, or GLM-5.2. With the current capability gap, this should occur within the next six months. »

Nathan Lambert, AI Policy Specialist

This prospect worries open source advocates. In response to rumors of restrictions, Nvidia, Microsoft, Meta, and roughly twenty other companies published an open letter warning Washington against « premature restrictions ». Meanwhile, Dario Amodei, CEO of Anthropic, clarified his company’s position.

« Anthropic has never advocated for banning open-weight models as a category. »

Dario Amodei, CEO of Anthropic

Amodei, however, highlighted two main concerns: the risk that authoritarian regimes develop more powerful models than those of the United States for military or repressive purposes, and the risk of misuse for cyberattacks or biological attacks.

« Open-weight models potentially pose higher risk than closed models, because it is very difficult to apply guardrails or monitor their use, and once weights are published, they cannot be taken back. »

Dario Amodei, CEO of Anthropic

China’s Position and Its Limitations

China itself acknowledges these risks. At the World AI Conference, President Xi Jinping emphasized the importance of open-weight models while insisting on the need to keep AI under strict human control. However, according to Graham Webster, a Chinese AI policy specialist at the Stanford Cyber Policy Center, Chinese regulations have historically focused on politically sensitive content, disinformation, and social stability, rather than on catastrophic risks such as offensive cybersecurity capabilities and biological misuse.

« American AI analysts generally worry more about existential catastrophic risks than the Chinese community does. »

Graham Webster, Stanford Cyber Policy Center

He noted that many Chinese policy researchers believe that if an existential risk truly emerges, American companies will probably encounter it first.

Chinese Advances Illustrate the Dynamics

Stanford’s AI Index found that the performance gap between the United States and China has narrowed to a few percentage points. Moonshot AI launched Kimi K3, a model with 2.8 trillion parameters and a one-million-token context window. Alibaba unveiled Qwen3.8-Max, performing at 2.4 trillion parameters. These models represent the most powerful open-weight systems to date.

ModelCompanyParametersOrigin
GLM-5.2Z.aiUndisclosedChina
Kimi K3Moonshot AI2.8 trillionChina
Qwen3.8-MaxAlibaba2.4 trillionChina
Claude Opus 4.7AnthropicConfidentialUnited States

Lambert warns against a ban that would be counterproductive: « If models are not also banned in China, it is very easy for a malicious actor to still use the open-weight model in question, which negates the security potential. » He instead advocates that an American company publish a competitive open model to redirect the debate: « This will shift the focus from ‘only China builds open models via distillation’ to ‘we are all in this together’. »

Open Weights Are Not Enough, Argues Stanford HAI

On the open-weights advocates’ side, Clem Delangue, CEO of Hugging Face, argued that these models enable companies to defend against attacks. « The same systems that helped stop an AI-powered cyberattack can now help defend against millions of daily cyberattacks, while helping us identify and fix vulnerabilities before attackers exploit them », he stated on social media.

Papadatos, however, nuance: « The main point in my view is that we should not simply accept that dangerous capabilities are easily accessible to anyone anywhere. »

Stanford HAI agrees. James Landay, director of the institute, advocates for truly open-source models that go beyond simply publishing weights:

« Open weights answers the question ‘Can I run this?’ Open source answers ‘Can I trust this, improve it, and build the next thing on top of it?’ Currently, almost all companies, including American and Chinese ones, answer the first question but are far from the second. »

James Landay, Director of Stanford HAI

Conclusion: Three Measures for Anthropic

Amodei concluded with three measures Anthropic advocates to address these risks: preventing the sale of advanced chips to China, combating large-scale industrial distillation, and mandatory security testing for all sufficiently capable models, whether open or closed. These proposals reflect the complexity of the debate: balancing innovation, security, and geopolitics. The outcome of these discussions will likely determine the future of open-weight AI in the coming months.

Sources

This article is published for informational and educational purposes. It does not constitute investment advice in any way. Do your own research (DYOR) before making any decisions.

Telemac
Telemachttp://cryptoinfo.ch
Passionné de nouvelles technologies, j’explore l’univers de la blockchain et des cryptomonnaies pour partager l’actualité et les innovations du secteur.

Lire la Suite

Articles