Open-weight AI models are rapidly closing the gap with frontier systems, but the safety divide is widening, according to multiple reports released this summer.
🔑 Key Takeaways
- GLM-5.2 from Z.ai refused zero offensive cybersecurity or biology tasks
- Safety guardrails become unenforceable once open weights are downloaded
- Hundreds of universal jailbreaks identified in closed models
- Discussions underway at the White House to restrict open-weight models
GLM-5.2: Comparable Performance, Zero Refusals
GLM-5.2, the open-weight model developed by Chinese company Z.ai, is dangerously close to OpenAI and Anthropic systems in terms of cybersecurity and biological capabilities. The evaluation conducted by SaferAI, a nonprofit organization specializing in AI safety, reveals that GLM-5.2 refused none of the offensive tasks submitted through its public API. By comparison, Anthropic’s Claude Opus 4.7 « refused so consistently that SaferAI could not complete CyberGym on this model », the report states. CyberGym is a benchmark evaluating cybersecurity capabilities of AI systems.
« The capability frontier is not the risk frontier, and so we must take into account the state of mitigation measures to properly assess risks. »
Henry Papadatos, Executive Director of SaferAI

The Core Problem: Guardrails Become Unenforceable
Once downloaded and run on proprietary hardware, open-weight models escape all control. Whereas companies like OpenAI and Anthropic rely on classifiers, refusal training, and API-level controls, these protections become totally inapplicable when weights are run locally. This fundamental limitation represents a major challenge for the governance of these systems.
SaferAI notes that Z.ai has published no security framework, no pre-deployment testing commitments, and no risk assessment for GLM-5.2. This lack of transparency complicates the assessment of risks by users and regulators alike.
Closed Models Are Not Infallible Either
Far.ai, an AI safety organization, identified hundreds of universal jailbreaks in closed models such as xAI’s Grok 4.5 and Google DeepMind’s Gemini 3.1 Pro. These vulnerabilities exploit a combination of techniques: role-playing, authority impersonation, falsified conversation history, and complementary prompts. This discovery nuances the picture: security issues are not exclusive to open models.
Washington Mulls Over Restrictions
The debate is shifting to the regulatory arena. According to Nathan Lambert, an AI policy specialist, discussions are underway at the White House regarding a new executive order that could restrict open-weight models, particularly those of Chinese origin and government uses.
« The most likely action would be to ban or indefinitely postpone any open-weight model significantly exceeding the capability level of GPT 5.5, Claude Opus 4.8, or GLM-5.2. With the current capability gap, this should occur within the next six months. »
Nathan Lambert, AI Policy Specialist
This prospect worries open source advocates. In response to rumors of restrictions, Nvidia, Microsoft, Meta, and roughly twenty other companies published an open letter warning Washington against « premature restrictions ». Meanwhile, Dario Amodei, CEO of Anthropic, clarified his company’s position.
« Anthropic has never advocated for banning open-weight models as a category. »
Dario Amodei, CEO of Anthropic
Amodei, however, highlighted two main concerns: the risk that authoritarian regimes develop more powerful models than those of the United States for military or repressive purposes, and the risk of misuse for cyberattacks or biological attacks.
« Open-weight models potentially pose higher risk than closed models, because it is very difficult to apply guardrails or monitor their use, and once weights are published, they cannot be taken back. »
Dario Amodei, CEO of Anthropic
China’s Position and Its Limitations
China itself acknowledges these risks. At the World AI Conference, President Xi Jinping emphasized the importance of open-weight models while insisting on the need to keep AI under strict human control. However, according to Graham Webster, a Chinese AI policy specialist at the Stanford Cyber Policy Center, Chinese regulations have historically focused on politically sensitive content, disinformation, and social stability, rather than on catastrophic risks such as offensive cybersecurity capabilities and biological misuse.
« American AI analysts generally worry more about existential catastrophic risks than the Chinese community does. »
Graham Webster, Stanford Cyber Policy Center
He noted that many Chinese policy researchers believe that if an existential risk truly emerges, American companies will probably encounter it first.
Chinese Advances Illustrate the Dynamics
Stanford’s AI Index found that the performance gap between the United States and China has narrowed to a few percentage points. Moonshot AI launched Kimi K3, a model with 2.8 trillion parameters and a one-million-token context window. Alibaba unveiled Qwen3.8-Max, performing at 2.4 trillion parameters. These models represent the most powerful open-weight systems to date.
| Model | Company | Parameters | Origin |
|---|---|---|---|
| GLM-5.2 | Z.ai | Undisclosed | China |
| Kimi K3 | Moonshot AI | 2.8 trillion | China |
| Qwen3.8-Max | Alibaba | 2.4 trillion | China |
| Claude Opus 4.7 | Anthropic | Confidential | United States |
Lambert warns against a ban that would be counterproductive: « If models are not also banned in China, it is very easy for a malicious actor to still use the open-weight model in question, which negates the security potential. » He instead advocates that an American company publish a competitive open model to redirect the debate: « This will shift the focus from ‘only China builds open models via distillation’ to ‘we are all in this together’. »
Open Weights Are Not Enough, Argues Stanford HAI
On the open-weights advocates’ side, Clem Delangue, CEO of Hugging Face, argued that these models enable companies to defend against attacks. « The same systems that helped stop an AI-powered cyberattack can now help defend against millions of daily cyberattacks, while helping us identify and fix vulnerabilities before attackers exploit them », he stated on social media.
Papadatos, however, nuance: « The main point in my view is that we should not simply accept that dangerous capabilities are easily accessible to anyone anywhere. »
Stanford HAI agrees. James Landay, director of the institute, advocates for truly open-source models that go beyond simply publishing weights:
« Open weights answers the question ‘Can I run this?’ Open source answers ‘Can I trust this, improve it, and build the next thing on top of it?’ Currently, almost all companies, including American and Chinese ones, answer the first question but are far from the second. »
James Landay, Director of Stanford HAI
Conclusion: Three Measures for Anthropic
Amodei concluded with three measures Anthropic advocates to address these risks: preventing the sale of advanced chips to China, combating large-scale industrial distillation, and mandatory security testing for all sufficiently capable models, whether open or closed. These proposals reflect the complexity of the debate: balancing innovation, security, and geopolitics. The outcome of these discussions will likely determine the future of open-weight AI in the coming months.
Sources
- TechCrunch – Open-weight AI models catching up, safety gap remains
- Interconnects – 6 months to live for open models
- MindStudio – Open weight vs closed frontier models
- Stanford HAI – Open weight models aren’t enough
- Anthropic – Position on open weights models
- New York Times – Open weight AI
This article is published for informational and educational purposes. It does not constitute investment advice in any way. Do your own research (DYOR) before making any decisions.

