VMTech
Discuss a project

SaferAI finds GLM-5.2 near frontier capability but short on safeguards

SaferAI finds GLM-5.2 near frontier capability but short on safeguards

Z.ai's open-weight GLM-5.2 is only a few months behind OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7 in cyber and biological capabilities, a SaferAI report found. Yet in tests conducted through Z.ai's public API, GLM-5.2 refused none of the offensive cybersecurity or dual-use biology tasks it received.

Claude Opus 4.7 behaved very differently: its refusals were so consistent that SaferAI could not complete the CyberGym cybersecurity benchmark. The comparison highlights a widening divide between model capability and the mitigations surrounding it.

Why open weights change the risk calculation

Z.ai can apply controls to its hosted API, but those protections become unenforceable when users download the model weights and run them on their own hardware. Operators can modify safeguards, change system prompts or fine-tune the model without the provider retaining control.

Closed-model developers such as OpenAI and Anthropic typically combine refusal training, classifiers and API-level controls. These measures are imperfect: Far.ai identified hundreds of reusable jailbreaks affecting frontier systems including xAI's Grok 4.5 and Google DeepMind's Gemini 3.1 Pro. Attackers combined roleplaying, impersonated authority, fabricated conversation histories and follow-up prompts to amplify weaknesses.

The stakes were also illustrated by controlled AI development after the Hugging Face breach, which placed model capability and controlled development in the context of the Hugging Face breach. Open-weight systems create a distinct challenge because safeguards applied by a provider need not survive local deployment.

Mitigations remain technically difficult

SaferAI executive director Henry Papadatos said risk assessments must consider mitigations as well as frontier capability. One proposed measure is pre-training data filtering, which removes hazardous material before training. Research suggests that this can reduce dangerous biological knowledge without degrading overall performance.

Cybersecurity is harder to separate from legitimate capability. A general model that is strong at coding may also be useful for hacking, while coding is commercially important to AI developers. Providers therefore use additional controls, including restrictions on specific forms of assistance, pre-deployment evaluations, published risk assessments and decisions to withhold weights.

Anthropic's Opus 5, for example, can search for vulnerabilities in uncompiled source code but not compiled software, its system card says. SaferAI said Z.ai published no safety framework, pre-deployment testing commitments or risk assessment for GLM-5.2. Z.ai did not respond to TechCrunch's question about internal or third-party frontier safety evaluations.

Defensive value and governance trade-offs

Open-weight advocates argue that access helps defenders identify vulnerabilities and prepare for attacks. Hugging Face used GLM-5.2 in its response to OpenAI's breach, and chief executive Clem Delangue said comparable systems could help stop attacks and fix weaknesses before exploitation.

Papadatos cautioned that defensive value does not justify making every dangerous capability broadly available. He noted that attackers can adopt tools faster than institutions: a ransomware group may change methods within a week, while a hospital cannot.

For businesses evaluating advanced models, benchmark performance should therefore be reviewed alongside refusal behavior, release documentation and the enforceability of controls. Sensitive deployments need scoped access, activity monitoring and an explicit misuse plan, especially when model weights can leave the provider's infrastructure.

#openweightai#aisafety#cybersecurity#aimodels
Open analytics
On the site 1 views
min read 4 05.08.2026
Instagram

SaferAI finds GLM-5.2 near frontier capability but short on safeguards

Open the post on Instagram ↗