Open-Weight AI Models Close Capability Gap as Safety Concerns Persist

A new SaferAI report says China’s GLM-5.2 is nearing frontier AI capabilities, but its lack of safety measures highlights growing concerns over open-weight models.

Aug 5, 2026 - 11:06
 2
Open-Weight AI Models Close Capability Gap as Safety Concerns Persist
IMAGE CREDITS: Z.AI

Open-weight artificial intelligence models are rapidly closing the performance gap with the industry’s most advanced closed systems, but a new report suggests safety practices are not keeping pace. According to AI safety nonprofit SaferAI, China’s GLM-5.2 model from Z.ai now trails leading systems such as OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 by only a matter of months in cyber and biological capabilities, while offering significantly fewer safeguards against misuse.

SaferAI evaluated GLM-5.2 through Z.ai’s public API and found that the model refused none of the offensive cybersecurity or dual-use biology tasks it was asked to perform. By contrast, Anthropic’s Claude Opus 4.7 consistently declined similar requests to the extent that SaferAI said it was unable to complete its CyberGym cybersecurity benchmark using the model.

The findings add new urgency to the debate over open-weight AI systems, which allow users to download and run model weights on their own hardware. While supporters argue that open-weight models accelerate research and defensive cybersecurity, critics warn that once the weights are released, there is little practical way to enforce safety controls or prevent malicious use.

Capability advances outpace safety measures

SaferAI Executive Director Henry Papadatos said the frontier of AI capability is no longer the only measure that matters. He argued that risk should also be evaluated based on the safeguards surrounding a model, noting that safety protections available through hosted APIs disappear once an open-weight model is run independently.

Unlike frontier AI developers such as OpenAI and Anthropic, which rely on refusal training, classifiers and API-level restrictions to limit dangerous requests, open-weight models can be modified by anyone. Users can remove safeguards, alter system prompts or fine-tune models without oversight.

Even those protections are not foolproof. AI safety organisation Far.ai has identified hundreds of reusable jailbreak techniques capable of bypassing restrictions in frontier models, including xAI’s Grok 4.5 and Google DeepMind’s Gemini 3.1 Pro. According to its research, attackers combine role-playing, authority impersonation, fabricated conversation history and carefully structured follow-up prompts to exploit weaknesses in model defences.

However, Papadatos noted that API safeguards at least provide a layer of protection for closed models, while open-weight releases eliminate those controls once downloaded.

Searching for practical safety solutions

One approach SaferAI highlighted is pre-training data filtering, where developers remove hazardous biological or cybersecurity content before training a model. Some studies suggest this can reduce dangerous biological knowledge without significantly affecting general performance.

Applying the same approach to cybersecurity is far more difficult. Modern AI systems are expected to excel at software development, and many of the skills that make a model a strong programmer also make it capable of identifying software vulnerabilities. Because coding has become one of AI’s most valuable commercial applications, developers face pressure to continue improving those abilities while limiting abuse.

Instead, many frontier AI companies have shifted toward deployment safeguards. Anthropic, for example, allows some defensive vulnerability analysis while restricting assistance involving compiled software, an approach intended to reduce offensive use. Other measures include extensive pre-deployment testing, published risk assessments and, when necessary, withholding model weights altogether.

SaferAI said Z.ai did not publish a formal safety framework, public risk assessment or documented pre-deployment testing commitments for GLM-5.2.

China’s AI strategy differs from the West

The report arrives as China continues promoting open-weight AI development. During lastmonth’ss World AI Conference, Chinese President Xi Jinping emphasised both the value of open-weight AI models and the importance of maintaining strict human oversight of artificial intelligence.

Graham Webster, a researcher at Stanford University’s Cyber Policy Centre, said China’s AI regulations have historically concentrated on politically sensitive content, misinformation and social stability rather than catastrophic risks such as offensive cyber operations or biological misuse. He added that many Chinese policy experts believe American AI companies are more likely to encounter novel frontier risks first because of their leading-edge research.

Webster also noted that China’s system of real-name internet use and regulatory oversight gives authorities greater confidence that AI technologies can be controlled domestically. At the same time, because Chinese AI companies often work closely with regulators behind closed doors, outside observers have limited visibility into the safety testing conducted before models are released.

Balancing openness and security

Supporters of open-weight AI argue that broader access also strengthens cybersecurity. Hugging Face relied on GLM-5.2 while responding to the autonomous AI cyberattack disclosed by OpenAI, and Hugging Face CEO Clem Delangue recently said publicly available models can help organisations detect vulnerabilities before attackers exploit them.

Papadatos acknowledged those benefits but argued they should not justify releasing dangerous capabilities without adequate safeguards. In his view, the industry should work toward making beneficial AI capabilities widely accessible while limiting access to functions that could significantly increase offensive cyber or biological risks.

He also cautioned that attackers often adopt new technologies much faster than defenders. While ransomware groups can quickly integrate new AI capabilities into their operations, hospitals, governments and large enterprises typically require far longer to update security systems and procedures, leaving a window where increasingly capable open-weight models may create new risks before corresponding defences are in place.

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Angry Angry 0
Sad Sad 0
Wow Wow 0
Shivangi Yadav Shivangi Yadav reports on startups, technology policy, and other significant technology-focused developments in India for TechAmerica.Ai. She previously worked as a research intern at ORF.