Chinese Open-Weight AI Model GLM-5.2 Closes Gap with Frontier Models on Capabilities but Lags on Safety
A new SaferAI evaluation found that China's open-weight AI model GLM-5.2, developed by Z.ai, has narrowed the capability gap with frontier closed-source models including OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7. However, a significant safety divergence persists. Unlike Claude Opus 4.7, which refused cyber and biological dual-use tasks at a high rate, GLM-5.2 refused none of the same offensive tasks during testing. The disparity highlights risks associated with open-weight model distribution: once weights are downloaded, users can remove safeguards, fine-tune models, or alter safety-focused system prompts. While frontier developers rely on classifiers, refusal training, and API-level controls to limit dangerous outputs, these protections are ineffective on open-weight models. SaferAI's executive director noted that pre-training data filtering—removing offensive cybersecurity information during model training—could reduce hazardous knowledge, though such techniques are less practical for cybersecurity training data.