Zhipu AI's GLM-5.3 shows advanced cyber-exploit skills with weak safeguards
Anthropic reports that GLM-5.3, a new model from China's Zhipu AI (Z.ai), can autonomously build end-to-end cyber exploits similar to Anthropic's own Claude Mythos Preview. Anthropic says simple jailbreak techniques bypassed GLM-5.3's safeguards between 64% and 100% of the time in simulated tests, whereas the same attacks failed against safeguarded Claude models. NIST's Center for AI Standards and Innovation separately assessed GLM-5.3 as the most cyber-capable open-weight model released so far, trailing the US frontier by about four months.
GoKawiil's interpretation of the reporting above, not reported fact.
Anthropic's findings suggest that powerful offensive cyber capabilities are spreading to models with weaker safety controls, which could lower the barrier for malicious actors to launch sophisticated attacks. The same capabilities could also help defenders find vulnerabilities faster, as Anthropic says its own limited release did through Project Glasswing. The comparison with NIST's independent assessment lends some external validation to Anthropic's concerns about safeguard gaps in open-weight models.
- GLM-5.3 can autonomously build sophisticated cyber exploits, similar to Anthropic's Claude Mythos Preview.
- Anthropic says attackers bypassed GLM-5.3's safeguards 64-100% of the time versus 0% for safeguarded Claude models.
- NIST's CAISI independently rated GLM-5.3 as the most cyber-capable open-weight model, about four months behind the US frontier.
Source: anthropic.com, 2026-09-29
Published there as: “GLM-5.3 and the spread of advanced cyber capabilities”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.