Anthropic finds open-weight GLM-5.3 nearly matches Claude Mythos Preview at writing exploits

On September 29, 2026, Anthropic published “GLM-5.3 and the spread of advanced cyber capabilities” by Andrew Fasano, Marius Fleischer, Cole McFaul, Robert Xiao and Tripp Gallagher. The team evaluated GLM-5.3, the open-weight model from Zhipu AI, on offensive security tasks and on how well its safeguards hold. Their conclusion is that it has autonomous exploitation ability comparable to Anthropic’s own restricted Claude Mythos Preview, but ships without meaningful safeguards.

On ExploitBench, built on Chrome V8 bugs, GLM-5.3 produced working end-to-end exploits in 12 percent of 410 attempts against 14 percent for Mythos Preview. On an internal binary exploitation benchmark it succeeded in 4 percent of 100 trials against 6 percent for Mythos Preview, where earlier models scored 0. The researchers used GLM-5.3 to find previously unknown vulnerabilities in a browser JavaScript engine and chain them into an exploit that stole SSH keys, and GLM-5.3-Flash built an N-day exploit for CVE-2026-11645 with 20 minutes of human attention and 8 hours of model work for 20.40 dollars. On safeguards, abliteration, which removes the refusal direction from the weights, cut refusal rates from 95 percent to 6 percent on JailbreakBench and HarmBench and 14 percent on StrongREJECT, at a cost of about 2,200 GPU hours, roughly 4,400 dollars. Even without it, deceptive prompts got 64 percent engagement and prefilled reasoning 92 percent; Anthropic reports its Claude models stayed at 0 percent under all tested conditions.

Why it matters: it complements the earlier US government CAISI assessment of the same model with a lab’s direct comparison to a frontier model held back for cyber risk. The gap between restricted frontier cyber capability and what anyone can download looks small, and refusal training in open weights can be removed for a few thousand dollars.

What it does not show: this is a competitor evaluating a rival model, and the comparison to Claude’s refusals was run by Anthropic itself. The authors note that their simulated environments do not perfectly represent real-world conditions, and success rates in the low teens are far from reliable autonomous attack.