CAISI finds GLM-5.3 the most cyber-capable open-weight model, about four months behind the US frontier

On September 17, 2026, the Center for AI Standards and Innovation (CAISI) at NIST, the US government’s AI evaluation body, published an assessment of the cyber capabilities of Z.ai’s GLM-5.3, an open-weight model released in August 2026. Its headline finding: “GLM-5.3 is the most cyber-capable open-weight model released to date,” but its cyber capabilities “are significantly lower than those of current U.S. frontier models,” lagging “by about four months in an aggregate measure of performance across CAISI cyber benchmarks.”

CAISI reported scores on four vulnerability discovery and exploitation benchmarks, each against the best US and best PRC model it had evaluated. On SEC-Bench Pro, where the model gets the source code of the V8 or SpiderMonkey JavaScript engines and must find a vulnerability and write code that triggers a crash, GLM-5.3 scored 40.4% against 90.2% for the US frontier best and 27.3% for the previous PRC best. On ExploitBench it scored 61.1% (US best 100.0%, PRC 32.2%), on ExploitGym 9.4% (44.4%, 2.6%) and on OSS-Fuzz 7.7% (23.2%, 2.4%). The prior PRC best named in the report is Moonshot’s Kimi K3.

To combine the benchmarks CAISI used item response theory (a one-parameter model) to build a “cyber capability index,” in which a 400-point increase equals a tenfold increase in the odds of solving a task. GLM-5.3’s index sits above Kimi K3 and below the current US frontier, with the gaps well outside 95% confidence intervals. CAISI notes it had previously assessed GLM-5.2.

Why it matters: this is the US government publishing model-by-model, benchmark-level measurements of a foreign open-weight model’s offensive cyber ability, and it shows open weights closing on the frontier within months rather than years - while also showing a PRC jump over Kimi K3 on every benchmark. What it does not show: the post does not name the US models it compares against, does not explain how the four-month figure was derived, and says it does not compare with models “developed but not yet released, which could have stronger capabilities.” Benchmark scores also measure capability under test conditions, not real-world misuse.