TL;DR: NIST's CAISI unit published on 17 September 2026 that Z.ai released GLM-5.3 on 14 August and open weights two weeks later. On CAISI's cyber suites GLM-5.3 leads prior Chinese open models but trails U.S. frontier systems tested with cyber safeguards disabled, versions not freely downloadable. Example scores: SEC-Bench Pro 40.4 percent (74 of 183 tasks) versus 90.2 percent for U.S. models and 27.3 percent for prior PRC open weights; ExploitBench 61.1 percent (9.8 of 16 best-of-three) versus 100 percent U.S. and 32.2 percent prior PRC; ExploitGym 9.4 percent versus 44.4 percent U.S.; OSS-Fuzz 7.7 percent versus 23.2 percent U.S. Do not mix Anthropic's 50-of-410 ExploitBench count with CAISI's 61.1 percent figure; they are different harnesses.
Status note: Checked 5 October 2026 against the NIST CAISI news release of 17 September 2026. Benchmarks move when vendors patch models or when labs update tasks. U.S. scores reflect evaluation builds with safeguards off, not necessarily what every API customer receives.
What CAISI is measuring
Congress stood up CAISI inside the National Institute of Standards and Technology to track how artificial intelligence models perform on standardized tasks, including cybersecurity scenarios relevant to national security and critical infrastructure. When a Chinese lab publishes open weights, CAISI can run the same tests U.S. frontier labs face, then publish the gap in plain numbers.
That mission differs from a vendor's internal red team. Anthropic's late-September 2026 blog post tested GLM-5.3 with its own jailbreak suites and reported scores such as 50 successful tasks out of 410 on its ExploitBench variant. CAISI's 17 September release uses government benchmark names and reports ExploitBench as 61.1 percent on a best-of-three subset of 16 items. The two percentages answer related questions but are not interchangeable copy-paste stats.
GLM-5.3 in the open-weight lane
Z.ai, the brand tied to Zhipu AI, shipped GLM-5.3 on 14 August 2026 and released downloadable weights about two weeks later, CAISI noted. Open-weight release means anyone with enough GPUs can run the model locally, wrap it in hacking assistants, or fine-tune away refusals without asking Beijing or San Francisco for permission.
CAISI labeled GLM-5.3 the strongest open-weight model it had evaluated for cyber use to date, a statement about the public download class, not about every closed U.S. system. Closed frontier models often stay behind APIs with monitoring, terms of use, and safety filters. CAISI still tests some U.S. models with cyber safeguards turned off for comparison, but those builds are not the ones hobbyists torrent alongside GLM-5.3.
Scoreboard highlights readers should know
On SEC-Bench Pro, which stresses security reasoning across 183 tasks, GLM-5.3 scored 40.4 percent, clearing 74 tasks. U.S. frontier models in the same table hit 90.2 percent, while earlier open-weight models from China averaged 27.3 percent in CAISI's comparison set. That jump shows rapid progress in the open PRC tier even with a large remaining gap to U.S. scores.
ExploitBench in the CAISI release is reported as 61.1 percent using a best-of-three run on 16 challenge items, against 100 percent for the U.S. frontier reference and 32.2 percent for prior PRC open weights. ExploitGym, which tests interactive exploitation skills, landed at 9.4 percent for GLM-5.3 versus 44.4 percent U.S. and 2.6 percent prior PRC open. OSS-Fuzz style fuzzing tasks scored 7.7 percent versus 23.2 percent U.S. and 2.4 percent prior PRC open.
Aggregating those suites, CAISI placed GLM-5.3 about four months behind the U.S. frontier on its internal item response timeline index. That lag is a lab estimate, not a calendar law, but it gives policymakers a shorthand for how fast open downloads are catching specialized cyber skills.
What this report is not
CAISI did not accuse Z.ai of committing a specific intrusion, did not order a takedown, and did not criminalize downloading weights. The release is an assessment meant to inform export controls, critical infrastructure defense, and diplomatic conversations about model release norms.
Headlines that treat the assessment as proof GLM-5.3 already powers a named breach go beyond the document. Likewise, posts that cite Anthropic's 50-of-410 ExploitBench result and CAISI's 61.1 percent ExploitBench line in the same breath without explaining different test sizes mislead readers. Keep each lab's harness in its own lane.
Where things stand
Washington now has a published government scoreboard showing GLM-5.3 leading open-weight cyber capability while trailing safeguard-off U.S. frontier models by a few months on CAISI's composite index. Vendors can patch, retrain, or pull models; CAISI can refresh scores. Open weights mean defenders assume copies already left the lab.
Watch for Commerce Department follow-ups, whether allies adopt the same benchmarks, and if Z.ai's next release closes the SEC-Bench or ExploitGym gaps without new safety tooling. For now, the actionable read is defensive: treat downloadable GLM-5.3 as a capable scripting assistant in adversary hands, and read CAISI's tables separately from Anthropic's red-team blog when you compare percentages.
Sources: NIST CAISI, 17 September 2026.