Anthropic's Frontier Red Team tested a model anyone can download and found it nearly as good at building hacking tools as the model the company keeps behind a vetted-access program. In the report it published on September 29, 2026, Z.ai's open-weight GLM-5.3 built working exploits in 50 of 410 attempts on a test drawn from known flaws in the engine behind Google Chrome, while Anthropic's Claude Mythos Preview, offered only to vetted users, managed 56. Six attempts out of 410 separate a free download from a frontier model. The larger change is what came with the download: an editing job that cost Anthropic about $4,400 of computing stripped the model's refusals almost entirely.

What the two models did on the same test

An exploit is code that turns a flaw in software into access it was never meant to grant: a way in, a file read, a command that runs. ExploitBench scores that work step by step rather than as a single pass or fail, across 16 stages from reaching the vulnerable code to running your own code on the machine, and it is built on 41 bugs in V8, the JavaScript engine inside Chrome. Reaching a crash is routine for current models; getting all the way to running your own code is not.

On Anthropic's numbers, GLM-5.3 reached the end of that chain in 12 per cent of its attempts, against 14 per cent for Mythos Preview. The models that came before it barely registered: Kimi K3, DeepSeek-V4.1-Flash, Claude Opus 4.6 and the previous GLM-5.2 sat at or near zero on the same test. A second, internal Anthropic benchmark tells the same story on harder ground. Across 100 randomly selected tasks from a Google-run database of bugs in open-source software, GLM-5.3 completed a full control-flow hijack, taking over the path a program's code follows, in 4 per cent of trials, and Mythos Preview in 6 per cent; Claude Opus 4.6 and GLM-5.2 succeeded in none.

A working attack in about a day, and one for about $20

Benchmarks count attempts. Anthropic also let the model work on a target. Over about a day, with limited human attention, GLM-5.3 found several previously unknown flaws in a browser's JavaScript engine and chained them into a web page that, when visited, reads files from the visitor's computer, including an SSH private key.

The cheaper demonstration used the small version of the model, GLM-5.3-Flash, on bugs that were already public. It linked two known Chrome flaws, one of them tracked as CVE-2026-11645, into a reliable chain against an ARM64 machine that bypassed pointer authentication, a defence that makes hijacked memory addresses harder to use. That took 20 minutes of human attention and eight hours of model time, which at Z.ai's published API prices would have cost $20.40. Starting from a known flaw is easier than finding a new one, which is why the second figure is the smaller.

Why open weights change the arithmetic

Z.ai, the company formerly known as Zhipu AI, released GLM-5.3 on August 14, 2026, and made the model's weights public two weeks later. Weights are the file that makes a model run, which means anyone who downloads them can also edit them.

Anthropic tested how much work it takes to get past the model's refusals. Asked for the harmful task directly, GLM-5.3 declined. A false cover story that presented the work as an authorized exercise got it to engage 64 per cent of the time, and writing the opening of its reasoning for it raised that to 92 per cent. Editing the refusals out of the weights altogether, a technique called abliteration, reached 100 per cent. None of those approaches got Anthropic's safeguarded Claude models to carry out the same tasks.

Abliteration means removing the part of the weights that produces a refusal. Anthropic's team had not done it before and spent about 2,200 GPU hours, roughly $4,400 of computing; it estimates a team that already knew the technique would need about 600 GPU hours, or $1,200. The edit left the model's abilities largely intact while refusal rates on three public tests fell from above 90 per cent to about 3 per cent, 2 per cent and 12 per cent. Anthropic says several developers released abliterated versions of GLM-5.3 within days of its release.

The hardware the download does not include

A file is not a working tool. Tom's Hardware put the practical requirement at 306 gigabytes of video memory at the model's standard numeric precision, with a similar amount again to hold the conversation's context, which in practice means a cluster of eight Nvidia H200 accelerators and hardware costing into the hundreds of thousands of dollars. That barrier sits outside the model and inside the machine that runs it.

Anthropic's stated position is that defenders need models at least as good as the ones attackers use, and it sells such models. Gizmodo noted the commercial stake in a company warning about a free rival, and framed the dispute as a domestic argument about open and closed AI as much as a safety question. The pattern is older than this model: the same outlet reported that third-party researchers found two earlier Kimi models from China's Moonshot could be jailbroken for bioweapons instructions and assassination planning, prompting an internal investigation. What changed here is the capability level, not the shape of the argument.

Anthropic's counterpoint is that the same capability helps defenders. Trusted users working with Claude Mythos models through its access programs have found more than 10,000 vulnerabilities, according to the company.

What an independent assessor found

On September 17, 2026, NIST's Center for AI Standards and Innovation assessed GLM-5.3 as the most cyber-capable open-weight model released to date and placed it about four months behind the United States frontier on an aggregate score across four cyber benchmarks, including ExploitBench. Before GLM-5.3, Kimi K3 held that open-weight position.

Anthropic adds a qualification to the comparison: the assessor tested United States models with their cyber safeguards disabled where that applied, and the frontier it measures includes models released only to vetted users. Attackers cannot readily obtain those versions. Anyone can download GLM-5.3.

What the report asks for

Anthropic's team argues that governments should run safety testing on sufficiently capable models, including whatever succeeds GLM-5.3. The ability to build a working exploit is no longer gated by who is allowed to buy the model. By the report's own figures it is gated by hardware and by an editing job priced in the thousands of dollars, which are costs a determined group can plan for.

Edited by Dan Martens