Six days after developers began sending code to a mystery model called Ox Alpha, Z.ai gave it a name: GLM-5.3-Flash. OpenRouter said the free preview processed more than 20 trillion tokens in those six days, the most of any model on its service.

When Onset covered the reveal earlier Wednesday, the promised download had not appeared. It has now: Z.ai published reusable model files under the permissive MIT license and opened paid access through its API.

Two independent test suites put Flash within sight of GLM-5.3. Artificial Analysis scored it 57 against 60 for GLM-5.3 on a nine-test mix of coding, reasoning, knowledge and other work. DeepSWE measured 63% against 69% across 113 long software jobs, with the uncertainty ranges overlapping.

A chart of 59 models highlights GLM-5.3-Flash’s score and cost
via @alcreon

The tests do not show that the models feel the same in everyday use: Artificial Analysis had no speed result, found Flash unusually verbose and measured weaker factual knowledge.

What it costs

Z.ai’s regular API rate is 15 cents per million input tokens and 50 cents per million output tokens, about one-ninth of GLM-5.3’s rates. A launch sale halves those prices through Sept. 9.

For developers whose work suits it, that gap can cut API bills. The downloadable files add another route: organizations with enough computing power can run or adapt Flash themselves instead of sending every job through Z.ai.