Onset

Z.AI Made Ox Alpha, Which Drew Twice DeepSeek V4 Flash’s Traffic

Z.AI Made Ox Alpha, Which Drew Twice DeepSeek V4 Flash’s Traffic

Z.AI says it will publish downloadable model files, which could let developers build with Ox Alpha beyond OpenRouter.

Illustration by Onset News: cover image for Z.AI made Ox Alpha, which drew twice DeepSeek V4 Flash’s traffic
Illustration · Onset News

Latest update

GLM-5.3-Flash arrives in AutoClaw

AutoClaw added GLM-5.3-Flash, the released name for Ox Alpha, in version 1.17.8, giving users another way to put the model to work on agentic tasks.

View 8 earlier updates

Ox Alpha triples Z.AI’s OpenRouter share

Ox Alpha accounted for 19% of all tokens routed through OpenRouter so far this week, Peter Walker’s analysis found, versus about 6% for Z.AI models in recent weeks.

Ox Alpha’s weights are now open

  • 320 billion total parameters, with 18 billion active per token
  • 34 Kimi Delta Attention layers and 11 DeepSeek Sparse Attention layers
  • 1 million-token context window

Ox Alpha’s preview ran on Chinese chips

Z.AI says the anonymous model, now GLM-5.3-Flash, ran entirely on Chinese AI chips. OpenRouter recorded 23.2 trillion Ox Alpha tokens from Aug. 20 through Aug. 25, more than twice DeepSeek V4 Flash’s 9.9 trillion. The preview therefore demonstrated large-scale serving capacity as well as demand for the model.

GLM-5.3-Flash enters Code Arena near the top

The model previewed as Ox Alpha scored 1,634 in Code Arena’s WebDev test, placing around fifth overall and second among open models.

The result is provisional. It came from AutoEval, which substitutes a reward model trained on Arena’s human-preference data for live votes. Arena said it will track whether the ranking holds as human votes arrive.

Z.AI releases Ox Alpha as GLM-5.3-Flash

  • GLM-5.3-Flash has 320 billion total parameters and activates 18 billion per request
  • The model is natively multimodal and supports a 1 million-token context window
  • Z.AI released the model under the MIT License and made its weights available

Z.AI discounts GLM-5.3-Flash API access

  • Input: $0.075 per 1 million tokens
  • Output: $0.25 per 1 million tokens
  • Cached input: $0.015 per 1 million tokens

Z.AI releases Ox Alpha as GLM-5.3-Flash

  • 320 billion total parameters, with 18 billion active
  • Native multimodal support and a 1 million-token context window
  • API price: $0.15 per 1 million input tokens

Z.AI releases Ox Alpha as GLM-5.3-Flash

  • 320 billion total parameters, with 18 billion active
  • 1 million-token context window and 131,000-token maximum output
  • Standard API price per 1 million tokens: $0.15 input, $0.50 output and $0.03 cached input

Full story

Developers had already spent six days handing code to Ox Alpha by the time they learned which company was on the other end. Onset first covered the model when it appeared Aug. 20 as a one-week free preview on OpenRouter, a marketplace for accessing many AI models. Its listing described it as built for coding and long software jobs.

By Wednesday, Ox Alpha’s weekly text volume on OpenRouter was about 2.1 times DeepSeek V4 Flash’s. The chart counts text sent to and returned by models, so it shows adoption on OpenRouter, not answer quality, request totals or user counts.

Z.AI, the Chinese AI company also known as Zhipu, told Bloomberg News that it made Ox Alpha and that the preview is a new model in its GLM family. The company said it would publish downloadable model files that night, a step that could let developers build products around Ox Alpha outside OpenRouter.

At the last check, no Ox Alpha download appeared in Z.AI’s official repositories; whether the promised release was complete remained unknown, and Z.AI had not disclosed the model’s size or license.

Z.AI has done this before

OpenRouter also carried Pony Alpha without a maker’s name in February. Its listing now identifies that preview as an early testing version of GLM-5.

Ox Alpha is the second GLM preview to follow that sequence: people put the model to work first, and Z.AI attached its name later.

The Conversation