Onset

Alibaba Releases a Downloadable Preview of Qwen4’s Design

Alibaba Releases a Downloadable Preview of Qwen4’s Design

The experimental files arrive before the hosted production model, putting the first usable version in the hands of server operators.

Promotional Alibaba Qwen graphic on a black-and-electric-blue abstract background. Large text reads “Qwen3.8-Flash Now Available,” with the subheading “Qwen3.8-Flash-Next: Now Open Weights,” alongside
via @Alibaba_Qwen

Latest update

Qwen opens the production API

Qwen3.8-Flash is now live on QwenCloud, ending the “coming soon” status that applied when Alibaba released the downloadable Flash-Next architecture preview. The hosted model costs $0.16 per million input tokens and $0.47 per million output tokens, giving developers access without running the data-center hardware required to deploy the open weights themselves.

Full story

Downloading the smaller Qwen3.8-Flash-Next package means pulling 186 gigabytes before the model answers a single prompt. Qwen’s validated setup for it starts with two Nvidia GB300 AI accelerators, hardware made for data centers rather than an ordinary laptop.

Alibaba’s Qwen team posted that package and a 360GB version on Aug. 26. vLLM and SGLang, two systems for running AI models, released launch-day deployment instructions, so organizations with suitable servers can inspect and run the model themselves.

This is not Qwen4. Qwen calls it an experimental preview of the architecture planned for that family; the hosted Qwen3.8-Flash production model is still forthcoming. Although the preview contains 176 billion parameters, only 6 billion, about 3%, activate for each piece of text.

Qwen’s benchmark results

Qwen’s table puts Flash-Next ahead of Claude Opus 4.6 Max on eight of the nine benchmarks where both models had a score, including tests of software engineering, office work, instruction following and scientific reasoning.

Qwen’s benchmark table comparing Flash-Next with four other models
via x.com

Whether it is genuinely competitive with frontier systems remains unknown because Qwen ran the published tests, Arena’s launch score used automated judging, and neither live Arena nor an independent DeepSWE evaluation includes it yet.

The weights, downloadable model files, are available for developers to examine and self-host. Companies offering model services or coding and office assistants need separate commercial permission under Qwen’s license.

The Conversation