← ALL NEWS

SIMON WILLISON · 27 Aug 2026

Qwen3.8-Flash-Next

Qwen3.8-Flash-Next is a new open weights artificial intelligence model developed by Qwen. It is a multimodal mixture of experts model and serves as an early preview of the architecture that will be used in the upcoming Qwen4 series.

The model is quite large in total capacity, containing 125 billion parameters or tokens, but it activates only 6 billion parameters at a time. This design allows the system to achieve a significant performance boost.

Testing of the model has involved running Unsloth quantized versions on hardware such as a DGX Spark. Specifically, users have experimented with the 72.5-gigabyte and 78.9-gigabyte variants, including configurations set to high reasoning effort.

Releasing models with advanced multimodal capabilities and preview architectures like Qwen3.8-Flash-Next matters because it provides researchers and developers with early access to cutting-edge open weights technology, enabling experimentation with efficient mixture of experts systems and quantized setups on specialized hardware.

Read the original ↗