← ALL NEWS

AI NEWS (SMOL.AI) · 26 Aug 2026

Z.ai Launches GLM-5.3-Flash with 1M Context Window and Open Weights

Z.ai formally launched GLM-5.3-Flash, revealing that a previously previewed model known as Ox Alpha was its public identity. The natively multimodal, MIT-licensed model features a one-million-token context window and a mixture-of-experts architecture with 320 billion total parameters and 18 billion active parameters. It is available through weights on Hugging Face, an API, chat platforms, a coding plan, and AutoClaw integration.

Independent evaluations from Artificial Analysis place the model’s Intelligence Index score at 57, tying it with models like GPT-5.6 Terra while offering significantly lower task costs at nine cents per task. The model achieves high efficiency partly through very low token pricing and relies heavily on reasoning tokens. While it scores well on agentic tasks, coding, and terminal benchmarks, its real-world factual knowledge accuracy trails competitors, and independent practitioners note that its native vision capabilities struggle with specialized tasks like object detection.

Architecturally, the model utilizes a hybrid attention system combining linear and sparse attention layers, along with advanced residual paths to reduce active parameters and attention compute. Z.ai also stated that the model runs entirely on Chinese AI chips, drawing widespread industry attention to its inference serving scale and domestic hardware capabilities.

Adoption was immediate, with platforms like Cline reporting rapid traffic growth and infrastructure providers like CoreWeave and Baseten adding day-zero support. The launch highlights a broader trend among open Chinese labs toward hyper-efficient, long-context architectures that rival proprietary models on practical workflows while driving down costs.

Read the original ↗