
Z.ai's GLM-5.3-Flash Pairs Sparse and Linear Attention in an Open-Source First — at $0.15 per Million Tokens
The MIT-licensed 320B model activates just 18B parameters per token, handles a 1M context natively across text and images, and cuts attention compute roughly 3x — undercutting frontier pricing by two orders of magnitude.
Z.ai's latest open-weight release is less notable for its benchmark scores than for its architecture — though the scores are strong too. GLM-5.3-Flash, released under the MIT license, is the first open-source frontier-class model to combine sparse attention with linear attention in a single hybrid design, a configuration aimed squarely at making million-token contexts economically viable.
The architecture
GLM-5.3-Flash is a 320-billion-parameter mixture-of-experts model that activates just 18 billion parameters per token. Unlike earlier Flash variants that grafted vision onto a text model, it is natively multimodal, trained from a new multimodal base with a 1M-token context window spanning text and images.
The hybrid attention scheme is the headline contribution: at long context, Z.ai reports roughly 3x lower attention compute and over 4x smaller KV cache versus the base GLM-5.3 — the two costs that dominate serving economics for agentic and long-document workloads.
Benchmarks and price
On Z.ai's DeepSWE software-engineering evaluation, GLM-5.3-Flash scores 63.4 against GLM-5.2's 46.2, with a larger margin on AutomationBench. Self-reported numbers deserve the usual skepticism, but early community testing has been notably positive — the model trended for days after appearing unannounced as a stealth checkpoint on evaluation platforms.
Then there is the price: $0.15 per million input tokens and $0.50 per million output, with cached input at $0.03. Against GPT-6 Astra's $10/$50 and Fable 5.1's identical frontier pricing, Z.ai is offering roughly 90 percent of the capability surface for coding and agentic work at around 1-2 percent of the cost.
The pattern
GLM-5.3-Flash extends the strategy that has defined China's leading labs all year: ship open weights fast, compete on serving efficiency rather than raw scale, and occupy the default slot in cost-sensitive stacks worldwide. With the GLM line already powering coding agents and now a Saudi national platform built on rival MiniMax's weights, the open-weight efficiency race is arguably where Chinese labs are most clearly ahead — and hybrid attention just raised the bar again.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.
More in Research

Google Ships Gemini 3.8 Flash and a 'Cyber' Twin That Hunts Vulnerabilities

Meta's Muse Spark 1.3 Reaches the Frontier: 75.4% on DeepSWE, 1M Context, and a Data-for-Discount Endpoint
