
DeepSeek V4 Stable Release Lands July 24 — and the Open Trillion-Parameter Class Gets Crowded
DeepSeek's V4 stable build arrives this week, squaring off against Kimi K3 and GLM-5.2 in a three-way fight among open trillion-scale MoE models — with benchmark parity against closed frontiers at a fraction of the serving cost.
DeepSeek is expected to ship the stable release of V4 on July 24, graduating its latest flagship out of preview in direct answer to Moonshot's Kimi K3 — and completing the most consequential open-model class ever assembled: three Chinese trillion-scale mixture-of-experts systems, all with published weights, all benchmarking against the closed frontier.
The trillion-scale trio
The class portrait, per benchmark roundups including MarkTechPost's comparison: DeepSeek V4 Pro at roughly 1.6 trillion parameters, Moonshot's Kimi K3 at approximately 2.8 trillion total parameters with a 1-million-token context window, and Zhipu's GLM-5.2, which continues to top open-source coding and agentic leaderboards. All three are sparse MoE designs activating a small fraction of parameters per token — the architecture that makes trillion-scale training and serving economically survivable.
What separates them is emphasis. K3 is built for long-horizon agentic coding, GLM-5.2 optimizes tool use and agent workflows, and V4's calling card is raw efficiency: frontier-adjacent output currently priced around $0.87 per million tokens, and demonstrated ability to run on Huawei-made processors — a detail with as much geopolitical weight as any benchmark.
Why the 'stable' label matters
Preview releases win Twitter; stable releases win procurement. Enterprises and inference platforms have been holding deployment decisions until V4 exits preview with fixed weights, a stable API surface and a license they can send to legal. The July 24 date also lands three days before Moonshot publishes K3's weights on July 27 under a modified MIT license — a sequencing that looks less like coincidence than a coordinated bid for the world's self-hosting budget in a single week.
The demand is demonstrably there: Moonshot suspended new K3 subscriptions after launch demand swamped its serving capacity, meaning open weights are now the primary channel for global access.
The research takeaway
A year ago, the open-versus-closed debate was about whether open models could reach the frontier. The trillion-scale trio has ended that argument and started a harder one: whether closed labs can maintain any capability premium at all once open MoE systems match them within months of every release. With US benchmark leadership now measured in weeks and priced at dollars per million tokens, the interesting research question has shifted from capability to efficiency — tokens per watt, per dollar, per domestic chip. On that axis, the July 24 release is less a model launch than a market repricing.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.


