
Benchmarking the Open-Weight War: Chinese Models Now Match the Entire US Open Ecosystem
A new independent analysis finds Qwen, Kimi, DeepSeek and GLM dominating the open-weight frontier, with Moonshot's K2 Thinking at Intelligence Index 67 — and America's best open model trailing on knowledge.
How wide is the open-weight gap between China and the United States? A detailed new analysis from Understanding AI, drawing on Artificial Analysis's Intelligence Index and independent testing, puts numbers on what developers have sensed for months: Chinese labs now field the strongest open models at essentially every size class.
The scoreboard
At the top sits Moonshot AI's Kimi K2 Thinking, a 1-trillion-parameter system scoring 67 on the Intelligence Index — the highest of any open-weight model tested — with reviewers singling out its writing quality and ability to sustain long tool-calling chains. DeepSeek V3.2 follows at 66, with its Speciale variant topping the MathArena benchmarks outright. Z.ai's GLM 4.6 line (Intelligence Index 56 at 357B parameters) has become a favorite for coding, with international usage up roughly tenfold in two months.
Alibaba's Qwen family remains the ecosystem's backbone: models from edge-sized to 235B parameters scoring 43–57, and the most downloaded model family in the world. As researcher Nathan Lambert put it: "Qwen alone is roughly matching the entire American open model ecosystem today."
America's thin bench
The best US open model, OpenAI's gpt-oss-120b, scores a respectable 61 with exceptional serving speed — beyond 3,000 tokens per second at some providers — but craters on factual knowledge, scoring 6.7–16.8 percent on SimpleQA where Chinese rivals reach 50–70 percent. The Allen Institute's OLMo 3 stands alone as a truly open-source option, shipping full training code, data and checkpoints, but trails comparable Qwen models out of the box.
The structural problem is breadth: no US lab offers a coherent ladder of open models across sizes the way Qwen, GLM or DeepSeek do — and enterprise hesitancy about Chinese-origin weights only partially offsets that, since much of the world outside the US has no such compunction.
What the numbers mean
Two caveats temper the headline. Chinese open models still trail the best closed US systems on the hardest agentic tasks, and adoption barriers — security reviews, data-handling concerns flagged around DeepSeek, procurement politics — are real in Western enterprises. But the trendline is unambiguous, and with Kimi K3's weights scheduled to ship within days, the open-weight frontier is about to move again — from Beijing.
For researchers and startups everywhere, the practical conclusion is the analysis's most striking finding: value-for-performance in open AI is now overwhelmingly a Chinese story, and building on that stack is becoming the global default.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.
More in Research

Kimi K3's Weights Are Out: 2.8T Parameters, a 1.4TB Download, and an Escalation Nobody Can Undo

MeetingToM: A Benchmark for the Social Skill AI Keeps Failing — Telling Real Agreement From Fake
