
HUMAIN's humain-m3: A 428B-Parameter Arabic Frontier Model Trained on a Trillion Arabic Tokens
Saudi Arabia's state-backed lab unveils a mixture-of-experts model built on MiniMax's M3 lineage that tops seven public Arabic benchmarks with an 89.37% average — now in research preview on HUMAIN Node.
The most capable Arabic language model ever published did not come from Silicon Valley or Beijing — though it owes something to both. HUMAIN, Saudi Arabia's PIF-backed AI company, unveiled humain-m3 at LEAP in Riyadh: a 428-billion-parameter mixture-of-experts model now available in research preview through HUMAIN Node, the company's model-access platform.
Built on MiniMax, remade in Arabic
Architecturally, humain-m3 descends from Chinese lab MiniMax's open M3 lineage. HUMAIN's contribution is the data and the continued pre-training: more than one trillion tokens of Arabic-native content, spanning dialects, classical texts, and domain corpora that generic frontier models see only in trace amounts.
The results, per HUMAIN's evaluation, are state-of-the-art for the language: an average of 89.37 percent across seven public Arabic benchmarks — the highest among the frontier models tested — with particular strength in Arabic language understanding and reasoning. As always with self-reported figures, independent replication is pending; the research preview exists partly to invite it.
The sovereign-AI template
Beyond the scores, humain-m3 is a template for how mid-sized powers may build "sovereign" AI without training from scratch: take a permissively licensed open base (increasingly Chinese), invest where you hold a unique advantage (national language data), and deploy on domestic infrastructure (HUMAIN's Nvidia- and AMD-powered data centers). The full stack costs a fraction of a frontier training run and yields a model that beats the giants on the benchmarks a nation actually cares about.
For the research community, the open question is depth: whether trillion-token continued pre-training on a linguistically distant base transfers reasoning ability, or mostly relocates surface fluency. Arabic — morphologically rich, diglossic, and underserved by web-scale corpora — is a demanding test case, which is exactly what makes humain-m3's benchmark claims worth scrutinizing.
The geopolitics are unsubtle: a US-allied kingdom shipping its national model on Chinese architecture, days after Washington moved again to wall off Chinese labs from American compute. In model weights, as in oil, Riyadh is keeping its options open.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.
More in Research

Google Ships Gemini 3.8 Flash and a 'Cyber' Twin That Hunts Vulnerabilities

Meta's Muse Spark 1.3 Reaches the Frontier: 75.4% on DeepSWE, 1M Context, and a Data-for-Discount Endpoint
