
Unbundling the Token Factory: What AMD and Cerebras' Split-Inference Bet Means for the Nvidia Economy
The July 23 AMD-Cerebras pact splits inference into specialist stages and claims up to 5x tokens per watt — a sign that the metric of the buildout has shifted from chips to electricity, with real consequences for Asia's memory makers and model labs.
On July 23, at AMD's Advancing AI event, Lisa Su and Cerebras chief Andrew Feldman announced something that would have sounded like plumbing a year ago: a "disaggregated" inference system in which AMD's Helios racks handle one half of answering an AI query and Cerebras' dinner-plate-sized Wafer-Scale Engine handles the other. The headline claim — up to 5x more tokens per second per watt — comes with a footnote worth reading. But the architecture itself is the story, because it marks the moment the inference buildout stopped being about buying the best chip and started being about industrial process engineering.
Why inference is being cut in two
Serving a large language model has two phases with opposite physics. Prefill — digesting the prompt and its context — is compute-bound: it wants raw FLOPS. Decode — generating the answer token by token — is memory-bandwidth-bound: it wants data moved fast, not math done fast. Running both on the same GPU means each phase wastes what the other needs.
The AMD-Cerebras split assigns prefill and long-context processing to Helios, AMD's rack-scale system, and hands decode to Cerebras' WSE, whose on-wafer SRAM makes it the fastest token generator in production. Cerebras will deploy Helios in its own data centers, with the joint offering reaching customers through Cerebras Cloud in the second half of 2026.
About that 5x: the fine print says AMD Performance Labs and Cerebras modeled tokens per second per watt against a Cerebras-only configuration — not against Nvidia — using Moonshot AI's trillion-parameter Kimi 2.6 as the workload. It is a projection, not a benchmark, and it flatters the comparison. The more aggressive competitive claim is separate: AMD says Helios delivers up to 30 percent more inference tokens per dollar than Nvidia's Vera Rubin NVL72 rack.
Nvidia already agreed with the thesis
Here is what should worry AMD's cheerleaders: disaggregation is not a rebel doctrine. Nvidia's Rubin CPX — a prefill-specialized chip that swaps expensive HBM for cheap GDDR7 memory, arriving late 2026 — is built on exactly the same insight. Nvidia's own architecture pairs CPX racks for prefill with VR200 NVL72 racks for decode, orchestrated by its Dynamo scheduler, and the company claims the platform can return $5 billion in token revenue on $100 million of capex.
So the argument is not whether inference gets unbundled — everyone now says yes. It is whether unbundling erodes the premium of an integrated stack. Nvidia holds an estimated 80 to 90 percent of the AI data-center market, defended by CUDA and NVLink, moats built for training. But once inference is decomposed into commodity-shaped stages — a throughput stage here, a latency stage there, a KV cache passed between them — each stage becomes separately contestable. AMD takes prefill with tokens-per-dollar; Cerebras takes decode with raw speed; Groq's LPX fights for the same lane. Mix-and-match is precisely the market structure Nvidia's bundle exists to prevent.
Tokens per watt is the tell
Note what the headline metric is not: it is not FLOPS, and it is not tokens per dollar. It is tokens per watt. That reflects where the binding constraint has moved. With multi-gigawatt campuses queuing for grid connections from Ohio to Johor, the scarce input in inference economics is increasingly electricity, not silicon. A 5x efficiency claim — even a modeled one — is aimed at operators who know exactly how many megawatts they can get and want to know how much revenue each one yields.
That framing matters most in Asia, where the token business is growing fastest and power is hardest to secure. Japan, Korea, India and Southeast Asia are all building sovereign inference capacity under grid constraints that make American-style brute force impossible. For a Malaysian or Indian cloud deciding what to rack in 2027, a disaggregated system that squeezes multiples more tokens from the same substation is not an optimization — it is the difference between a viable business and a stranded one.
Second-order effects: memory, models, margins
Two quieter implications deserve attention. First, memory. Prefill-specialized silicon — Rubin CPX with GDDR7, Helios in the throughput role — deliberately economizes on HBM, the product carrying SK Hynix and Samsung's record profits. Decode still devours bandwidth, so HBM demand is not going anywhere. But if the industry standardizes on architectures where a meaningful fraction of inference compute runs on conventional memory, the HBM share of each incremental data-center dollar shrinks at the margin. For Korean memory makers priced for an unbroken supercycle, architecture is now a demand variable.
Second, models. The benchmark workload both companies chose was not GPT or Claude — it was Kimi 2.6, a Chinese open-weight model. That is a quiet acknowledgment of where the volume inference market actually lives: open-weight, trillion-parameter, Asia-originated models served by whoever offers the cheapest, fastest tokens. Hardware challengers and Chinese model labs are natural allies here; neither can win inside Nvidia-CUDA-plus-closed-model orthodoxy, and both gain if the serving stack fragments.
The honest bottom line: a modeled 5x against your own baseline does not dent 90 percent market share, and Nvidia has already hedged the architectural shift. But the direction of travel is set. Inference is industrializing into specialized stages competing on watts and dollars, and every step in that direction converts the AI hardware market from a monopoly on genius into a business of margins. Those are the markets Asia's manufacturers and power-constrained clouds have always been best at fighting in.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.


