AI inference: the CS-4's real edge is routing, not 30x

Cerebras launched the CS-4, promising up to 30 times GPU token speed and deals with OpenAI and AMD. We think the headline is the wrong story. The AMD partnership splits workloads across chips, proving the value in AI inference now lives in the coordination layer, not the silicon.

TL;DR: The CS-4's headline speed numbers are real, but the quiet story is disaggregation — the value is moving from the chip to the layer that decides which chip does what.

Key takeaway: Cerebras just made AI inference faster and more power-efficient, yet its most telling move is splitting workloads across its own and AMD's hardware — a coordination problem, not a chip problem.

Why it matters: As power becomes the binding constraint on scaling AI, the winners won't own the fastest silicon. They'll own the handoffs between silicon.

What happened

Cerebras used a company event led by CEO Andrew Feldman to unveil the CS-4 and a set of partnerships aimed squarely at AI inference. You can read the original coverage of the CS-4 launch and roadmap for the full rundown.

The specs are genuinely large. According to HPCwire, the CS-4 delivers up to 30 times the tokens per second per user of GPU-based systems and up to 10 times the throughput per watt of the previous CS-3, with first shipments scheduled this quarter.

Source: HPCwire, 2026

Alongside the chip, two alliances landed. Cerebras signed a multi-year deal with OpenAI to deploy 750 megawatts of wafer-scale systems for OpenAI customers, and a separate partnership with AMD lets customers split inference workloads across both companies' systems.

Source: Cerebras / Axios, 2026

Most people will read this as a shot at Nvidia's GPU crown

The consensus take writes itself: a specialist challenger posts eye-watering speed numbers, bags OpenAI and AMD as validators, and threatens the GPU incumbents. Faster tokens, cheaper tokens, a new king of inference. It's a clean narrative and it isn't wrong, exactly — the throughput figures are real and the OpenAI commitment is serious money.

But "our chip is faster" is the oldest story in computing. Reading the CS-4 purely as a horse race misses what the AMD deal is actually admitting.

We think the coordination layer is the real product in AI inference

Look at the AMD partnership again. It doesn't say "buy our chip instead of theirs." It says split the work — route the latency-sensitive parts to one system and the throughput-heavy parts to another. That's disaggregated inference, and it's a confession: no single piece of silicon wins every job.

Which is exactly our house view. A monolith has to be right about everything. An ecosystem only has to be right about the handoffs. The moment you're routing workloads across Cerebras and AMD hardware, the product stops being the wafer and becomes the router — the layer that decides what runs where, when, and at what cost per token.

That layer is unglamorous. There's no keynote slide for "we got the scheduling right." But in our experience building agents, the scheduling is where the value and the risk both live. Anyone can wire two fast systems together in a demo. Keeping the handoff correct under real load, across two vendors, with a power budget as the hard ceiling — that's the hard part, and it's the part that compounds.

Here's why power changes the maths. The IEA projects that global data-centre electricity use will more than double to around 945 TWh by 2030 — roughly Japan's entire current consumption.

Source: IEA, 2025

When electricity is the constraint, "tokens per watt" isn't a vanity metric — it's the budget line. And optimising tokens per watt across a mixed fleet is a coordination job, not a hardware spec. The same logic runs straight through the marketing stack: the tools were never the expensive part. The stitching is. We build our AI agents for marketing around that belief — get the handoffs right and the underlying models become swappable commodities, which is precisely what Cerebras and AMD are betting on at the silicon layer.

So the interesting question isn't "is the CS-4 faster than a GPU." It's "who owns the routing when you're running three vendors and paying by the watt." That owner captures the margin. Everyone else supplies parts.

What this means for marketing teams

  • Stop benchmarking models in isolation. Measure cost per finished output — a published post, a qualified lead — across your whole workflow, and review it monthly.
  • Treat your model provider as swappable. Build so you can change the engine behind an agent in under a day, not a quarter.
  • Instrument the handoffs. Log where each task is routed and what it costs; the expensive failures hide between tools, not inside them.
  • Put a number on efficiency. Track output per pound spent, the way infrastructure teams now track tokens per watt.
  • If you're mapping this out, our transparent pricing for agent builds shows where the stitching costs actually sit.

Frequently asked questions

What is the Cerebras CS-4?

The CS-4 is Cerebras's wafer-scale AI system for inference, delivering up to 30 times the per-user token speed of GPU-based systems and up to 10 times the throughput per watt of its predecessor, the CS-3, per Cerebras. First shipments begin this quarter.

Why does the Cerebras AMD partnership matter?

It lets customers split AI inference workloads across Cerebras and AMD systems, routing each task to the best-suited hardware. This disaggregated approach signals that coordination between chips, not any single chip, is becoming the source of competitive advantage.

Why is tokens per watt important for AI inference?

Electricity is becoming the main limit on scaling AI. Tokens per watt measures useful output per unit of power, so improving it directly lowers cost and lets providers serve more demand within a fixed energy budget.

Written by the Anjin team - we build AI marketing systems and remain professionally unimpressed by hype.

Continue reading