cd /news/artificial-intelligence/with-taalas-amd-can-bake-ai-inferenc… · home topics artificial-intelligence article
[ARTICLE · art-88385] src=nextplatform.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

With Taalas, AMD Can Bake AI Inference Directly Into Its Chippery

AMD acquired AI inference startup Taalas for an undisclosed sum this week, aiming to bake AI inference directly into its chips. The move follows Nvidia CEO Jensen Huang's GTC 2026 disclosure that GPU architectures alone cannot deliver the low-latency inference needed for agentic AI, prompting a split between prefill on GPUs and decode on specialized SRAM-heavy engines. AMD's acquisition mirrors Nvidia's $20 billion acquihire of Groq and its partnership with Cerebras for disaggregated inference.

read6 min views1 publishedAug 7, 2026
With Taalas, AMD Can Bake AI Inference Directly Into Its Chippery
Image: Nextplatform (auto-discovered)

Jensen Huang, the chief executive officer and co-founder of Nvidia, let the cat out of the bag back at GTC 2026 in March that if hyperscalers, cloud builders, AI model builders, and other enterprises as well as sovereigns needed low latency inference if they want to use agentic AI or if they just want to charge a premium for chattybot inference, they could not do this on GPU architectures.

What Huang showed was that to get the best low latency performance across a wider range of latency profiles, what companies needed to do was to split their AI inference engines and process the input tokens of a context window on GPU clusters, which are very good at this work and which is called prefill, and then offload the decode part – where the model gives its responses – on a massively parallel, more deterministic, SRAM heavy matrix math engine like those created by Grok, SambaNova Systems, Cerebras Systems, and others.

The reason is simple enough. A GPU is an excellent, high performance, high bandwidth general purpose parallel processing engine, but its performance on the decode part of inference is less deterministic than is ideal. Huang’s charts at GTC showed that a system comprised of its “Grace” CG100 CPUs and “Hopper” H100 GPUs (presumably an eight GPU node) could handle a load of maybe 100 tokens per second per user by batching up those users and sort of petered out after that.

He called this a medium tier of performance. The NVL72 hybrid CPU-GPU platforms, which have 18 CPUs and 36 GPUs based on the same Grace CPUs and Nvidia’s “Blackwell” B300 GPUs, offers around 3.5X the performance (expressed in tokens per second per megawatt) at the 100 TPS interactivity level. An impending NVL72 rackscale system based on the impending “Vera” CV100 CPUs and “Rubin” R200 GPUs could do around 2X the TPS/MW of the Grace-Blackwell system and that means it delivers 7X that of the Grace-Hopper system.

The Grace-Hopper node could not really do higher interactivity than this, but you could push a Grace-Blackwell system to support what Huang called a High tier with 200 TPS/user at a reasonable overall throughput, but Vera-Rubin NVL72s would do 3X better. Out at a Premium tier, with 400 TPS interactivity, the TPS/MW goes very close to zero for the Grace-Blackwell rackscale system, but a Vera-Rubin could do 10X better than that and deliver a reasonable overall TPS/MW.

Again, these are all using GPUs for both the prefill and decode parts of the inference. But look what happens when Nvidia split the inference load, putting prefill on the GPUs but decode on what we presume was a rack of the current Grok LP30 accelerators:

The whole new Ultra tier of inference, which is probably a baseline of performance for agentic workloads with upwards of 1,000 TPS/user of interactivity, opens up. Not only that, the Premium tier on the Vera-Rubin NVL72 plus Grok now has 35X the performance of the Grace-Blackwell NVL72 without Grok (compared to 10X without Grok accelerators). You can plainly see that GPU accelerators alone really can’t drive more than 400 TPS/user at a reasonable overall system throughput.

Of course, there are limited to how far a Grok cluster can scale on the decode part, too. You can’t link together a million of these devices to drive interactivity further. (We do not know what those limits are, of course.)

This pair of charts not only explain why Nvidia did a $20 billion aquihire of the Grok team and technology back in December, but it also explains why AMD is partnering with Cerebras for disaggregated inference and why it is also acquired – not some dubious, scrutiny-avoiding acquihire, but a real acquisition – AI inference startup Taalas for an undisclosed sum this week. AMD’s GPUs have the same limitations when it comes to inference decode as do Nvidia’s GPUs.

The Cerebras partnership is a good thing for both AMD and for Cerebras, which will put what we presume will be the fourth generation of waferscale compute engines (WSE-4) that are expected to be announced later this year alongside AMD’s “Helios” clusters using “Verano” Epyc CPUs and “Altair” MI455X GPUs doing the same kind of disaggregated inference as Nvidia was showing with its future Vera-Rubin-LPU combination.

The only trouble is, AMD does not control Cerebras technology because the latter company raised a ton of money and became very expensive for AMD to acquire before it went public earlier this year. AMD increasingly likes to control its own stack, and right now Cerebras has a market capitalization of $50.9 billion, which means AMD might have to shell out $60 billion for a company that is making just under $200 million in revenues per quarter right now. Paying 300X annual revenue is very risky in a world that pays maybe one-tenth of that premium.

Hence, it looks like AMD will partner for now – Cerebras was its only choice with Graphcore being eaten by Softbank, the owner of Arm, Nvidia essentially buying Grok, and SambaNova partnering with Intel – and look for a new approach for future inference acceleration.

I did a deep dive on Taalas back in February this year when the company dropped out of stealth mode, and I am not going to repeat all of that here. The company was formed by a bunch of engineers who founded or worked at AI accelerator and RISC-V champion Tenstorrent, and its innovative inside is to take the AI inference models and their weights and hard coding them into ROM circuits and link them to giant SRAM blocks that act as an on-chip KV cache. The current HC1 generation of Taalas chips can store a model with 8 billion parameters, and the next-gen HC2 was slated to do 20 billion parameters. A few tens of chips interlinked could, in theory, hold the models and weights of inference models spanning a trillion parameters. In initial benchmark tests, Taalas is showing extremely low latencies and much lower costs per token for inference compared to the Nvidia Blackwell B200 GPUs.

The trick with Taalas is that each model would require a variant of the HC1 chiplet – the SRAM and components around it are the same, but two layers of metal in the chip that encode the model and weights has to change. But, then again, Taalas says it takes 100X as much money to train a new GenAI model as it does to customize an HC chip from Taalas and buy it in reasonable volumes (a few hundred thousand units is a good guess for this comparison).

AMD is not saying much about its plans for Taalas except that “AMD plans to integrate the technology into its accelerator roadmap and develop system-level solutions with AMD Instinct GPUs.”

Welcome to the world of Model Specific Architectures.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @amd 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/with-taalas-amd-can-…] indexed:0 read:6min 2026-08-07 ·