{"slug": "ai-s-inference-era-of-ferment-by-ben-bajarin", "title": "AI's Inference Era of Ferment – By Ben Bajarin", "summary": "At Hot Chips 2026, semiconductor vendors presented divergent designs for AI inference, signaling an era of ferment where no dominant architecture has emerged, according to analyst Ben Bajarin. Companies disagree on whether memory bandwidth, data movement, or processor design is the key constraint, leading to varied approaches such as Samsung's zHBM, d-Matrix's custom DRAM, and OXMIQ's flash-based HBF. This technical diversity, framed by Anderson and Tushman's 1990 theory, suggests the industry is testing competing solutions before a standard design prevails.", "body_md": "Our hottest take from Hot Chips 2026 is that we may be in one of the semiconductor industry’s most inventive periods in decades, if not ever. Every presenter treated inference at scale as the common engineering problem: the system has to produce more useful tokens at a speed and cost customers will accept. The disagreement began with what keeps the system from doing that. One vendor sees memory bandwidth or capacity as the limit. Another focuses on how far data has to move, while others want to change the processor or use software to place work more carefully. Solving one constraint moves cost or complexity somewhere else, which is why the companies arrived at such different designs. That range of technical bets is why we believe inference has entered an era of ferment.\n\n## Why This Looks Like an Era of Ferment\n\nIn their seminal 1990 paper, * Technological Discontinuities and Dominant Designs: A Cyclical Model of Technological Change*, management scholars Philip Anderson and Michael Tushman formalized the idea of an “era of ferment.” They used the term for the period after a major technical break, when an industry tests competing approaches before one becomes the common architecture.\n\n**Their research found that the design that wins often, though not necessarily, trails the technical frontier because adoption also depends on cost, manufacturing scale, software support and ease of use.**(\n\n[Anderson and Tushman, 1990](https://www.edegan.com/pdfs/Anderson%20Tushman%20%281990%29%20-%20Technological%20Discontinuities%20and%20Dominant%20Designs.pdf))\n\nHot Chips gave us the clearest evidence yet for applying that framework to inference. NVIDIA has already set the merchant standard for training. Inference is producing a much wider range of technical bets, allowing for some competitive pressure, because vendors disagree about where data should live and which processor should handle each part of a request. Those choices change how much work software must carry. More specialized hardware can improve speed or cost, although the gain has to justify the added work. We highlight a few examples below before examining each bet in the full report.\n\n## Memory Is Becoming Part of Compute Design\n\nMemory gave us some of the clearest examples of these different views. Micron and SK hynix are extending standard HBM through taller stacks and better bonding. Samsung is turning the HBM base die into a control point, a move we believe will become table stakes for next-generation HBM, and eventually wants to place DRAM over logic with zHBM. d-Matrix bonds custom DRAM under its accelerator. OXMIQ’s HBF uses flash as a cheaper home for model state that stays cold, while XCENA and Samsung use CXL memory to hold older KV pages and reduce them before sending a result back to the GPU. HBM supports a broad range of workloads at higher package cost. HBF and CXL ask software to identify state that can sit farther away, while d-Matrix accepts a smaller local memory pool to keep the useful data close.\n\n## Ethernet Still Provides the Common Network Base\n\nThe networking talks showed more agreement at one layer. Ethernet is becoming the common protocol for scale-out networks. Copper remains practical over the shortest links, while optics—[and the variety of approaches showing up](https://www.thediligencestack.com/p/optics-wont-scale-as-fast-as-the)— takes over as distance and speed increase. The scale-up fabric connecting accelerators into one tightly coupled system remains less settled. Broadcom provides merchant NICs and gives customers more choice over cables and optics. NVIDIA controls the DPU, switch and software, and is bringing optics into the switch. Microsoft changes how work moves over Ethernet through software. Google uses optical circuit switching to shape the network around the workload. These systems share a protocol while putting the physical links and traffic control in different places.\n\n## Accelerators Carry the Widest Disagreement\n\nAccelerators showed the widest range of views. We saw it in the public presentations and in meetings with six private AI accelerator companies, where each team had a firm view of why its architecture maps better to inference. The public presentations gave us several examples of how those views lead to very different hardware.\n\nGoogle has landed on an ASIC roadmap that separates TPU designs for training and inference, while Meta is building one MTIA platform for recommendation and generative AI. Microsoft uses software to place data explicitly across Maia. OpenAI keeps KV state local and changes the active compute mix inside Jalapeño as a request moves through its phases. NVIDIA is bringing GPU throughput and Groq’s low-latency LPU into one platform. SambaNova maps the model into a persistent dataflow system, while Cerebras makes the wafer the compute unit. Each architecture also carries a forecast about how models will behave years from now. A design built around sparse access, local KV state or deterministic execution keeps its advantage only while that workload assumption holds.\n\n## NVIDIA Is Closest to the Merchant Standard\n\nOur conviction is that the dominant design will form at the system level because useful inference depends on software coordinating compute and memory to produce the required output at an acceptable speed and cost. As specialized processors are added, that coordination becomes part of the architecture itself. The system has to make different processors and memory designs work as one pool of capacity, which pushes vendors toward deeper co-design around the workload.\n\n[NVIDIA](https://www.nvidia.com/en-us/networking/) is the clearest current example of that thesis. It has the most complete in-house control of the merchant AI system, from accelerators and CPUs through networking and fleet software. CUDA ties those parts together. Bringing Groq’s deterministic LPU alongside the GPU shows how NVIDIA can add a specialized inference engine without asking customers to operate a second system.\n\nAMD’s [Helios](https://www.amd.com/en/products/rackscale-solutions/helios.html) follows the same direction through an open reference design that partners turn into products. [Google](https://docs.cloud.google.com/ai-hypercomputer/docs/overview) and [Amazon](https://aws.amazon.com/ai/machine-learning/trainium/) already integrate custom chips with their cloud infrastructure around workloads they control. Hyperscalers have enough internal demand to support those systems, while NVIDIA is building for customers that need to buy the full stack. That gives NVIDIA the clearest path to becoming the broadly used standard outside the largest clouds.\n\n## How We Will Know the Market Is Settling\n\nThis period could last longer than earlier hardware shifts because AI models can change several times during one chip-development cycle, while inference includes workloads with very different speed and memory needs. A design that fits today’s models may lose its advantage before the next chip reaches volume. Specialized hardware only improves the economics when its performance gain exceeds the added software work and operating cost. If we use the era of ferment framework we would expect this part of the cycle to end when competition among rival designs produces a dominant design, after which the industry shifts toward incremental improvement around that common architecture. We detail in the full report why some nuance and variations will still exist but it will not be nearly as diverse as we see in the market today.\n\nWhat intrigues us about this moment is what that long testing period could produce. And a follow on question to who ultimately benefits the most in longevity of uncertainty in models and inference architecture standards. Several 800-pound gorillas already control much of the market, yet the range of architectures gives the industry a way to test very different answers against real workloads. Some will fall short once their software and operating costs are included. Others could create a larger opening for a smaller company or become part of an incumbent’s broader roadmap. In the rest of this report, we examine the technical bets from vendors large and small that stood out at Hot Chips 2026, along with the workload assumptions that have to hold for each one to work.\n\n### Paid subscribers get the full technical breakdown\n\nHow HBM scaling, Samsung zHBM, d-Matrix 3D DRAM, HBF and CXL processing near memory each relocate the memory bottleneck, including the workload behavior required for their economics to hold\n\nThe architecture choices that remain open above Ethernet, including the handoff from copper to optics and the division of fabric control between hardware and software\n\nThe model and software forecasts embedded in Google TPU8, Meta MTIA, Microsoft Maia, OpenAI Jalapeño, NVIDIA/Groq, SambaNova and Cerebras, along with the assumptions that would strengthen or weaken each technical bet\n\nWhy NVIDIA is currently closest to the merchant dominant design and how hyperscalers or specialized accelerators could still sustain alternative systems\n\nThree original Creative Strategies exhibits mapping each architecture’s technical bet, required workload behavior, main constraint and the proof points needed to separate benchmark gains from durable inference economics", "url": "https://wpnews.pro/news/ai-s-inference-era-of-ferment-by-ben-bajarin", "canonical_source": "https://www.thediligencestack.com/p/ais-inference-era-of-ferment", "published_at": "2026-08-27 21:07:48+00:00", "updated_at": "2026-08-27 21:18:36.979202+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-chips", "ai-infrastructure"], "entities": ["Ben Bajarin", "Hot Chips 2026", "NVIDIA", "Micron", "SK hynix", "Samsung", "d-Matrix", "OXMIQ"], "alternates": {"html": "https://wpnews.pro/news/ai-s-inference-era-of-ferment-by-ben-bajarin", "markdown": "https://wpnews.pro/news/ai-s-inference-era-of-ferment-by-ben-bajarin.md", "text": "https://wpnews.pro/news/ai-s-inference-era-of-ferment-by-ben-bajarin.txt", "jsonld": "https://wpnews.pro/news/ai-s-inference-era-of-ferment-by-ben-bajarin.jsonld"}}