{"slug": "d-matrix-raptor-drops-the-memory-phy-32gb-at-100-tb-s-against-hbm4s-192gb-at-18", "title": "d-Matrix Raptor Drops the Memory PHY: 32GB at 100 TB/s Against HBM4’s 192GB at 18", "summary": "D-Matrix unveiled Raptor, a datacenter inference accelerator that bonds 3D DRAM directly to the logic die, eliminating the memory PHY, at Hot Chips 2026. The card delivers 32GB of capacity at over 100 TB/s bandwidth and 0.37 pJ/bit, compared to a 192GB HBM4 configuration at roughly 18 TB/s and 2-3 pJ/bit, offering 5.6x bandwidth and five to eight times better energy efficiency. Built on TSMC N4 with Alchip, Raptor targets LLM decode workloads where memory bandwidth is the bottleneck, but shipping details remain unannounced.", "body_md": "*Hot Chips 2026 got its memory-wall heresy: bond the DRAM straight to the logic die and throw the PHY away entirely.*\n\nd-Matrix spent Monday at Stanford [showing early silicon](https://www.servethehome.com/d-matrix-raptor-3d-dram-accelerator-for-generative-inference-at-hot-chips-2026/) for Raptor, a datacenter inference accelerator that does something no shipping part does today. Instead of parking HBM stacks beside the compute die and shouting across an interposer, d-Matrix Raptor bonds a custom 3D DRAM die face to face with the logic die sitting directly on top of it. The numbers that came off the slide are lopsided in both directions, which is exactly what makes them worth reading.\n\nOne Raptor card carries 32GB of 3D DRAM and moves data at [more than 100 TB/s](https://videocardz.com/newz/raptor-3d-dram-combines-32gb-capacity-with-over-100-tb-s-bandwidth), at 0.37 picojoules per bit. The comparison d-Matrix chose is a 192GB HBM4 configuration, which manages roughly 18 TB/s at 2 to 3 pJ/bit. That works out to 5.6x the bandwidth, five to eight times better energy per bit and around seven times the I/O density, bought with one sixth of the capacity.\n\n## Deleting the PHY is most of the trick\n\nHBM bandwidth is expensive because the data has to leave the die. Every stack needs a physical interface, drivers sized to push signals across an interposer, and the power budget that comes with them. Bonding face to face at a 36-micron pitch removes that trip: the logic die, built on TSMC N4, sits immediately above the DRAM, with connections short enough that the PHY stops being necessary at all. What is left is an enormous number of very short wires, which is why the energy figure falls off a cliff rather than merely improving.\n\nIt is also why the density claim holds up. You are no longer limited by how many high-speed lanes can be routed to a beachfront along the edge of a die, because the entire face of the die is the interface.\n\n## Why 32GB might be enough\n\nDecode is the target. When a large language model generates text, it re-reads its weights and KV cache for every single token, and on a modern accelerator that phase spends most of its time waiting on memory rather than doing arithmetic. If the working set fits, bandwidth is effectively the only number that matters, and a part that reads five to six times faster while burning a fraction of the power moves the cost of a token a long way.\n\nIf it does not fit, you are partitioning across cards. Each Raptor card is built from up to four packages, each holding four chiplets, and d-Matrix is betting that scaling out across many small fast pools beats one large slow one. That is a real architectural wager rather than a spec-sheet victory, and it is the part worth watching.\n\n## The problem Corsair could not solve\n\nRaptor succeeds Corsair, which chased the same idea using SRAM: a card pair reportedly reached something like 300 TB/s at roughly 1 ns latency but held only about 4GB. Fast enough, nowhere near big enough to hold anything interesting. Swapping SRAM for 3D DRAM gives back some bandwidth and some latency and buys an eightfold capacity increase, which lands on a far more usable point of the curve.\n\nThe DRAM and packaging work is being done with Alchip, a partnership the two companies [announced in November 2025](https://www.d-matrix.ai/announcements/d-matrix-and-alchip-announce-collaboration-on-worlds-first-3d-dram-solution-to-supercharge-ai-inference/), and the underlying in-memory compute was proved out first on test silicon called Pavehawk. Early silicon at Hot Chips is still early silicon: d-Matrix has not said when Raptor ships, in what volume, or what it costs, and a startup comparing its unreleased part against a competitor shipping product deserves the usual scepticism. But at a moment when memory has grown into roughly a quarter of the cost of an AI rack, a credible attempt to route around HBM entirely is not a small thing.", "url": "https://wpnews.pro/news/d-matrix-raptor-drops-the-memory-phy-32gb-at-100-tb-s-against-hbm4s-192gb-at-18", "canonical_source": "https://hwbusters.com/news/d-matrix-raptor-drops-the-memory-phy-32gb-at-100-tb-s-against-hbm4s-192gb-at-18/", "published_at": "2026-08-24 15:10:32+00:00", "updated_at": "2026-08-24 17:14:23.733963+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-chips", "ai-infrastructure", "ai-research"], "entities": ["d-Matrix", "Raptor", "Alchip", "TSMC", "Hot Chips 2026", "Corsair", "Pavehawk", "HBM4"], "alternates": {"html": "https://wpnews.pro/news/d-matrix-raptor-drops-the-memory-phy-32gb-at-100-tb-s-against-hbm4s-192gb-at-18", "markdown": "https://wpnews.pro/news/d-matrix-raptor-drops-the-memory-phy-32gb-at-100-tb-s-against-hbm4s-192gb-at-18.md", "text": "https://wpnews.pro/news/d-matrix-raptor-drops-the-memory-phy-32gb-at-100-tb-s-against-hbm4s-192gb-at-18.txt", "jsonld": "https://wpnews.pro/news/d-matrix-raptor-drops-the-memory-phy-32gb-at-100-tb-s-against-hbm4s-192gb-at-18.jsonld"}}