cd /news/ai-chips/d-matrix-tells-lds-how-to-judge-rapt… · home › topics › ai-chips › article
[ARTICLE · art-142603] src=letsdatascience.com ↗ pub= topic=ai-chips verified=true sentiment=· neutral

d-Matrix tells LDS how to judge Raptor before it ships

D-Matrix co-founder and CTO Sudeep Bhoja told Let's Data Science that Raptor's early 3D-DRAM silicon uses 0.37 picojoules per bit moved, roughly five to eight times better than the 2 to 3 picojoules per bit he cites for HBM4, while the company's simulated 72-card Raptor system serving Kimi K3 reaches about 1,000 tokens per second per user at a one-million-token context. Bhoja separated those company-reported component and modeled figures from measured whole-system performance, noting Raptor tape-out is expected by the end of 2026 and full racks in Q4 2027, with complete-system speed, power and cost-per-request still unmeasured. He advised buyers to count only requests meeting speed and quality requirements, include idle hardware and data transfers, and identify which NVIDIA dependencies a proposed system retains.

read10 min views1 publishedSep 30, 2026
d-Matrix tells LDS how to judge Raptor before it ships
Image: Letsdatascience (auto-discovered)

In written answers to LDS, d-Matrix co-founder and CTO Sudeep Bhoja separates early memory-silicon measurements from modeled Raptor performance. Full racks are expected in Q4 2027. Complete-system speed and power remain unmeasured, and Raptor cost-per-request figures are not yet available. His practical advice: count requests that meet speed and quality requirements, include idle hardware and data transfers, and understand which NVIDIA dependencies the proposed system retains.

An AI service can generate tokens quickly and still be expensive to run. The bill includes machines waiting for work, the movement of data between processors and requests that fail to meet the service's requirements.

That is the buying problem Sudeep Bhoja addresses in an original written interview with Let's Data Science. The d-Matrix co-founder and CTO explains how to assess Raptor, the company's planned inference accelerator, while separating three kinds of evidence: measurements on early memory silicon, simulations of a future system and results from its existing Corsair product.

The distinction matters because these numbers describe different things. A promising memory result does not yet establish the electricity bill for a complete rack. A modeled token-generation rate does not tell a team what each useful response will cost.

"Some of our results come from real hardware, some come from modeling, and some are still targets."

Bhoja's answer provides a more useful starting point than treating every performance figure as a finished-product benchmark. Raptor is expected to complete tape-out, the handoff of its chip design for fabrication, by the end of 2026. The company expects full racks in Q4 2027. Neither milestone means the complete system is available today.

What d-Matrix has measured, and what it has modeled

Raptor brings memory and compute together in a stacked package. The aim is to reduce the cost of moving data while providing the memory bandwidth needed for fast token generation.

Bhoja says early 3D-DRAM silicon uses 0.37 picojoules per bit moved. A picojoule is a unit of energy; the figure concerns moving data, not serving an entire AI request. The company's Hot Chips presentation identifies this as an I/O energy figure. Bhoja compares it with about 2 to 3 picojoules per bit for HBM4, putting the claimed advantage at roughly five to eight times for that comparison.

Those are company-reported component figures. They should not be read as a five-to-eightfold reduction in whole-system power. Compute, other memory activity, networking and the rest of the system also consume energy.

For speed, Bhoja describes a simulated 72-card Raptor system serving Kimi K3 at about 1,000 tokens per second per user with a one-million-token context. The context is the material the model can work with during the request. It is not a statement that the system generates a million output tokens in one second. The public slides add useful conditions. Their sizing assumptions include 4-bit weights and an 8-bit key-value cache. The Kimi K3 chart shows 988 tokens per second per user at a decode batch of eight users, falling to 785 at 32 users. These are modeled figures, as clarified in the interview, rather than measurements from a complete Raptor rack. Concurrency changes the result.

Evidence What Bhoja reports What it does not establish
Early 3D-DRAM silicon 0.37 picojoules per bit for moving data The power draw of a complete Raptor system
Raptor simulation About 1,000 tokens per second per user for the stated Kimi K3 scenario Measured production latency, power or cost per request
Existing Corsair hardware A reported cost-per-token improvement in a separate speculative-decoding setup The same saving on Raptor or on a different workload
Product roadmap Tape-out expected by the end of 2026; full racks expected in Q4 2027 Current availability or a guaranteed delivery date

Bhoja explicitly says d-Matrix has not measured end-to-end speed or power on a full Raptor system. The September 10 announcement likewise places initial availability of the NVIDIA-integrated Raptor rack in Q4 2027.

Why one request may use two kinds of processor

The proposed split is easier to understand by following a request. First, the model processes the prompt and other input. This is prefill. It then generates the response through repeated decoding steps.

In the configuration described in our questions and Bhoja's answers, GPUs handle prefill and d-Matrix accelerators handle decode. The reasoning is that the two stages can have different needs: processing a large input can demand substantial computation, while generating tokens quickly can depend heavily on moving model data from memory.

The handoff is not free. Bhoja highlights the key-value cache, or KV cache: stored attention data that lets generation reuse work associated with earlier tokens. Passing that state from a GPU to an accelerator takes time and energy. The memory available on each accelerator also limits how many users it can serve at once.

An operator therefore has to size and manage two hardware pools. If the GPUs or the accelerators are underused, the idle capacity still contributes to cost. Buying a fast second stage does not by itself produce an efficient complete service.

Count completed requests that meet the requirement

Bhoja's proposed comparison begins with the service outcome:

"Buyers should look at cost per completed request, not cost per chip."

His method takes the total cost of running the system for an hour, including hardware, networking, power and operations, and divides it by the requests that met the speed and quality targets during that hour.

Cost per qualifying request = total system cost for the period / completed requests meeting the agreed speed and quality requirements.

Failed attempts and retries still add to the bill. They do not become successful outcomes merely because they consumed tokens. A fair comparison also needs the same workload and acceptance criteria on both systems.

This is a proposed evaluation method, not a Raptor price quotation. Bhoja says d-Matrix will publish Raptor cost-per-request figures once it has measured silicon. The interview does not supply those figures today.

For an AI engineering team, the practical implication is to define an acceptable result before comparing infrastructure. Specify how quickly a request must finish, what makes its output acceptable and how repeat attempts are counted. Then test the traffic pattern the service actually experiences, including quiet periods.

The 72% result belongs to Corsair

The closest hardware example Bhoja supplies uses the company's existing Corsair product in a different arrangement: speculative decoding.

A smaller helper model proposes tokens, and a larger model checks them. In the example he describes, Corsair runs Qwen3 1.7B as the helper, while four NVIDIA H200 GPUs run Qwen3 235B to perform the checking.

Bhoja reports a measured 72% reduction in cost per token compared with the GPUs alone. This is d-Matrix's reported Corsair result. It is not a measured Raptor saving, nor evidence that every completed application task would become 72% cheaper.

He says the larger model's verification preserves output quality in that configuration. Whether the token-cost improvement translates into a lower cost per completed request still depends on meeting the speed requirement and keeping the system well utilized. LDS did not reproduce the test or independently assess its output quality or cost accounting.

Bhoja describes Corsair as being in full production and says it validates the company's software stack and experience operating accelerators alongside GPUs. He also states what it cannot validate: Raptor's 3D-DRAM performance, its NVLink Fusion behavior or its operation at rack scale.

When the split can cost more

Bhoja identifies three situations where the architecture offers less benefit: requests dominated by long inputs and short outputs, batch work without a demanding response-time requirement, and light or uneven traffic.

The first spends less of its work generating output. The second may not need to pay for very fast interactive responses. The third risks leaving the additional accelerator capacity unused.

"When traffic is light or bursty, idle XPU capacity can push cost per completed request above a GPU-only setup, so the split pays off only when both pools stay well utilized."

For a platform team, that is a reason to examine sustained utilization alongside peak speed. A configuration selected for a busy demonstration may behave differently during a normal week of uneven demand. This is an evaluation consideration drawn from Bhoja's explanation, not a measured customer outcome supplied to LDS.

Accelerator choice still leaves infrastructure dependencies

The proposed NVIDIA integration divides responsibility across several parties. Bhoja says d-Matrix supplies Raptor chips, trays and its software. NVIDIA supplies the rack design, NVLink connections and switches, CPUs and networking. Astera Labs supplies connectivity. The customer decides how much hardware to use, routes requests and monitors the service.

NVIDIA's description of the collaboration confirms that the plan puts d-Matrix accelerators within its rack and networking ecosystem. It gives buyers another accelerator option inside that infrastructure; it does not remove the infrastructure dependency.

Bhoja says the partnership is non-exclusive. Corsair already runs as a PCIe card in standard servers, and d-Matrix's roadmap includes an Ethernet-based option. He says moving to another platform would require changes to the interconnect and orchestration layer, while the company's software and kernel library move with the workload. The interview does not demonstrate that migration or quantify its cost.

What to ask before committing

The interview suggests a practical checklist for teams evaluating specialized inference hardware:

  • •Separate component measurements, system simulations and production-system measurements. Record which product each result concerns.
  • •Keep the model, numerical precision, input and output lengths, user concurrency and quality requirement consistent across comparisons.
  • •Measure complete-request latency as well as token-generation speed, including data transfers and retries.
  • •Include both hardware pools, networking, energy, operations and idle time in the cost calculation.
  • •Identify the hardware, software and operational work needed to deploy, run or migrate the service.

Raptor's memory approach and planned NVIDIA integration give teams an architecture to evaluate. Bhoja's answers also make the next evidence requirement clear: a measured complete system running the intended workload. Until then, the useful comparison is one that keeps the component result, modeled performance and finished-request cost separate.

Reporting note

This LDS Exclusive is based on three written answers attributed to Sudeep Bhoja, co-founder and CTO of d-Matrix, supplied directly to Let's Data Science through Aircover Communications. LDS reviewed the company's Hot Chips slides and the public d-Matrix and NVIDIA announcements as supporting context. All reported hardware measurements, simulations, cost improvements and availability targets are attributed to d-Matrix. LDS did not test Corsair or Raptor, reproduce the benchmarks or independently verify their costs. The practical evaluation checklist is LDS's interpretation of the interview.

Key Points #

  • 1The 0.37-picojoule figure concerns early memory I/O silicon. Raptor's near-1,000-token result is modeled; complete-system speed and power remain unmeasured, with cost-per-request figures still to come.
  • 2Bhoja recommends cost per request meeting speed and quality requirements, including idle capacity, networking, transfers and retries. Light traffic can make a split system more expensive.
  • 3The reported 72% token-cost reduction comes from a separate Corsair setup. Raptor racks are expected in Q4 2027 and the proposed configuration retains NVIDIA infrastructure dependencies.

Scoring Rationale #

Original technical answers give AI engineering teams a practical framework for separating component evidence, modeled speed and complete-request costs.

Sources #

Original reporting, with the public references used alongside it.

LDS Exclusive

Reporting based on written answers given directly to Let's Data Science by Sudeep Bhoja, Co-founder and CTO, d-Matrix.

Practice interview problems based on real data

1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.

Try 250 free problems

── more in #ai-chips 4 stories · sorted by recency
── more on @d-matrix 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/d-matrix-tells-lds-h…] indexed:0 read:10min 2026-09-30 · —