# SK Hynix and Etched Frontier Inference Clusters

> Source: <https://www.blocksandfiles.com/ai-ml/2026/07/24/sk-hynix-and-etched-frontier-inference-clusters/5278155>
> Published: 2026-07-24 14:10:45+00:00

# SK Hynix and Etched Frontier Inference Clusters

SK Hynix has participated in a funding round for [Etched](https://www.etched.com/), a US-based AI Inferencing startup developing rack-scale inferencing systems using Low Voltage Inference (LVI) and Cluster Scale Memory (CSM) concepts.

Etched is trying to fix the problem emerging from the realization that classic GPUs are over-powered and have too little memory for running large-scale inferencing workloads. Consequently complicated, multi-tier KV Caching schemes have had to be developed. We need processors better tuned for inferencing work that draw less electricity, and have access to much more memory than a GPU’s limited capacity HBM, thus avoiding KV caching latency lengthening and complexity.

This news closely follows the emergence from stealth of Tel Aviv-based [Majestic Labs](https://www.blocksandfiles.com/ai-ml/2026/07/23/majestic-labs-wants-to-solve-memory-bound-gpu-problem/5277163) which is developing AI inferencing server technology involving custom processors and a large-scale pool of shared memory as well.

Etched was started up in 2023 by three Harvard students who left Harvard before completing their undergraduate degrees and moved to San Jose in the Bay area. They are CEO Gavin Uberti, President and maybe COO Robert Wachen and Chris Zhu who has an unspecified role. They later became Thiel Fellows, being recipients of a 2-year Thiel fellowship with its $250,000 grant. These are designed to help college dropouts build something new and helped the trio startup Etched.

The trio saw that they needed to develop custom prefill processing using LVI and a shared pool of memory; CSM, for the decide phase of inferencing, with both being part of a full system, a frontier inference cluster with co-designed chips, packages, PCBs, cold plates, and interconnects, that enables the running of inference workloads at unimagined scale.

Zhu [posted](https://www.linkedin.com/feed/update/urn:li:activity:7477738639771193345/ ) “Today, AI chips can't scale FLOPs without thermal throttling. As FLOPs utilization increases, AI chips draw more power and downregulate clock speed. This often results in sustained inference throughput under half of peak FLOPs. … We’ve designed a new architecture to run our chip’s math blocks at under half the voltage of most AI chips. This enables multiple times the FLOPs density of AI chips today.”

He says “AI chips using HBM can’t achieve SRAM-level decode speeds due to memory subsystem and interconnect bottlenecks. SRAM-only chips have lower FLOPs density and memory capacity, sacrificing throughput.”

This leads to another problem: “When running large MoE models, token routing across experts requires sending data through a deep memory hierarchy and a networking switch to reach a destination expert. Each memory layer inherently adds latency; thus, the best layer is no layer.”

Concerning CSM, he posted: “We’ve designed a new architecture that creates a shared low-latency memory pool across the entire scale-up domain. We use a proprietary ultra-low-latency, high-bandwidth interconnect to enable dramatically faster memory access across chips.

“Our HBM/SRAM hybrid design solves both memory capacity and mem2mem latency, enabling high throughput and interactivity simultaneously. CSM improves latency and avoids today's cost, reliability, yield, thermal, and compute tradeoffs of SRAM-only chips, 3D DRAM chips, or optics.”

This is intriguing as it suggests that Etched has found a way to decouple HBM stacks from a GPU and its shoreline limitations, group them together in a pool and achieve SRAM speed with its new interconnect.

A picture of an Etched system tray shows an unusually large number of cables; 16 of them with woven casing linking rear ports to 8 copper-coloured units;

Etched systems are designed to run traditional LLMs, mixture-of-experts models and non-transformer systems like Mamba. The company aims to develop and sell full systems.

Etched funding history:

2022 - founded

2023- $5.4 million seed round

2024 - $120 million A-round

2025 - $500 million B-round

2026 - $300 million C-round

Total funding to date is $925.4 million.

The C-round was led by Sequoia Capital alongside Andreessen Horowitz, Jane Street, Diffusion, Argo, and SK Hynix, with Etched valued it at $10.3 billion, effectively doubling the December 2024 $500 million B-round valuation of $5 billion.

This annual funding cadence surely follows convincing product development thresholds that satisfy early customer prospects and convince investors to open their wallets and pour more cash into Etched.

Wachen [posted](https://www.linkedin.com/feed/update/urn:li:share:7486073411245371392/): ”This round accelerates production of our inference clusters, and we've opened an 80,000-sqft, 10-MW facility 15 minutes from our office to expedite production and prototyping.” That’s in MIlpitas.

Chief Architect Saptadeep Pal says Etched now employs more than 400 people and has more than $1 billion in customer demand. It's set up a a factory in Taiwan, to support 24/7 engineering cycles and production, and also an office there to colocate with suppliers and speed up testing. All-in-all it's scaling up at a furious rate.

The obvious markets for its systems will be large enterprises, neocolouds and hyperscalers who can sell inferencing-as-a-service.

##### Comment

If Majestic Labs and Etched inference processing racks don’t need KV Caching schemes then the storage array playing field is levelled. They’ll need arrays which can load data into their memory pools at high speed and low latency, and lots of suppliers can do that. We expect the two companies will be getting approaches from the usual suspects: Dell, DDN, Everpure, HPE, MinIO, NetApp, WEKA, VAST Data, and others. Have an array, they’ll say. Let’s hook it up to your servers, and see how fast we can make it go.
