cd /news/ai-infrastructure/xcena-and-samsung-put-3072-risc-v-co… · home topics ai-infrastructure article
[ARTICLE · art-115635] src=runtimewire.com ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

XCENA and Samsung put 3,072 RISC-V cores on a CXL memory card

XCENA and Samsung Electronics demonstrated MX1, a CXL memory expansion card with 3,072 RISC-V cores and up to 2TB of DDR5, at Hot Chips 2026 on August 25, targeting production in late 2026. The card, built on Samsung's 4nm process, aims to lower AI inference costs by enabling near-data processing, and follows XCENA's $135 million Series B in May.

read7 min views3 publishedAug 30, 2026
XCENA and Samsung put 3,072 RISC-V cores on a CXL memory card
Image: Runtimewire (auto-discovered)

Jin Kim's prototype combines 2TB of DDR5, SSD-backed memory and near-data processing as XCENA targets late-2026 production.

By RuntimeWire Staff · Published

Primary source: Chips and Cheese

Why it matters #

XCENA is turning memory into a programmable compute tier, a bet that could lower AI inference costs if MX1 survives production and software-integration tests.

Jin Kim and XCENA, the South Korean computational-memory developer he co-founded, used Hot Chips 2026 on August 25 to demonstrate a memory expansion card that can execute selected jobs itself. MX1 combines up to 2TB of DDR5, SSD-backed memory and 3,072 small RISC-V cores on silicon manufactured with Samsung Electronics' 4nm process.

Kim co-founded XCENA with CTO Dohun Kim and CPO Harry Juhyun Kim in early 2022 under the name MetisX; XCENA rebranded in late 2024. At the official Hot Chips session, Harry Kim presented MX1 with Samsung. The three founders came from Samsung and SK hynix, where they worked across memory architecture, system-on-chip design and software. Jin Kim previously served as a corporate vice president at SK hynix and led its next-generation architecture development team.

That background explains the bet. Kim has argued that processor development left memory behind as CPUs and GPUs became more capable. "CPUs and GPUs have both gotten smarter over the decades. Memory never did," he told TechCrunch in May.

XCENA is trying to change memory from a component that feeds processors into a programmable tier that can run selected jobs where the data sits. The Hot Chips presentation added hardware detail to the thesis behind XCENA's $135 million Series B in May.

XCENA is based in Pangyo, a technology district in Seongnam, and also has an office in Sunnyvale, California. TechCrunch reported that XCENA employed more than 90 people across the two locations as of May.

What Hot Chips added

The August presentation exposed how XCENA divides work inside the MX1 computational-memory device. Its 3,072 in-order RISC-V cores run at 1.1GHz. Groups of 32 cores form clusters, four clusters make a 128-core subsystem, and the chip has 24 subsystems that can accept separate jobs. Two Arm Cortex-A53 cores manage the device. Chips and Cheese reported power consumption of 40 watts for the compute chip and about 90 watts for a board populated with four DIMMs.

MX1 connects to its host over a PCIe 6 and CXL 3.2 x8 interface, providing 128GB per second of aggregate host bandwidth. Eight downstream PCIe lanes connect SSDs. The card can carry up to 2TB of DDR5.

XCENA calls the SSD tier Infinite Memory. It exposes attached storage to the host as byte-addressable memory and uses the card's DDR5 as a faster cache. The design can provide considerably more capacity than DRAM alone. Its utility will depend on whether caching and prefetching can hide enough of the SSDs' higher latency for a given workload.

The thousands of cores serve a narrower purpose than the core count might suggest. MX1 provides roughly 3 TFLOPS of dot-product throughput through its vector engines, a qualified measure that is not directly comparable with a general-purpose accelerator rating. XCENA built the device for bandwidth-bound work such as analytics kernels, vector search, memory compression and KV-cache handling, where processing data on the card can reduce transfers over the CXL link.

XCENA's software design is central to that pitch. Host applications and MX1 can operate on the same virtual addresses, reducing pointer translation and data copying. XCENA's software stack includes C/C++ and Rust support, drivers, simulation tools and a MapReduce-style runtime called the Parallel Xceleration Library.

That software layer is where Kim's team has to make unfamiliar silicon usable. Hyperscalers are unlikely to rewrite mature inference and data-processing systems solely to adopt a memory card. XCENA needs the MX1 offload path to fit into existing frameworks and produce savings large enough to justify another programmable device in the server.

Samsung turns MX1 into a rack-scale pitch

Samsung used the presentation to show how MX1 might operate beyond a single add-in card. According to ServeTheHome's account of the Hot Chips slides, Samsung and XCENA presented a reference design connecting GPU servers and MX1 devices through a CXL switch. The shared pool targeted 20TB of capacity and 2.7TB per second of aggregate bandwidth.

Samsung and XCENA reported results from two selected AI workloads, but Samsung and XCENA have not published independent reproductions of those tests. In a retrieval-augmented generation test, Samsung said 10 near-memory devices delivered 64 times the query throughput and 65 times the energy efficiency of a host CPU accessing a CXL memory pool. The test used a 512-million-vector FAISS index derived from LAION-5B.

For long-context inference, ServeTheHome reported that Samsung tested an INT8 version of Llama 3.1 70B with a 100,000-token context across two servers. Each server contained five MX1 devices and Nvidia RTX Pro 6000 Blackwell GPUs. Samsung reported 17.7 tokens per second with near-memory processing, compared with 5.50 for the baseline, or a 3.35-fold increase. Samsung also reported an increase from 1.12 to 4.31 tokens per kilojoule, a 3.84-fold gain. The MX1 configuration processed missed KV-cache pages near memory instead of returning them to the GPUs. Both tests concentrated on data movement and near-memory processing, the cases most favorable to MX1's design. The disclosed hardware, model, context length and baselines give prospective customers a starting point for evaluating the claims, while leaving open how the card performs across broader production workloads.

XCENA separately reported benchmark results of up to 4.7 times the throughput and 18.7 times the energy efficiency of host processing over CXL on selected analytics kernels. Against local DDR5, XCENA claimed gains of up to 2 times in throughput and 6.2 times in energy efficiency. ServeTheHome reported that the CXL comparison used an Intel Xeon 6767P reference system and excluded idle power. Excluding idle consumption narrows what the energy result says about equipment that remains powered between bursts of useful work.

The production test comes next

MX1 remains a prototype. Kim told TechCrunch in May that mass-production chips were scheduled to roll off Samsung Foundry lines by the end of 2026, with XCENA expecting revenue starting in 2027. The August presentation showed functioning silicon and a detailed programming model, leaving XCENA little time between its conference demonstration and its stated production target.

MX1 was among the FMS 2025 Best of Show winners in the "Most Innovative Technology - Computational Memory" category. The award and functioning samples give XCENA visibility, while production yield, reliability and deployment economics remain unproven.

Kim has identified Astera Labs and Marvell as XCENA's closest competitors, while describing Panmnesia and UnifabriX as more focused on CXL fabrics.

Astera Labs' Leo, a family of CXL smart-memory controllers, supports memory expansion, pooling and sharing. Astera Labs lists production-qualified Aurora A-Series add-in cards with four DDR5 RDIMM slots and up to 2TB of capacity. Leo focuses on connectivity and memory management rather than thousands of programmable near-memory cores.

Marvell and SK hynix have taken a closer architectural route with CMM-Ax, a CXL 2.0 device using 16 Arm Neoverse V2 cores and four DDR5 channels. XCENA's design instead uses a much wider array of lightweight RISC-V cores, PCIe 6 and CXL 3.2, and SSD-backed memory in the same product.

Investors have already funded the production push. Atinum Investment and IMM Investment co-led XCENA's $135 million Series B at a reported $570 million valuation, bringing total funding to $185 million. Corstone Asia, SBI Investment and Mirae Asset Capital also participated. XCENA said the capital would support customer deployments, international expansion and the next generation of its computational-memory technology.

The Samsung relationship gives Kim's team manufacturing credibility and a route into rack-scale system design. Commercial adoption will depend on production yields, software integration and repeatable performance outside the benchmark cases chosen for Hot Chips. MX1 now has enough technical detail to be judged as a system component. XCENA's next milestone is proving that customers will deploy it as one.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @xcena 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/xcena-and-samsung-pu…] indexed:0 read:7min 2026-08-30 ·