# AMD ROCm 10: Bringing ROCm.AI’s AI-Native Developer Experiences to AMD Platforms

> Source: <https://newsroom.amd.com/news/rocm-10-software-ai-native-developer-experiences/>
> Published: 2026-08-27 00:00:00+00:00

**Three things to know:**

[ROCm.AI](http://ROCm.AI)is now generally available with AMD ROCm™ 10 software, bringing an AI-driven development platform to help developers build, run and optimize AI workloads.ROCm.AI combines ROCm Hyperloom, AMD Skills and ROCm CLI, bringing AMD expertise, simpler workflows and agentic optimization into the tools developers already use.

ROCm 10 advances the broader ROCm software stack, with a more modular ROCm Core SDK and enhancements across libraries, compilers, frameworks, model support, performance and hardware platforms.

Today, AMD released AMD ROCm™ 10, marking 10 years of the AMD software stack and making [ROCm.AI](http://ROCm.AI) generally available for users. First introduced at [Advancing AI 2026](https://newsroom.amd.com/press-kits/advancing-ai-2026-all-news/), ROCm.AI is an AI-native software experience designed to accelerate development velocity and performance optimization on AMD hardware.

ROCm.AI brings together three core developer experiences: AMD Skills, the ROCm CLI and AMD Hyperloom. Together, they introduce AMD expertise and agentic workflows into the tools developers already use, enabling innovation to evolve at the pace of AI.

Through AI-driven optimization of kernels, memory management and scheduling, a system configured with ROCm.AI delivers an average 3.3x inference improvement and 2.4x training improvement over ROCm™ 7 on the same hardware.[1](#fn-1)[2](#fn-2)

**A New Developer Experience for AI on AMD**

ROCm.AI supports developers across three parts of the AI development process: optimizing inference performance, building with AMD expertise and running AI workloads.

**Optimize End-to-End Inference with AMD ROCm Hyperloom**

ROCm Hyperloom, a key component of ROCm.AI, is an autonomous agentic system for optimizing end-to-end inference workloads across host code and GPU kernels.

Hyperloom profiles workloads, identifies bottlenecks, explores optimization options, implements targeted changes, benchmarks the results, and validates performance and correctness.

With ROCm 10, Hyperloom expands support across AMD Instinct™ GPUs, with support for vLLM and SGLang. Developers can target optimizations across HIP, Triton and FlyDSL, with reports describing proposed code changes and measured or expected performance improvements.

Hyperloom is available through standalone workflows and AMD Skills, giving developers multiple ways to incorporate agentic optimization into their development process.

**Build with AMD Expertise**

AMD Skills brings curated AMD knowledge and validated workflows into supported AI coding agents like Claude Code, Cursor and Codex, giving developers AMD-specific guidance within the tools they already use. With ROCm 10, the AMD Skills catalog expands across three areas:

**Client-native workflows** for local AI and application integration.**Cross-stack workflows** for diagnostics, routing, replay analysis and optimization.**Server-native workflows** for AMD Instinct GPUs and AMD EPYC™ processors, including serving, profiling and performance analysis.

The skills previewed at Advancing AI are now available through the Claude Code, Codex and Cursor marketplaces, as well as through the open catalog on [GitHub](https://github.com/amd/skills). Each shipped skill passes structural and behavioral testing before release to help provide consistent, reliable workflows.

**Run with Simpler Workflows**

The ROCm CLI, a Technology Preview component of ROCm.AI, provides a stable, unified command-line interface for setting up, managing and operating AI workloads on AMD hardware.

Developers can use the same workflows manually, through an AI coding agent or in continuous integration (CI) environments. From a single interface, they can inspect systems, install and manage ROCm environments, serve models, run diagnostics, update components, and control runtimes.

The CLI is available on Windows as well as Linux as a prebuilt binary and does not require an existing ROCm installation. It delivers managed ROCm environments, supports multiple side-by-side runtimes including runtime activation and rollback, as well as integrated model serving and engine management. Adapters today support Lemonade on select AMD client systems and vLLM for AMD Instinct GPU serving.

The ROCm Console (formerly “dash”), included with the CLI, provides a real-time view of system status and workload activity. Developers can monitor ROCm runtime health, model serving, GPU utilization and benchmark telemetry, including metrics such as high bandwidth memory (HBM) usage, power consumption and tokens per watt on supported AMD Instinct systems.

Available as a Technology Preview, ROCm CLI provides a version-agnostic experience across ROCm releases, beginning with ROCm 7.13 software, with ROCm 10 official support coming soon.

**Building on the ROCm 10 Software Stack**

These developer experiences build on the broader ROCm 10 software stack that includes the ROCm Core SDK, providing a more modular foundation for building and running AI workloads across AMD platforms.

ROCm 10 also includes updates to libraries, compilers, frameworks, tools, model support, performance and hardware platforms.

For a deeper look at the ROCm Core SDK, including libraries, compilers, tools, framework and model support, performance improvements and platform updates, read the full [ROCm 10 technical deep dive](https://advanced-micro-devices-rocm-blogs--118.com.readthedocs.build/projects/preview/en/118/ecosystems-and-partners/rocm-x-blog/README.html).

-
(MI350-81) Testing by AMD Performance Labs as of July 7, 2026, measuring the inference performance in tokens per second (TPS) of a system configured with an AMD Instinct MI355x 8x GPU platform and AMD ROCm 7.0 software vs a similarly configured system using a preview version of AMD

[ROCm.ai](http://ROCm.ai)(ROCm 7.2.2 with optimizations such as Optimized Kernels, Parallelism and Scheduling) running GLM-5, Kimi-K2.5, and DeepSeekk-R1-0528 models.Stated performance uplift is expressed as a combined average TPS over across the (3) models tested.

Hardware Configuration

Supermicro AS -4126GS-NMR-LCC (board H14DSG-OD)

8x AMD Instinct MI355X. BIOS AMI v1.4a (2025-04-16), GPU firmware SMC 04.86.11.02, TA RAS 27.69.00.10, TA XGMI 32.00.00.20, RLC43, MEC36, SDMA12, Ubuntu 22.04.2 LTS, kernel 5.15.0-70-generic, amdgpu driver 6.16.6, HOST ROCm 7.1.0

Software Configuration(s)

GLM-5 ROCm Docker Image: rocm/sgl-dev:v0.5.8.post1-rocm700-mi35x-20260219

PYTorch Version 2.8.0, SGLang v0.5.8

Kimi ROCm Docker Image: vllm/vllm-openai-rocm:v0.16.0, vLLM version 0.16.0

DeepSeek-R1 Docker Image: rocm/7.0:...sgl-dev-v0.5.2-rocm7.0-mi35x-20250915, SGLang version V0.5.13

vs

GLM-5 ROCm Docker Image: rocm/atom:rocm7.2.2_ubuntu24.04_py3.12_pytorch_release_2.10.0_

[atom0.1.2.post](http://atom0.1.2.post), ATOM[v0.1.2.post](http://v0.1.2.post)Kimi ROCm Docker Image: vllm/vllm-openai-rocm:v0.22.0, vLLM version V0.22.0

DeepSeek-R1 Docker Image: lmsysorg/sglang-rocm:v0.5.13-rocm720-mi35x-20260612, SGLang version V0.5.13

Server manufacturers may vary configurations, yielding different results. Performance may vary based on configuration, software, vLLM version, and the use of the latest drivers and optimizations. (MI350-81)

-
(MI350-82) Testing by AMD Performance Labs as of July 7, 2026, measuring the training performance in tokens per second (TPS) of AMD ROCm 7.0 software vs a preview version of AMD

[ROCm.ai](http://ROCm.ai)(ROCm 7.2.2 with optimizations such as Optimized Kernels, Parallelism and Scheduling),) using Megatron -LM on a system with 8x AMD Instinct MI355x 8x GPUs GPU platform running DeepSeek-V2-Lite, DeepSeek-V3-16B, and Qwen3-30B-A3B models.Stated performance uplift is expressed as the combined average TPS over across the (3) models tested.

Hardware Configuration

Supermicro AS -4126GS-NMR-LCC (board H14DSG-OD)

8x AMD Instinct MI355. BIOS AMI v1.4a (2025-04-16), GPU firmware SMC 04.86.11.02, TA RAS 27.69.00.10, TA XGMI 32.00.00.20, RLC 43, MEC 36, SDMA 12, Ubuntu 22.04.2 LTS, kernel 5.15.0-70-generic, amdgpu driver 6.16.6, HOST ROCm 7.1.0

Software Configuration(s)

DeepSeek-V2-Lite, ROCm 7.2.1 + Primus v26.3

DeepSeek-V3-16B, ROCm 7.2.1 + Primus v26.3

Qwen3-30B-A3B, ROCm 7.2.1 + Primus v26.3

Server manufacturers may vary configurations, yielding different results. Performance may vary based on configuration, software, and the use of the latest drivers and optimizations. (MI350-82)

Press inquiries: [corporate.pressinquiry@amd.com](mailto:corporate.pressinquiry@amd.com)
