cd/sources/rocm-auto-discovered· home sources Rocm (auto-discovered)
cat /sources/rocm-auto-discovered.feed | wc -l → 45

Rocm (auto-discovered)

articles 45 domain rocm.blogs.amd.com → page 2/3 feed RSS
00:00
2026-07-16
rocm.blogs.amd.com
artificial-intelligence

Multi-Accelerator Support for AIMs and AMD Solution Blueprints

AMD released version 2.2 of its enterprise AI reference stack, introducing multi-accelerator support for AMD Inference Microservices (AIMs) and AMD Solution Blueprints across AMD Instinct GPUs (MI300X…

00:00
2026-07-13
rocm.blogs.amd.com
artificial-intelligence

Triton-Based Optimization of Video Sparse Attention on ROCm

AMD has released a Triton-based optimization for video sparse attention on its ROCm platform, targeting Diffusion Transformers (DiTs) used in video generation. The implementation reduces the quadratic…

00:00
2026-07-13
rocm.blogs.amd.com
artificial-intelligence

GEAK Agent-Driven Optimization of the DeepSeekV4 MLA Kernel

AMD's open-source GEAK agent-driven framework automated the optimization of the DeepSeekV4 MLA kernel, achieving a 2.10x improvement in end-to-end throughput and a 3.71x reduction in time-to-first-tok…

00:00
2026-07-13
rocm.blogs.amd.com
artificial-intelligence

QuickReduce INT3 Quantization and Benchmarking on MI355

AMD's QuickReduce library now supports INT3 quantization for all-reduce communication in multi-GPU LLM inference, achieving a 22% reduction in on-wire data volume compared to INT4 on AMD Instinct MI35…

00:00
2026-07-13
rocm.blogs.amd.com
artificial-intelligence

Serving NVFP4 Models on AMD Instinct™ MI355 Accelerators

AMD has integrated an NVFP4 emulation pipeline into vLLM that enables AMD Instinct MI355 accelerators to serve standard NVFP4 quantized checkpoints directly, dequantizing weights to BF16 on-the-fly at…

00:00
2026-07-08
rocm.blogs.amd.com
artificial-intelligence

SGLang-ATOM: Bring ROCm-Native Acceleration to SGLang Serving

AMD introduced SGLang-ATOM, a bridge connecting the SGLang serving framework with ATOM's ROCm-native execution path to accelerate large language model inference on AMD Instinct GPUs. The integration u…

00:00
2026-07-08
rocm.blogs.amd.com
machine-learning

Towards Feature Complete Triton Support in JAX-Triton

AMD contributed a compatibility update to JAX-Triton that supports most Triton features, enabling users to run virtually any Triton or Gluon kernel inside JAX with minimal changes. The update includes…

00:00
2026-07-06
rocm.blogs.amd.com
machine-learning

Primus Tuning Agent: Closing the Configuration-Search Loop

AMD has released the Primus Tuning Agent, a tool that automatically searches for optimal training configurations for large language models by using a projection engine as a fast scoring oracle. In a c…

19:03
2026-06-30
rocm.blogs.amd.com
large-language-models

Accelerating LLM Inference on AMD GPUs with Low-Latency GEMMs

AMD announced a new kernel family, LDS-Pipelined Split-K GEMM, that accelerates LLM inference on AMD GPUs by optimizing decode-time GEMMs with small M and large N/K dimensions. The technique achieves …

← prev page 2 / 3 next →