cd/sources/rocm-auto-discovered· home› sources› Rocm (auto-discovered)
cat /sources/rocm-auto-discovered.feed | wc -l → 69

Rocm (auto-discovered)

articles 69 domain rocm.blogs.amd.com → page 4/4 feed RSS
19:03
2026-06-30
rocm.blogs.amd.com
large-language-models

Accelerating LLM Inference on AMD GPUs with Low-Latency GEMMs

AMD announced a new kernel family, LDS-Pipelined Split-K GEMM, that accelerates LLM inference on AMD GPUs by optimizing decode-time GEMMs with small M and large N/K dimensions. The technique achieves …

00:00
2026-06-29
rocm.blogs.amd.com
machine-learning

OpenXLA and JAX - ROCm Support and the State of CI

The OpenXLA compiler stack and JAX now run upstream on AMD ROCm, with XLA gating every pull request on real AMD Instinct silicon through GitHub Actions and JAX running hardware tests on every ROCm PR.…

00:00
2026-06-24
rocm.blogs.amd.com
machine-learning

DP Attention and TBO for DeepSeek-V4 on MI355X

AMD introduces DP Attention and Two-Batch Overlap (TBO) optimizations for DeepSeek-V4 inference on MI355X GPUs, using a coordinated prefill scheduler called PrefillDelayer to reduce padding waste and …

00:00
2026-06-19
rocm.blogs.amd.com
large-language-models

A Practical Guide to Running LLMs on AMD Radeon™ GPUs

AMD Radeon GPUs, both integrated and discrete, now support running large language models locally through open-source tools like Lemonade, LM Studio, Ollama, and llama.cpp. A new guide provides step-by…

← prev page 4 / 4