cd/entity/ROCm· home› entities› ROCm
grep -l @rocm /news/*.json | wc -l → 146

ROCm

mentions 146 type Organization page 2/8 feed RSS

// recent coverage 146 mentions

00:00
2026-09-08
rocm.blogs.amd.com
artificial-intelligence

veRL on AMD: Production-Ready RL Post-Training on ROCm

AMD and the veRL project released a production-ready reinforcement learning post-training container for AMD Instinct GPUs, supporting MI300 and MI355 series on ROCm, with AITER-accelerated vLLM and SG…

09:26
2026-09-07
vllm.ai
large-language-models

Speculative Decoding in vLLM on AMD GPUs

AMD's experiments with speculative decoding in vLLM on AMD Instinct MI300X and MI355X GPUs using the ROCm platform show that output-token throughput gains vary by drafting method, proposal length, mod…

13:00
2026-09-06
vettedconsumer.com
large-language-models

Why Is My Local LLM So Slow? The 6 Bottlenecks, in Order

A new guide from Vetted Consumer identifies six bottlenecks that slow local large language models, ranked by impact, with the top cause being the model not fitting in fast memory, which can drop speed…

08:57
2026-09-04
forum.level1techs.com
artificial-intelligence

ROCm 10.0 on Polaris

A developer has ported ROCm 10.0 to AMD's gfx803 architecture (Polaris GPUs), enabling vLLM and llama.cpp to run on older cards like the RX 580. The port, which builds on earlier work with ROCm 6.4.4 …

22:10
2026-09-03
forum.level1techs.com
artificial-intelligence

Franken Strix Halo: 2x R9700s + 128gb Strix Halo unified memory

An enthusiast experiment combined a Strix Halo APU (Ryzen AI MAX+ 395 with 128 GB unified memory and Radeon 8060S iGPU) with two Radeon R9700 32 GB discrete GPUs via an external PCIe switch, achieving…

19:28
2026-09-03
forum.level1techs.com
artificial-intelligence

I need an AI sanity check (9700 Pro x2)

A Level1Techs forum user reports that running two AMD Radeon RX 9700 Pro GPUs on Windows with llama.cpp and Qwen 3.6 27B Q4_K_M can double token generation speed for complex prompts, reaching 37.8 tok…

00:00
2026-09-01
rocm.blogs.amd.com
artificial-intelligence

Optimizing ATOM and vLLM-ATOM for High-Interactivity Inference

AMD's ATOM and vLLM-ATOM inference engines were optimized for high-interactivity LLM inference, targeting responsiveness for single users rather than batch throughput. The optimizations focus on remov…

03:13
2026-08-29
forum.level1techs.com
artificial-intelligence

MoE with little models

A forum user comparing CUDA and ROCm for local AI inference reports that ROCm issues have diminished and is considering an all-AMD build with dual R9700 GPUs by 2027, citing AMD's lower cost. Another …

20:23
2026-08-25
discuss.huggingface.co
artificial-intelligence

Building Local: My 2026 Headless AI Server Journey

A developer reports that running Qwen 3.8 27B at Q5_K_M quantization on a dual AMD Radeon RX 7900 XT and 7800 XT setup achieves 20 tokens per second with a 256k context window, enabling autonomous mul…

00:30
2026-08-25
servethehome.com
artificial-intelligence

AMD MI400 GPU at Hot Chips 2026

AMD detailed the Instinct MI400 series GPU architecture at Hot Chips 2026, highlighting the MI455X silicon with 432 GB of HBM4 memory, 23.3 TB/s bandwidth, and a peak MXFP4 performance of 40.26 petafl…

07:00
2026-08-24
hiraditya.github.io
artificial-intelligence

A Bug Is a Violation of a Specification

A bug is a violation of a specification, and no specification exists that prefix caching's variable logits violate, according to an analysis of vLLM and SGLang issues. The vLLM PR #34046 adds an opt-i…

22:10
2026-08-22
forum.level1techs.com
artificial-intelligence

What i run with my Strix Halo

A user reports running large language models on a 128GB Bosgame M5 Strix Halo mini PC purchased used for 1800€ on eBay, achieving 30 tokens per second decode with Qwen 3.8 27B and 50 tokens per second…

16:53
2026-08-20
jdagostino.github.io
artificial-intelligence

AI at Home Part 2: Multi-GPU Drifting

A developer building a home AI server from e-waste GPUs details the process of optimizing multi-GPU performance for running large language models, focusing on llama.cpp settings and existing technique…

← prev page 2 / 8 next →
// co-occurs with top 8 entities
// topics top 6 topics