cd/entity/AITER· home entities AITER
grep -l @aiter /news/*.json | wc -l → 12

AITER

mentions 12 type Organization feed RSS

// recent coverage 12 mentions

00:00
2026-09-08
rocm.blogs.amd.com
artificial-intelligence

veRL on AMD: Production-Ready RL Post-Training on ROCm

AMD and the veRL project released a production-ready reinforcement learning post-training container for AMD Instinct GPUs, supporting MI300 and MI355 series on ROCm, with AITER-accelerated vLLM and SG…

00:00
2026-09-01
rocm.blogs.amd.com
artificial-intelligence

Optimizing ATOM and vLLM-ATOM for High-Interactivity Inference

AMD's ATOM and vLLM-ATOM inference engines were optimized for high-interactivity LLM inference, targeting responsiveness for single users rather than batch throughput. The optimizations focus on remov…

19:57
2026-08-20
forum.level1techs.com
artificial-intelligence

DeepSeek V4 Flash on 8× AMD gfx1201: packaged TP=8 deployment

DeepSeek V4 Flash, a 284B-parameter mixture-of-experts model with 256 routed experts and FP4 expert weights, was successfully deployed on eight AMD Radeon AI PRO R9600D GPUs (32 GB each, 256 GB total)…

13:10
2026-08-04
sourcefeed.dev
artificial-intelligence

DeepSeek V4 Flash on One AMD GPU Took Nine Patches

A single AMD MI300X GPU with 192 GB of HBM3 now serves DeepSeek's 284B-parameter DeepSeek-V4-Flash-0731 checkpoint in mixed FP4+FP8 format, requiring nine patch overlays against a vLLM ROCm nightly pl…

10:00
2026-08-04
github.com
artificial-intelligence

DeepSeek V4 Flash on a Single AMD MI300X

A production configuration for running DeepSeek-V4-Flash-0731 on a single AMD MI300X GPU achieves 168.6 tok/s median single-stream decode and 542 tok/s aggregate across 8 concurrent streams, with the …

00:00
2026-07-23
rocm.blogs.amd.com
artificial-intelligence

Serve Kimi-K2.5-MXFP4 on MI355X with ATOM

AMD shows how to serve the pre-quantized amd/Kimi-K2.5-MXFP4 checkpoint on AMD Instinct MI355X GPUs using ATOM, a lightweight vLLM-like framework that integrates AITER kernels and exposes an OpenAI-co…

19:03
2026-06-30
rocm.blogs.amd.com
large-language-models

Accelerating LLM Inference on AMD GPUs with Low-Latency GEMMs

AMD announced a new kernel family, LDS-Pipelined Split-K GEMM, that accelerates LLM inference on AMD GPUs by optimizing decode-time GEMMs with small M and large N/K dimensions. The technique achieves …

11:20
2026-06-21
dev.to
large-language-models

AMD ATOM + ATOMesh: Prefill/decode Disaggregation on ROCm

AMD shipped ATOM + ATOMesh, a ROCm-native LLM serving stack for Instinct GPUs that implements prefill/decode disaggregation, splitting the two inference phases onto separate GPU pools to optimize for …

// co-occurs with top 8 entities
// topics top 6 topics