cd/entity/AITER· home entities AITER
grep -l @aiter /news/*.json | wc -l → 9

AITER

mentions 9 type Organization feed RSS

// recent coverage 9 mentions

13:10
2026-08-04
sourcefeed.dev
artificial-intelligence

DeepSeek V4 Flash on One AMD GPU Took Nine Patches

A single AMD MI300X GPU with 192 GB of HBM3 now serves DeepSeek's 284B-parameter DeepSeek-V4-Flash-0731 checkpoint in mixed FP4+FP8 format, requiring nine patch overlays against a vLLM ROCm nightly pl…

10:00
2026-08-04
github.com
artificial-intelligence

DeepSeek V4 Flash on a Single AMD MI300X

A production configuration for running DeepSeek-V4-Flash-0731 on a single AMD MI300X GPU achieves 168.6 tok/s median single-stream decode and 542 tok/s aggregate across 8 concurrent streams, with the …

00:00
2026-07-23
rocm.blogs.amd.com
artificial-intelligence

Serve Kimi-K2.5-MXFP4 on MI355X with ATOM

AMD shows how to serve the pre-quantized amd/Kimi-K2.5-MXFP4 checkpoint on AMD Instinct MI355X GPUs using ATOM, a lightweight vLLM-like framework that integrates AITER kernels and exposes an OpenAI-co…

19:03
2026-06-30
rocm.blogs.amd.com
large-language-models

Accelerating LLM Inference on AMD GPUs with Low-Latency GEMMs

AMD announced a new kernel family, LDS-Pipelined Split-K GEMM, that accelerates LLM inference on AMD GPUs by optimizing decode-time GEMMs with small M and large N/K dimensions. The technique achieves …

11:20
2026-06-21
dev.to
large-language-models

AMD ATOM + ATOMesh: Prefill/decode Disaggregation on ROCm

AMD shipped ATOM + ATOMesh, a ROCm-native LLM serving stack for Instinct GPUs that implements prefill/decode disaggregation, splitting the two inference phases onto separate GPU pools to optimize for …

// co-occurs with top 8 entities
// topics top 6 topics