Enabling Physical AI Agents with Lemonade
AMD and Robotec.ai demonstrated a fully local physical AI agent for robot control by combining the RAI framework, O3DE simulation, ROS 2 Jazzy, and the Lemonade SDK, enabling a robotic arm to pick, pl…
AMD and Robotec.ai demonstrated a fully local physical AI agent for robot control by combining the RAI framework, O3DE simulation, ROS 2 Jazzy, and the Lemonade SDK, enabling a robotic arm to pick, pl…
On a single 8-GPU AMD Instinct MI355X node, AMD served Kimi Linear 48B-A3B with context lengths from 1024 tokens to 64Mi (67,108,864), using vLLM with FP8 KV cache and tensor parallelism 8, recording …
XGBoost, an open-source library implementing gradient-boosted decision trees with a C++ core and CPU/CUDA/HIP backends, is examined in a technical deep dive covering its mathematical foundations and s…
AMD's ROCm blog published a walkthrough for running verl's asynchronous reinforcement learning examples on AMD Instinct MI355X GPUs, covering GRPO on Qwen2.5-VL-7B-Instruct with the Geometry3k dataset…
AMD has launched a multi-part study on memory instruction scheduling for lock-stepped kernels on its Instinct MI300X GPU, beginning with a tiled GEMM kernel example. The series aims to address bandwid…
Anthropic's Claude Code can now run on-premises with AMD Instinct GPUs, serving GLM 5.2 at full quality via SGLang and LiteLLM, eliminating cloud API dependency and per-token costs. The setup uses an …
AUP Learning Cloud, a JupyterHub deployment on Kubernetes accelerated by AMD ROCm, aims to streamline AI education by providing pre-configured GPU-ready notebook environments on AMD hardware, from Ryz…
AMD Quark now supports SVDQuant and native Hugging Face Diffusers integration for diffusion models, enabling 4-bit quantization of both weights and activations. On an AMD Instinct MI350 GPU, plain nat…
AMD introduced VSA (Video Sparse Attention), a hardware-efficient sparse attention mechanism implemented with CK Tile, achieving a 3.31× attention kernel-time speedup at 70% sparsity over FlashAttenti…
AMD introduced the CDNA 5 architecture, the Instinct MI455X GPU with 432 GB HBM4 memory and up to 4x greater AI compute throughput, and the Helios rackscale solution that unifies 72 GPUs into a single…
AMD's Instinct MI450 GPU achieves 85% of peak HBM bandwidth on attention decode kernels using the Gluon kernel optimization framework, according to a technical guide published by AMD. The MI450 series…
The Poro 2 Long family of models, developed by LumiOpen in collaboration with TurkuNLP and the OpenEuroLLM project, extends the context window of the Poro 2 base model to 128k tokens while maintaining…
AMD AI Workbench v2.0.0 now lets users onboard and deploy custom models from Hugging Face or private registries directly through its GUI, serving them as OpenAI-compatible endpoints via the same AMD I…
AMD shows how to serve the pre-quantized amd/Kimi-K2.5-MXFP4 checkpoint on AMD Instinct MI355X GPUs using ATOM, a lightweight vLLM-like framework that integrates AITER kernels and exposes an OpenAI-co…
AMD introduced ROCm Hyperloom, an open-source agentic system that automates inference workload optimization for AMD Instinct GPUs, reducing optimization time from weeks to hours. The system uses a mul…
AMD introduced ROCm AMD Infinity Context (ROCm AIC), a purpose-built KV cache storage tier for distributed inference on AMD Instinct GPUs, starting with MI300X and MI350 Series. ROCm AIC uses low-late…
AMD has open-sourced Spur, a modern GPU job scheduler written in Rust under Apache 2.0, to address the complexity of scheduling AI and HPC workloads on GPU clusters. Spur provides Slurm-compatible CLI…
AMD announced support for Radeon AI PRO R9700 and Radeon PRO W7900 GPUs in technical preview for its enterprise AI reference stack and GPU Operator, enabling deployment of AI workloads including LLM s…
AMD Instinct MI355X GPUs with ATOM and ATOMesh achieve competitive inference performance for MiniMax-M3, a 428-billion-parameter multimodal MoE model, outperforming NVIDIA B200 and B300 in per-GPU thr…
AMD introduced ROCm Infera, an open-source distributed inference reference solution for large-scale deployments, claiming it can improve goodput per GPU for agentic AI workloads by up to 2.6× in inter…