AI at Home Part 2: Multi-GPU Drifting
A developer building a home AI server from e-waste GPUs details the process of optimizing multi-GPU performance for running large language models, focusing on llama.cpp settings and existing technique…
A developer building a home AI server from e-waste GPUs details the process of optimizing multi-GPU performance for running large language models, focusing on llama.cpp settings and existing technique…
After one week of running local AI on AMD's Strix Halo hardware in mid-August 2026, the user reports that the hardware is fine but the software stack is far from ready, with crashes and performance is…
PantheonGPU, a new GPU health-check tool, runs 45+ targeted tests on NVIDIA CUDA and AMD ROCm hardware to detect stability issues and configuration bottlenecks that telemetry tools miss. The tool stre…
PantheonGPU 1.0.14, a GPU health testing and AI workload benchmarking tool, is now available as a Debian package and portable bundle for Linux. The tool tests GPU compute, memory, cache, interconnect,…
AMD reported a 4x improvement in AI energy efficiency from 2024 to mid-2026, exceeding its projected 3x target, and remains on track for its 20x rack-scale goal by 2030. The company's 20x30 initiative…
The KV cache has no standard ABI, with vLLM's FlashAttention backend alone reporting its cache shape as a four-dimensional tensor that varies by backend, attention variant, and model family, complicat…
A user reports that ComfyUI on an AMD RX 9070XT with ROCm and PyTorch remains the best combination, and describes creating a bash build script to generate Docker images with specified versions of ROCm…
AMD's ROCm blog published a walkthrough for running verl's asynchronous reinforcement learning examples on AMD Instinct MI355X GPUs, covering GRPO on Qwen2.5-VL-7B-Instruct with the Geometry3k dataset…
GEEKOM, a leading innovator in high-performance Mini PCs, has deployed DeepSeek V4 Flash across four GEEKOM A9 Mega systems, creating a distributed AI cluster that connects via USB4 instead of a data-…
AMD's Ryzen AI Halo may outperform NVIDIA's DGX Spark for local AI development, according to a hardware comparison. The Ryzen AI Halo offers better power efficiency and thermal performance, while the …
AMD and PyTorch upstreamed FP8 training optimizations for AMD Instinct GPUs into TorchAO and TorchTitan, delivering a 13.4% throughput gain over BF16 on Llama3-8B dense models and recovering 89% of FP…
A developer's third installment in a GPU optimization series details implementing distributed training for large language models using CUDA and ROCm, covering All-Reduce, Ring-AllReduce, and ZeRO shar…
A developer's analysis of community reports on running MiniMax H3 locally reveals that GPU model alone is insufficient to predict performance, with VRAM requirements varying widely based on resolution…
AMD announced a definitive agreement on August 6, 2026, to acquire Toronto-based AI-chip startup Taalas, adding model-specific inference silicon to its AI roadmap; financial terms were not disclosed. …
A developer has published a practical guide for setting up a local AI troubleshooting and support environment on macOS and Windows using Open WebUI, Ollama, and Hugging Face. The recommended architect…
AMD reported record second-quarter revenue of $11.5 billion, up 50% year over year, with data center revenue reaching $6.7 billion, up 107%, as CEO Lisa Su said open-source contributions to its ROCm s…
AMD executives say the chip company's open-source strategy gives it an edge over Nvidia in the AI chip race, with CEO Lisa Su noting that open-source contributions to its ROCm software have increased …
Advanced Micro Devices (AMD) forecast third-quarter revenue of roughly $13 billion, surpassing analyst expectations of $12.52 billion, driven by surging demand for its AI data-center chips. The compan…
A single AMD MI300X GPU with 192 GB of HBM3 now serves DeepSeek's 284B-parameter DeepSeek-V4-Flash-0731 checkpoint in mixed FP4+FP8 format, requiring nine patch overlays against a vLLM ROCm nightly pl…
AMD unveiled the Instinct MI455X accelerator at its Advancing AI 2026 event, a flagship GPU with 320 billion transistors, 432GB of HBM4 memory, and up to 40 petaflops of FP4 AI compute, designed to ch…