{"slug": "amd-publishes-machine-readable-isa-so-frontier-models-can-write-its-gpu-kernels", "title": "AMD publishes machine-readable ISA so frontier models can write its GPU kernels", "summary": "AMD has published a machine-readable ISA for its Instinct GPUs and partnered with Anthropic and OpenAI to let frontier models natively write and optimize low-level kernels, claiming a 38% inference speedup over baseline via automated profiling tools. The move directly attacks the CUDA performance moat by enabling PyTorch and JAX workloads on AMD hardware to approach hand-tuned throughput through prompts rather than kernel engineers.", "body_md": "[Hacker News](https://www.theregister.com/ai-and-ml/2026/07/24/amd-vibe-codes-its-way-past-the-cuda-moat-with-rocmai/5278580)\n\n### AMD publishes machine-readable ISA so frontier models can write its GPU kernels\n\nWhich summary reads better? Pick one — models revealed after.Both summaries are AI-generated.\n\nAMD is now shipping machine-readable ISA specs plus an agentic optimization tool (Hyperloom) that plugs into Claude Code, Codex, Cursor, etc., and claims a 38% inference speedup over baseline by auto-generating and tuning kernels on Instinct hardware. If it holds up, this directly attacks the \"CUDA moat\" performance gap—meaning your existing PyTorch/JAX workloads on AMD could get close to hand-tuned throughput via prompts rather than kernel engineers, changing the cost calculus of moving inference off NVIDIA. Treat the 38% as vendor-claimed and benchmark your own model/rack config before betting deployments on it.\n\nAMD has published its machine-readable GPU ISA and partnered with Anthropic and OpenAI to let frontier models natively write and optimize low-level AMD Instinct kernels, yielding a 38 percent performance boost over baseline using automated profiling tools. In production, this allows your deployment agents to bypass manual CUDA conversion Bottlenecks and auto-tune non-Nvidia hardware on the fly. This significantly lowers the engineering barrier to scaling down inferencing costs on AMD hardware without sacrificing custom kernel-level performance optimizations.", "url": "https://wpnews.pro/news/amd-publishes-machine-readable-isa-so-frontier-models-can-write-its-gpu-kernels", "canonical_source": "https://www.snipvote.com/story/cms1gwtiv0004bmn1s01g2dcq", "published_at": "2026-07-26 07:23:08.989536+00:00", "updated_at": "2026-07-26 07:23:10.887965+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-chips", "ai-tools", "ai-infrastructure"], "entities": ["AMD", "Anthropic", "OpenAI", "Instinct", "CUDA", "PyTorch", "JAX", "Hyperloom"], "alternates": {"html": "https://wpnews.pro/news/amd-publishes-machine-readable-isa-so-frontier-models-can-write-its-gpu-kernels", "markdown": "https://wpnews.pro/news/amd-publishes-machine-readable-isa-so-frontier-models-can-write-its-gpu-kernels.md", "text": "https://wpnews.pro/news/amd-publishes-machine-readable-isa-so-frontier-models-can-write-its-gpu-kernels.txt", "jsonld": "https://wpnews.pro/news/amd-publishes-machine-readable-isa-so-frontier-models-can-write-its-gpu-kernels.jsonld"}}