cd /news/artificial-intelligence/amd-publishes-machine-readable-isa-s… · home topics artificial-intelligence article
[ARTICLE · art-74045] src=snipvote.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

AMD publishes machine-readable ISA so frontier models can write its GPU kernels

AMD has published a machine-readable ISA for its Instinct GPUs and partnered with Anthropic and OpenAI to let frontier models natively write and optimize low-level kernels, claiming a 38% inference speedup over baseline via automated profiling tools. The move directly attacks the CUDA performance moat by enabling PyTorch and JAX workloads on AMD hardware to approach hand-tuned throughput through prompts rather than kernel engineers.

read1 min views1 publishedJul 26, 2026
AMD publishes machine-readable ISA so frontier models can write its GPU kernels
Image: Snipvote (auto-discovered)

Hacker News

AMD publishes machine-readable ISA so frontier models can write its GPU kernels

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

AMD is now shipping machine-readable ISA specs plus an agentic optimization tool (Hyperloom) that plugs into Claude Code, Codex, Cursor, etc., and claims a 38% inference speedup over baseline by auto-generating and tuning kernels on Instinct hardware. If it holds up, this directly attacks the "CUDA moat" performance gap—meaning your existing PyTorch/JAX workloads on AMD could get close to hand-tuned throughput via prompts rather than kernel engineers, changing the cost calculus of moving inference off NVIDIA. Treat the 38% as vendor-claimed and benchmark your own model/rack config before betting deployments on it.

AMD has published its machine-readable GPU ISA and partnered with Anthropic and OpenAI to let frontier models natively write and optimize low-level AMD Instinct kernels, yielding a 38 percent performance boost over baseline using automated profiling tools. In production, this allows your deployment agents to bypass manual CUDA conversion Bottlenecks and auto-tune non-Nvidia hardware on the fly. This significantly lowers the engineering barrier to scaling down inferencing costs on AMD hardware without sacrificing custom kernel-level performance optimizations.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @amd 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/amd-publishes-machin…] indexed:0 read:1min 2026-07-26 ·