13:54
2026-08-25
github.com
artificial-intelligence
Show HN: Open-source AMDGCN kernels for optimizing LLM inference
Netra Kernel, an open-source ahead-of-time GPU kernel compiler for AMD, achieved 78,498.66 output tokens per second mean throughput serving Qwen3.6-35B-A3B-FP8 on eight AMD MI350X GPUs, within 0.32% oโฆ