PTXBench: What about just CUDA-PTX? PTXBench, a benchmark released August 17, 2026 with code on GitHub, evaluates how well frontier LLMs can write architecture-specific CUDA-PTX for H100 and B200 GPUs, finding GEMM nearly solved while attention lags. Claude Opus 4.8 reached 1.012x cuBLAS performance on Blackwell and Gemini 3.1 Pro reached 0.892x, measured against cuBLAS 13.1, cuDNN 9.20.0 and FlashInfer 0.6.14. Adapting Qwen3.6-27B with a 158-record Fixit dataset solved all five primary evaluation problems, though the model still fails on head-dimension-96 backward attention. PTXBench: What about just CUDA-PTX? August 17, 2026 The code for PTXBench is available here https://github.com/zhang677/PTXBench . Why not directly generate CUDA-PTX? PTX is the lowest-level GPU interface that CUDA programmers can explicitly control, so generating CUDA with inline, architecture-specific PTX offers the shortest path from a new hardware feature to a working kernel. This approach has traditionally looked unattractive: PTX is difficult to program and validate, while abstractions such as Triton and CuTeDSL provide productivity and portability. Yet those abstractions must continually absorb new instructions, layouts, and synchronization mechanisms through compiler engineering. As GPU architectures evolve faster and LLMs become better at code generation and iterative repair, directly generating CUDA-PTX becomes appealing as a way to use new hardware capabilities before the higher-level software stack fully catches up. What does PTXBench ask? PTXBench asks how well current LLMs can reason about architecture-specific PTX on H100 and B200 GPUs, not merely whether they can emit a fast CUDA kernel. A model receives an architecture-specific knowledge pack, writes CUDA-PTX, and revises it over multiple turns using compilation, sanitization, correctness, and performance feedback. The evaluation framework separately checks whether the kernel is functionally correct, whether the requested instruction family actually executes at runtime, and whether the kernel is competitive with frontier libraries. This separation also points to the techniques that will matter next: execution-grounded repair, targeted post-training, runtime instruction verification, and much stronger testing infrastructure. Takeaway 1: GEMM is close; attention is not Frontier LLMs are beginning to make architecture-specific PTX work, but capability falls sharply as the workload becomes more complex. GEMM is closest to being solved: Claude Opus 4.8 reaches 1.012x cuBLAS performance on Blackwell, while Gemini 3.1 Pro reaches 0.892x. Attention remains substantially harder, especially on Blackwell, and backward attention is harder still. PTXBench measures performance against frontier libraries: cuBLAS 13.1 for GEMM, cuDNN 9.20.0 for the primary attention workloads, and FlashInfer 0.6.14 for GQA. Speedup is the reference-library latency divided by the generated kernel’s latency, so 1.0x means matching the corresponding performance baseline and values above 1.0x mean surpassing it. Figure 1 uses two different kinds of ratios. The x-axis p is a speedup threshold relative to the library baseline; the y-axis, Fast