{"slug": "jaxbench-benchmarking-autonomous-tpu-kernel-optimization", "title": "JAXBench: Benchmarking Autonomous TPU Kernel Optimization", "summary": "A new benchmark suite called JAXBench, comprising 50 JAX workloads for AI-generated kernel optimization on Google Cloud TPUs, shows that target-specific documentation context matters more than model scale for correctness on the Pallas DSL. Conditioning on curated TPU documentation raised per-sample correctness from 5.8% to 37.3% and solved 48 of 50 benchmarks at a 1.28x geomean speedup, while Autocomp's beam-search pipeline achieved a 1.36x geomean speedup over XLA across the full suite. The benchmark, evaluation harness, and baseline results are released to support open source contributions.", "body_md": "arXiv:2607.20466v1 Announce Type: new\nAbstract: Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equivalent exists for TPUs. We present JAXBench, a TPU-native benchmark suite for AI-generated kernel optimization on Google Cloud TPUs. JAXBench comprises 50 JAX workloads that are both relevant and provide headroom for optimization. We extract 17 production ML operators from architectures in the public MaxText library such as Llama-3.1, DeepSeek-V3, Mixtral, Mamba-2, and AlphaFold2, and translate 33 operators from KernelBench that are validated for correctness and set with new problem sizes that achieve high TPU v6e MXU utilization. Eight of the 17 production operators ship with hand-optimized Pallas kernels from the public Tokamax library and block-size tuned to establish an expert upper-bound baseline. We evaluate four feedback-driven methods on generating candidate Pallas kernels for JAXBench. Across the full suite with Gemini 3 Flash, we find that target-specific context matters more than model scale on a sparsely-documented DSL like Pallas. Conditioning on curated TPU documentation raises per-sample correctness from 5.8% to 37.3% and solves 48 of 50 benchmarks at a 1.28x geomean speedup. Search structure yields significant gains once correctness is achieved, with Autocomp's beam-search pipeline reaching a 1.36x geomean speedup over XLA. On the 8 hand-tuned kernels, Autocomp reaches 1.60x geomean over XLA, recovering most of the 2.08x Tokamax upper bound but trailing on the specialized paged and ragged attention operators. High-quality TPU kernel optimization remains a challenging task, and we release the JAXBench benchmark, evaluation harness, and baseline results to support open source contributions.", "url": "https://wpnews.pro/news/jaxbench-benchmarking-autonomous-tpu-kernel-optimization", "canonical_source": "https://arxiv.org/abs/2607.20466", "published_at": "2026-07-24 04:00:00+00:00", "updated_at": "2026-07-24 04:06:45.932381+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-infrastructure", "ai-tools"], "entities": ["JAXBench", "Google Cloud TPU", "MaxText", "Llama-3.1", "DeepSeek-V3", "Mixtral", "Mamba-2", "AlphaFold2"], "alternates": {"html": "https://wpnews.pro/news/jaxbench-benchmarking-autonomous-tpu-kernel-optimization", "markdown": "https://wpnews.pro/news/jaxbench-benchmarking-autonomous-tpu-kernel-optimization.md", "text": "https://wpnews.pro/news/jaxbench-benchmarking-autonomous-tpu-kernel-optimization.txt", "jsonld": "https://wpnews.pro/news/jaxbench-benchmarking-autonomous-tpu-kernel-optimization.jsonld"}}