KernelBench: Can LLMs Write GPU Kernels? – Benchmark and Toolkit, Torch –> CUDA Stanford's Scaling Intelligence lab released KernelBench, a benchmark and toolkit for evaluating whether large language models can generate correct and efficient CUDA kernels from PyTorch programs, with the HuggingFace dataset updated to v0.1. KernelBench spans 4 levels — 100 single-kernel operators, 100 simple fusion patterns, 50 full model architectures, and HuggingFace model architectures — and scores models with the fast_p metric, the fraction of tasks that are both correct and exceed a speedup threshold p over the PyTorch reference (fast_1 = faster than baseline, fast_2 = at least 2x faster, fast_0 = correctness rate). The repo provides src/eval.py for correctness and timing checks and scripts/run_and_check.py for local or Modal evaluation, and the team says it is extending KernelBench to other DSLs beyond CUDA and to AMD GPU support. A benchmark and environment for evaluating LLMs' ability to generate efficient GPU kernels Specifically we task LLM to generate correct and efficient CUDA / DSL kernels for PyTorch programs on a target GPU. arXiv https://arxiv.org/html/2502.10517v1 | blog post https://scalingintelligence.stanford.edu/blogs/kernelbench/ | HuggingFace Dataset https://huggingface.co/datasets/ScalingIntelligence/KernelBench The latest stable version will be on main branch. We continue to update and improve the repo. The Huggingface dataset https://huggingface.co/datasets/ScalingIntelligence/KernelBench is updated to v0.1. This repo provides core functionality for KernelBench and an easy-to-use set of scripts for evaluation. It is not intended to provide complex agentic scaffolds that solve this task; we recommend cloning and modifying this repo for your experiment, or using it as a git submodule. We structure the problem for LLMs to transpile operators described in PyTorch to CUDA kernels, at whatever level of granularity they desire. We construct KernelBench to have 4 Levels of categories: - Level 1 🧱 : Single-kernel operators 100 Problems The foundational building blocks of neural nets Convolutions, Matrix multiplies, Layer normalization - Level 2 🔗 : Simple fusion patterns 100 Problems A fused kernel would be faster than separated kernels Conv + Bias + ReLU, Matmul + Scale + Sigmoid - Level 3 ⚛️ : Full model architectures 50 Problems Optimize entire model architectures end-to-end MobileNet, VGG, MiniGPT, Mamba - Level 4 🤗 : Level Hugging Face Optimize whole model architectures from HuggingFace We are actively extending KernelBench to other DSLs beyond cuda as well see below , as well as AMD GPU support. To evaluate model-generated kernels, we need to check if they: - are correct ✅ : check against reference torch operators n correctness times on randomized inputs. - are performant ⏱️ : compare against reference torch operators n trial times to measure speedup between runtimes. Check out src/eval.py for details on how we implement correctness check and timing and EVAL.md for notes on evaluation and benchmarking guidelines WIP . We provide a convenient script scripts/run and check.py to evaluate one single sample source code against a reference source code, check correctness and compute speedup. You can use this to evaluate a kernel either locally or remotely by setting eval mode=local or eval mode=modal . Since we need to capture both correctness and performance, we define a metric fast p : fraction of tasks that are both correct and have a speedup greater than threshold p ; speedup is computed as the ratio of PyTorch reference wall-clock time to generated kernel time. Some examples to illustrate this metric that filters based on speedups: - fast 1 is the fraction of tasks that LM-generated kernels are both correct and faster than PyTorch baseline - fast 2 is the fraction of tasks that LM-generated kernels are both correct and at least 2x faster than PyTorch baseline - fast 0 is the fraction of tasks that LM-generated kernels are correct . same as correctness rate You can increase speedup threshold p to make the task more challenging. We provide a script scripts/greedy analysis.py to compute the overall benchmark performance. Since we need to capture both correctness and performance, we use a metric fast p : fraction of tasks that are both correct and have a speedup greater than threshold p ; speedup is computed as the ratio of PyTorch reference wall-clock time to generated kernel time. We organize the repo into the following structure: KernelBench/ ├── assets/ ├── KernelBench/ Benchmark dataset files ├── src/kernelbench/ KernelBench logic code │ ├── unit tests/ │ ├── prompts/ │ ├── .... ├── scripts/ helpful scripts to run the benchmark ├── results/ baseline times across hardware ├── runs/ where your runs will be stored ├── notebooks/ example notebooks for analysis ├── pyproject.toml Project configuration and dependencies We have transitioned to using pyproject.toml and uv for dependency management. Install uv https://docs.astral.sh/uv/getting-started/installation/ if you haven't already Install base dependencies works without a local GPU uv sync Install with AMD ROCm backend ROCm =7.1 is required uv add torch --index pytorch=https://download.pytorch.org/whl/rocm7.1 Install with GPU dependencies for local GPU evaluation uv sync --extra gpu Run commands with uv which invoke the right env uv run python scripts/