OpenGEMM: Open-source B200 GEMM kernels OpenGEMM, an open-source project from developer aramesh10, released CUDA GEMM kernels for NVIDIA's B200 GPU (sm_100a), installable via pip and requiring PyTorch 2.8+ and CUDA 12.9+. The library emits standalone CUDA kernels for dtypes including bf16, f16, tf32, s8, u8, e4m3, e5m2, e3m2, e2m3, and e2m1, plus block-scaled nvfp4, mxfp8, and mxfp4 formats, and compiles its two libraries on the first gemm() call in about 25 seconds. Optimized configs are stored in configs.json, with unoptimized shapes autotuned and saved locally to ./opengemm-configs/tuned_configs.json or the OPENGEMM_CONFIGS environment variable. GEMM kernels for NVIDIA B200 sm 100a in CUDA. python import opengemm as og c = og.gemm a, b C M, N = A M, K @ B N, K .T c = og.gemm a, b, sfa, sfb block-scaled: nvfp4, mxfp8, mxfp4 og.emit kernel a, b, file="k.cu" emits .cu/.cuh for this shape c = og.run kernel "k.cu", a, b compiles emitted kernel and runs it Check API.md https://github.com/aramesh10/OpenGEMM/blob/main/API.md for documentation. From PyPI pip install opengemm From a clone: git clone https://github.com/aramesh10/OpenGEMM.git cd OpenGEMM pip install -e . Requirements: - sm 100a - PyTorch 2.8+ - CUDA 12.9+ The kernels are compiled into two libraries on the first gemm call and takes ~25s. Run python -m opengemm or from python og.prebuild to pay the cost at install time instead. Give your agent this prompt to use OpenGEMM as a tool: OpenGEMM emits standalone CUDA GEMM kernels for B200 sm 100a , no GPU needed to emit: python -c " import opengemm as og S = dict m=1024, n=1024, k=1024 og.emit kernel S, atype='bf16', file='k' writes k.cu and k.cuh og.emit kernel S, atype='e4m3', btype='e5m2' mixed, names itself og.emit kernel S, atype='e2m1', sftype='ue4m3' block-scaled nvfp4 src, hdr = og.emit kernel S, atype='bf16' the text, always returned print src, hdr " atype / btype: bf16 f16 tf32 s8 u8 e4m3 e5m2 e3m2 e2m3 e2m1 sftype block-scaled : ue4m3 nvfp4 or ue8m0 mxfp8, mxfp4 C M, N = A M, K @ B N, K .T . Both operands are row-major with K innermost. | GEMM | atype / btype | sftype | output | torch.dtype in → out | |---|---|---|---|---| | bfloat16 | bf16 | — | f32 | bfloat16 → float32 | | float16 | f16 | — | f32 | float16 → float32 | | tf32 | tf32 | — | f32 | float32 → float32 | | int8 | s8 | — | s32 | int8 → int32 | | uint8 | u8 | — | s32 | uint8 → int32 | | fp8 | e4m3 | — | f32 | float8 e4m3fn → float32 | | fp8 | e5m2 | — | f32 | float8 e5m2 → float32 | | mixed fp8 | e4m3, e5m2 | — | f32 | float8 e4m3fn , float8 e5m2 → float32 | | fp6 | e3m2 | — | f32 | uint8 → float32 | | fp6 | e2m3 | — | f32 | uint8 → float32 | | fp4 | e2m1 | — | f32 | uint8 → float32 | | nvfp4 | e2m1 | ue4m3 per 16 | bf16 | float4 e2m1fn x2 , float8 e4m3fn → bfloat16 | | mxfp8 | e4m3 | ue8m0 per 32 | bf16 | float8 e4m3fn , float8 e8m0fnu → bfloat16 | | mxfp4 | e2m1 | ue8m0 per 32 | bf16 | float4 e2m1fn x2 , float8 e8m0fnu → bfloat16 | Note: fp6 and fp4 have no torch dtype. They arrive densely packed in uint8 and are named - gemm a, b, atype="e2m1" Use btype= when the two operands differ. Output is M, N , row-major, like torch.mm . There is no heursitic to choose the config. Optimized configs are stored in configs.json . If a particular shape has not been optimized, the library autotunes and returns and saves the best config locally to ./opengemm-configs/tuned configs.json or to OPENGEMM CONFIGS env variable. CUDA VISIBLE DEVICES=0 python scripts/tune.py --dtype f16 --shape 4096 4096 4096 CUDA VISIBLE DEVICES=0 python scripts/benchmark.py --dtype bf16 e4m3 vs cuBLAS CUDA VISIBLE DEVICES=0 python scripts/test.py correctness tune.py ablates every compiled configuration for a shape and records the best performing config to configs.json python scripts/emit kernel.py --dtype e4m3 --shape 4096 4096 4096 --file emitted/e4m3 4k.cu python scripts/run kernel.py emitted/e4m3 4k.cu correctness, then timing vs cuBLAS OpenGEMM can also emit the optimized CUDA files for a kernel given a shape and dtype. It can be ran with scripts/run kernel.py or built with nvcc : nvcc -O3 -std=c++20 -gencode=arch=compute 100a,code=sm 100a --expt-relaxed-constexpr -shared -Xcompiler -fPIC -lcuda