cd /news/ai-infrastructure/opengemm-open-source-b200-gemm-kerne… Β· home β€Ί topics β€Ί ai-infrastructure β€Ί article
[ARTICLE Β· art-128450] src=github.com β†— pub= topic=ai-infrastructure verified=true sentiment=Β· neutral

OpenGEMM: Open-source B200 GEMM kernels

OpenGEMM, an open-source project from developer aramesh10, released CUDA GEMM kernels for NVIDIA's B200 GPU (sm_100a), installable via pip and requiring PyTorch 2.8+ and CUDA 12.9+. The library emits standalone CUDA kernels for dtypes including bf16, f16, tf32, s8, u8, e4m3, e5m2, e3m2, e2m3, and e2m1, plus block-scaled nvfp4, mxfp8, and mxfp4 formats, and compiles its two libraries on the first gemm() call in about 25 seconds. Optimized configs are stored in configs.json, with unoptimized shapes autotuned and saved locally to ./opengemm-configs/tuned_configs.json or the OPENGEMM_CONFIGS environment variable.

read3 min views2 publishedSep 13, 2026
OpenGEMM: Open-source B200 GEMM kernels
Image: Michielbdejong (auto-discovered)

GEMM kernels for NVIDIA B200 (sm_100a) in CUDA.

import opengemm as og

c = og.gemm(a, b)                    # C[M, N] = A[M, K] @ B[N, K].T
c = og.gemm(a, b, sfa, sfb)          # block-scaled: nvfp4, mxfp8, mxfp4

og.emit_kernel(a, b, file="k.cu")    # emits .cu/.cuh for this shape
c = og.run_kernel("k.cu", a, b)      # compiles emitted kernel and runs it

Check API.md for documentation.

From PyPI

pip install opengemm

From a clone:

git clone https://github.com/aramesh10/OpenGEMM.git
cd OpenGEMM
pip install -e .

Requirements:

  • sm_100a
  • PyTorch 2.8+
  • CUDA 12.9+

The kernels are compiled into two libraries on the first gemm() call and takes ~25s. Run python -m opengemm or from python og.prebuild() to pay the cost at install time instead.

Give your agent this prompt to use OpenGEMM as a tool:

OpenGEMM emits standalone CUDA GEMM kernels for B200 (sm_100a), no GPU
needed to emit:

python -c "
import opengemm as og
S = dict(m=1024, n=1024, k=1024)

og.emit_kernel(**S, atype='bf16', file='k')         # writes k.cu and k.cuh
og.emit_kernel(**S, atype='e4m3', btype='e5m2')     # mixed, names itself
og.emit_kernel(**S, atype='e2m1', sftype='ue4m3')   # block-scaled (nvfp4)
src, hdr = og.emit_kernel(**S, atype='bf16')        # the text, always returned
print(src, hdr)
"
atype / btype: bf16 f16 tf32 s8 u8 e4m3 e5m2 e3m2 e2m3 e2m1
sftype (block-scaled): ue4m3 (nvfp4) or ue8m0 (mxfp8, mxfp4)

C[M, N] = A[M, K] @ B[N, K].T. Both operands are row-major with K innermost.

GEMM atype /btype sftype output torch.dtype (in β†’ out)
bfloat16 bf16 β€” f32 bfloat16 β†’float32
float16 f16 β€” f32 float16 β†’float32
tf32 tf32 β€” f32 float32 β†’float32
int8 s8 β€” s32 int8 β†’int32
uint8 u8 β€” s32 uint8 β†’int32
fp8 e4m3 β€” f32 float8_e4m3fn β†’float32
fp8 e5m2 β€” f32 float8_e5m2 β†’float32
mixed fp8 e4m3, e5m2 β€” f32 float8_e4m3fn ,float8_e5m2 β†’float32
fp6 e3m2 β€” f32 uint8 β†’float32
fp6 e2m3 β€” f32 uint8 β†’float32
fp4 e2m1 β€” f32 uint8 β†’float32
nvfp4 e2m1 ue4m3 (per 16) bf16 float4_e2m1fn_x2 ,float8_e4m3fn β†’bfloat16
mxfp8 e4m3 ue8m0 (per 32) bf16 float8_e4m3fn ,float8_e8m0fnu β†’bfloat16
mxfp4 e2m1 ue8m0 (per 32) bf16 float4_e2m1fn_x2 ,float8_e8m0fnu β†’bfloat16

Note: fp6 and fp4 have no torch dtype. They arrive densely packed in uint8 and are named - gemm(a, b, atype="e2m1") Use btype= when the two operands differ.

Output is [M, N], row-major, like torch.mm.

There is no heursitic to choose the config. Optimized configs are stored in configs.json. If a particular shape has not been optimized, the library autotunes and returns and saves the best config locally to ./opengemm-configs/tuned_configs.json or to OPENGEMM_CONFIGS env variable.

CUDA_VISIBLE_DEVICES=0 python scripts/tune.py --dtype f16 --shape 4096 4096 4096
CUDA_VISIBLE_DEVICES=0 python scripts/benchmark.py --dtype bf16 e4m3    # vs cuBLAS
CUDA_VISIBLE_DEVICES=0 python scripts/test.py                           # correctness

tune.py ablates every compiled configuration for a shape and records the best performing config to configs.json

python scripts/emit_kernel.py --dtype e4m3 --shape 4096 4096 4096 --file emitted/e4m3_4k.cu
python scripts/run_kernel.py emitted/e4m3_4k.cu       # correctness, then timing vs cuBLAS

OpenGEMM can also emit the optimized CUDA files for a kernel given a shape and dtype. It can be ran with scripts/run_kernel.py or built with nvcc:

nvcc -O3 -std=c++20 -gencode=arch=compute_100a,code=sm_100a --expt-relaxed-constexpr -shared -Xcompiler -fPIC -lcuda <KERNEL_FILE>.cu -o <KERNEL_FILE>.so

The entry point is extern "C" void mm_<dtype>_<M>_<N>_<K>(a, b, c, stream), or smm_<dtype>_<M>_<N>_<K>(a, b, sfa, sfb, c, stream)

emit_kernel reads only shapes and dtypes, so meta tensors work: emit_kernel(torch.empty(4096, 4096, dtype=torch.bfloat16, device="meta"), ...).

── more in #ai-infrastructure 4 stories Β· sorted by recency
── more on @opengemm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/opengemm-open-source…] indexed:0 read:3min 2026-09-13 Β· β€”