GEMM kernels for NVIDIA B200 (sm_100a) in CUDA.
import opengemm as og
c = og.gemm(a, b) # C[M, N] = A[M, K] @ B[N, K].T
c = og.gemm(a, b, sfa, sfb) # block-scaled: nvfp4, mxfp8, mxfp4
og.emit_kernel(a, b, file="k.cu") # emits .cu/.cuh for this shape
c = og.run_kernel("k.cu", a, b) # compiles emitted kernel and runs it
Check API.md for documentation.
From PyPI
pip install opengemm
From a clone:
git clone https://github.com/aramesh10/OpenGEMM.git
cd OpenGEMM
pip install -e .
Requirements:
- sm_100a
- PyTorch 2.8+
- CUDA 12.9+
The kernels are compiled into two libraries on the first gemm() call and takes ~25s. Run python -m opengemm or from python og.prebuild() to pay the cost at install time instead.
Give your agent this prompt to use OpenGEMM as a tool:
OpenGEMM emits standalone CUDA GEMM kernels for B200 (sm_100a), no GPU
needed to emit:
python -c "
import opengemm as og
S = dict(m=1024, n=1024, k=1024)
og.emit_kernel(**S, atype='bf16', file='k') # writes k.cu and k.cuh
og.emit_kernel(**S, atype='e4m3', btype='e5m2') # mixed, names itself
og.emit_kernel(**S, atype='e2m1', sftype='ue4m3') # block-scaled (nvfp4)
src, hdr = og.emit_kernel(**S, atype='bf16') # the text, always returned
print(src, hdr)
"
atype / btype: bf16 f16 tf32 s8 u8 e4m3 e5m2 e3m2 e2m3 e2m1
sftype (block-scaled): ue4m3 (nvfp4) or ue8m0 (mxfp8, mxfp4)
C[M, N] = A[M, K] @ B[N, K].T. Both operands are row-major with K innermost.
| GEMM | atype /btype |
sftype |
output | torch.dtype (in β out) |
|---|---|---|---|---|
| bfloat16 | bf16 | β | f32 | bfloat16 βfloat32 |
| float16 | f16 | β | f32 | float16 βfloat32 |
| tf32 | tf32 | β | f32 | float32 βfloat32 |
| int8 | s8 | β | s32 | int8 βint32 |
| uint8 | u8 | β | s32 | uint8 βint32 |
| fp8 | e4m3 | β | f32 | float8_e4m3fn βfloat32 |
| fp8 | e5m2 | β | f32 | float8_e5m2 βfloat32 |
| mixed fp8 | e4m3, e5m2 | β | f32 | float8_e4m3fn ,float8_e5m2 βfloat32 |
| fp6 | e3m2 | β | f32 | uint8 βfloat32 |
| fp6 | e2m3 | β | f32 | uint8 βfloat32 |
| fp4 | e2m1 | β | f32 | uint8 βfloat32 |
| nvfp4 | e2m1 | ue4m3 (per 16) | bf16 | float4_e2m1fn_x2 ,float8_e4m3fn βbfloat16 |
| mxfp8 | e4m3 | ue8m0 (per 32) | bf16 | float8_e4m3fn ,float8_e8m0fnu βbfloat16 |
| mxfp4 | e2m1 | ue8m0 (per 32) | bf16 | float4_e2m1fn_x2 ,float8_e8m0fnu βbfloat16 |
Note: fp6 and fp4 have no torch dtype. They arrive densely packed in uint8 and are named - gemm(a, b, atype="e2m1")
Use btype= when the two operands differ.
Output is [M, N], row-major, like torch.mm.
There is no heursitic to choose the config. Optimized configs are stored in configs.json.
If a particular shape has not been optimized, the library autotunes and returns and saves the best config locally to ./opengemm-configs/tuned_configs.json or to OPENGEMM_CONFIGS env variable.
CUDA_VISIBLE_DEVICES=0 python scripts/tune.py --dtype f16 --shape 4096 4096 4096
CUDA_VISIBLE_DEVICES=0 python scripts/benchmark.py --dtype bf16 e4m3 # vs cuBLAS
CUDA_VISIBLE_DEVICES=0 python scripts/test.py # correctness
tune.py ablates every compiled configuration for a shape and records the best performing config to configs.json
python scripts/emit_kernel.py --dtype e4m3 --shape 4096 4096 4096 --file emitted/e4m3_4k.cu
python scripts/run_kernel.py emitted/e4m3_4k.cu # correctness, then timing vs cuBLAS
OpenGEMM can also emit the optimized CUDA files for a kernel given a shape and dtype. It can be ran with scripts/run_kernel.py or built with nvcc:
nvcc -O3 -std=c++20 -gencode=arch=compute_100a,code=sm_100a --expt-relaxed-constexpr -shared -Xcompiler -fPIC -lcuda <KERNEL_FILE>.cu -o <KERNEL_FILE>.so
The entry point is extern "C" void mm_<dtype>_<M>_<N>_<K>(a, b, c, stream),
or smm_<dtype>_<M>_<N>_<K>(a, b, sfa, sfb, c, stream)
emit_kernel reads only shapes and dtypes, so meta tensors work:
emit_kernel(torch.empty(4096, 4096, dtype=torch.bfloat16, device="meta"), ...).