NanoGEMM – Bare-Metal AVX2 GEMM in ~100KB, Faster Than NumPy NanoGEMM, a bare-metal AVX2/FMA matrix multiplication engine, claims to outperform NumPy 2.2.3 on small-to-medium matrices, achieving a 2.83x speedup on 16x16 float32 matrices (1.23 µs vs 3.21 µs) but falling behind on larger sizes, with 128x128 and 256x256 matrices running at 0.60x and 0.28x the speed, respectively. The open-source project, available via pip and GitHub, uses register tiling and cache blocking to eliminate BLAS overhead for sub-microsecond CPU inference. NanoGEMM is a minimalist, bare-metal General Matrix Multiplication GEMM engine designed for sub-microsecond CPU inference and high-performance computing in Python. Built with direct AVX2 / FMA 256-bit SIMD assembly-level register tiling and cache blocking, NanoGEMM eliminates the heavy function-call dispatch, thread-pool barriers, and memory-packing overhead of heavyweight BLAS libraries OpenBLAS, MKL for small-to-medium tensors. Measured on Intel/AMD x86-64 CPU AVX2 + FMA against NumPy 2.2.3 single-precision float32 : | Matrix Dimension | NumPy 2.2.3 Latency | NanoGEMM Latency | Speedup Factor | NanoGEMM Throughput | |---|---|---|---|---| | 16 x 16 | 3.21 µs | 1.23 µs C: 0.65 µs | 🚀 2.83x FASTER | 2.83 GFLOPS | | 32 x 32 | 5.75 µs | 2.74 µs C: 2.18 µs | 🚀 2.26x FASTER | 15.13 GFLOPS | | 64 x 64 | 18.70 µs | 17.76 µs C: 16.39 µs | 🚀 1.10x FASTER | 27.08 GFLOPS | | 128 x 128 | 114.07 µs | 182.46 µs | 0.60x | 23.42 GFLOPS | | 256 x 256 | 426.24 µs | 1501.24 µs | 0.28x | 22.35 GFLOPS | 💡 Why is NanoGEMM faster on small/medium matrices? Traditional BLAS engines incur 3–10 µs of fixed overhead per invocation due to dynamic runtime dispatch, argument sanitization, thread synchronization, and packing buffers. NanoGEMM utilizes a zero-allocation, direct register-tiled microkernel that executes in sub-microsecond time immediately upon invocation. - Register allocation: Utilizes 12 ymm registers ymm0 – ymm11 as 256-bit floating-point accumulators storing a$6 \times 16$ tile of matrix$C$ . - Vector broadcast & FMA: Two ymm registers load vectors from$B$ , while individual scalar elements of$A$ are broadcast across ymm using mm256 set1 ps and accumulated via fused multiply-add mm256 fmadd ps . - Zero Spilling: Fits completely inside the 16 available x86-64 YMM registers without stack eviction. - $L 1$ /$L 2$ Cache Tiling:$M c = 64, N c = 128, K c = 128$ to maintain maximum L1/L2 data cache hit ratios and eliminate memory bus thrashing. - Vectorized Edge Handling: Arbitrary matrix dimensions non-multiples of 6 or 16 are processed using boundary SIMD edge loops without padding or buffer allocations. Matrix A M x K Matrix B K x N . . . . . . . . . . . ymm0 . . . . . . . . . . . . . . ymm1 . . . a0 a1 a2 a3 . . x . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . │ │ └──────────────┬──────────────┘ ▼ Matrix C 6 x 16 Tile ymm0 ymm1 - Row 0 ymm2 ymm3 - Row 1 ymm4 ymm5 - Row 2 ymm6 ymm7 - Row 3 ymm8 ymm9 - Row 4 ymm10 ymm11 - Row 5 pip install nanogemm git clone https://github.com/eminsk/nanogemm.git cd nanogemm pip install -e . python import nanogemm as ng import numpy as np Verify SIMD hardware acceleration print "Active ISA:", ng.get simd isa Output: Active ISA: AVX2+FMA 256-bit SIMD, 6x16 register tiling Allocate input matrices A = np.random.randn 32, 64 .astype np.float32 B = np.random.randn 64, 128 .astype np.float32 Direct hardware-accelerated MatMul: C = A @ B C = ng.matmul A, B Or with pre-allocated zero-copy output buffer for maximum performance: out = np.empty 32, 128 , dtype=np.float32 ng.matmul A, B, out=out Standard BLAS SGEMM interface: C = alpha A @ B + beta C res = ng.sgemm A, B, alpha=2.0, beta=0.5, c=out Run the comprehensive correctness test suite comparing NanoGEMM with NumPy reference outputs across random uniforms, normals, non-square dimensions, and prime shapes: python tests/test correctness.py Run the official benchmark against your installed NumPy BLAS: python benchmarks/bench vs numpy.py NanoGEMM is developed by @eminsk https://github.com/eminsk as part of an engineering ecosystem focused on low-level hardware performance, assembly programming, and native desktop computing: - 🎥 screenvideo https://github.com/eminsk/screenvideo — Lightweight desktop screen recorder featuring WASAPI loopback audio and a standalone pure x64 Flat Assembler FASM native edition. - 📊 xlsx vievers https://github.com/eminsk/xlsx vievers — Desktop spreadsheet processor with 80+ formula functions, Chart Wizard, and hardware-accelerated SIMD SSE2 math engine. - 📈 yfinance-ta-patterns https://github.com/eminsk/yfinance-ta-patterns — Candlestick pattern scanner and AI ranking suite powered by TA-Lib and quantitative backtesting. - 🔍 StackOverflowAPI https://github.com/eminsk/StackOverflowAPI — Desktop client for Stack Overflow built with CustomTkinter and native FASM x64 search client. MIT License — Copyright c 2026 eminsk https://github.com/eminsk .