# Compilers for Machine Learning

> Source: <https://gist.github.com/geohot/4768597d9dc536446ee2d5de1f29e89d>
> Published: 2026-10-01 08:44:17+00:00

A hands-on one semester course where students build their own compiler from scratch, starting from elementwise programs and ending with training SOTA LLMs on GPUs.

`uops` · `elementwise ops` · `symbolic` · `rewrites` · `renderer`

Students become familiar with the `UOp` and write a compiler capable of compiling and simplifying elementwise programs. We introduce the graph structure and rewriting, and they can compile simple CPU programs to C.

`loops` · `movement ops` · `reduction` · `rangeify`

Now we introduce movement ops and loops. Without memory hierarchies, things are slow — but this compiler is now capable of compiling *any* model, splitting it into kernels, and compiling it to C.

`memory hierarchies` · `upcasting` · `fast GEMMs` · `convs`

Now things get fast. Still only on CPU, but we can now produce SOTA-competitive C code for GEMMs, reduces, and convs.

`call` · `GPUs` · `hardware accelerators` · `tensor cores`

Here we add kernels, GPUs, axis mappings, and tensor cores. This can produce decent torch-competitive CUDA/HIP code now.

`real models` · `autodiff`

Here we implement a real LLM + autodiff to train models.

Use your compiler to implement a paper, port it to strange hardware, mostly anything

| Weeks | Topics | Milestone | 
|---|---|---|
| 1–2 | UOps, elementwise ops, symbolic, rewrites | Compile & simplify elementwise programs to C | 
| 3–4 | Loops, movement ops, reduction, rangeify | Compile any model to C (slow, but correct) | 
| 5–6 | Memory hierarchies, upcasting, GEMMs, convs | SOTA-competitive CPU code | 
| 7–8 | Call, GPUs, accelerators, tensor cores | Torch-competitive CUDA/HIP code | 
| 9–10 | Real models, autodiff | Train a real LLM | 
| 11+ | Student projects | Choose a project and extend your compiler |
