A hands-on one semester course where students build their own compiler from scratch, starting from elementwise programs and ending with training SOTA LLMs on GPUs.
uops · elementwise ops · symbolic · rewrites · renderer
Students become familiar with the UOp and write a compiler capable of compiling and simplifying elementwise programs. We introduce the graph structure and rewriting, and they can compile simple CPU programs to C.
loops · movement ops · reduction · rangeify
Now we introduce movement ops and loops. Without memory hierarchies, things are slow — but this compiler is now capable of compiling any model, splitting it into kernels, and compiling it to C.
memory hierarchies · upcasting · fast GEMMs · convs
Now things get fast. Still only on CPU, but we can now produce SOTA-competitive C code for GEMMs, reduces, and convs.
call · GPUs · hardware accelerators · tensor cores
Here we add kernels, GPUs, axis mappings, and tensor cores. This can produce decent torch-competitive CUDA/HIP code now.
real models · autodiff
Here we implement a real LLM + autodiff to train models.
Use your compiler to implement a paper, port it to strange hardware, mostly anything
| Weeks | Topics | Milestone |
|---|---|---|
| 1–2 | UOps, elementwise ops, symbolic, rewrites | Compile & simplify elementwise programs to C |
| 3–4 | Loops, movement ops, reduction, rangeify | Compile any model to C (slow, but correct) |
| 5–6 | Memory hierarchies, upcasting, GEMMs, convs | SOTA-competitive CPU code |
| 7–8 | Call, GPUs, accelerators, tensor cores | Torch-competitive CUDA/HIP code |
| 9–10 | Real models, autodiff | Train a real LLM |
| 11+ | Student projects | Choose a project and extend your compiler |