{"slug": "compilers-for-machine-learning", "title": "Compilers for Machine Learning", "summary": "A hands-on one-semester course walks students through building a machine-learning compiler from scratch, progressing from elementwise operations to training a real LLM on GPUs. Over roughly ten weeks, students implement UOps, loop and movement ops, memory hierarchies, fast GEMMs and convolutions, then GPU kernels, tensor cores and autodiff, with the final weeks reserved for student projects. The course is structured around milestones that take the compiler from simple C output to torch-competitive CUDA/HIP code.", "body_md": "A hands-on one semester course where students build their own compiler from scratch, starting from elementwise programs and ending with training SOTA LLMs on GPUs.\n\n`uops` · `elementwise ops` · `symbolic` · `rewrites` · `renderer`\n\nStudents become familiar with the `UOp` and write a compiler capable of compiling and simplifying elementwise programs. We introduce the graph structure and rewriting, and they can compile simple CPU programs to C.\n\n`loops` · `movement ops` · `reduction` · `rangeify`\n\nNow we introduce movement ops and loops. Without memory hierarchies, things are slow — but this compiler is now capable of compiling *any* model, splitting it into kernels, and compiling it to C.\n\n`memory hierarchies` · `upcasting` · `fast GEMMs` · `convs`\n\nNow things get fast. Still only on CPU, but we can now produce SOTA-competitive C code for GEMMs, reduces, and convs.\n\n`call` · `GPUs` · `hardware accelerators` · `tensor cores`\n\nHere we add kernels, GPUs, axis mappings, and tensor cores. This can produce decent torch-competitive CUDA/HIP code now.\n\n`real models` · `autodiff`\n\nHere we implement a real LLM + autodiff to train models.\n\nUse your compiler to implement a paper, port it to strange hardware, mostly anything\n\n| Weeks | Topics | Milestone | \n|---|---|---|\n| 1–2 | UOps, elementwise ops, symbolic, rewrites | Compile & simplify elementwise programs to C | \n| 3–4 | Loops, movement ops, reduction, rangeify | Compile any model to C (slow, but correct) | \n| 5–6 | Memory hierarchies, upcasting, GEMMs, convs | SOTA-competitive CPU code | \n| 7–8 | Call, GPUs, accelerators, tensor cores | Torch-competitive CUDA/HIP code | \n| 9–10 | Real models, autodiff | Train a real LLM | \n| 11+ | Student projects | Choose a project and extend your compiler |", "url": "https://wpnews.pro/news/compilers-for-machine-learning", "canonical_source": "https://gist.github.com/geohot/4768597d9dc536446ee2d5de1f29e89d", "published_at": "2026-10-01 08:44:17+00:00", "updated_at": "2026-10-01 09:46:10.869986+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-infrastructure", "mlops", "developer-tools"], "entities": [], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/compilers-for-machine-learning", "markdown": "https://wpnews.pro/news/compilers-for-machine-learning.md", "text": "https://wpnews.pro/news/compilers-for-machine-learning.txt", "jsonld": "https://wpnews.pro/news/compilers-for-machine-learning.jsonld"}}