{"slug": "agentic-gpu-programming-for-mlsys", "title": "Agentic GPU Programming for MLSys", "summary": "Carnegie Mellon University plans to integrate a new open-source book, \"Agentic GPU Programming for MLSys,\" into its Machine Learning Systems course series, teaching a compiler-driven harness that lets coding agents write, diagnose, benchmark, and optimize GPU kernels for attention, matrix multiplication, and fused operators. The book is a companion to \"Modern GPU Programming for MLSys\" and is organized into three parts covering agentic programming elements, the TIRx compiler harness with the KCoral benchmark server, and a hands-on tutorial using Grouped GEMM and recorded Kimi Delta Attention runs. Contributions are accepted through the mlc-ai GitHub repository.", "body_md": "# Agentic GPU Programming for MLSys[#](#agentic-gpu-programming-for-mlsys)\n\nMachine learning systems depend on fast GPU kernels for training and serving. Attention, matrix multiplication, and fused operators account for substantial work in these systems. Improving their implementations can reduce the time and resources needed to run a model.\n\nMaking these kernels fast requires in-depth knowledge of algorithms, GPU hardware, and programming models. Coding agents can increasingly carry out this work: finding relevant implementations, writing kernels, interpreting diagnostics, and choosing what to try next. Using these capabilities effectively requires a workflow that coordinates the work and an environment that gives agents access to the knowledge and feedback they need.\n\nThis book introduces the main elements of agentic GPU programming. You will learn how a compiler-driven harness equips an agent to express optimization ideas, diagnose errors, measure performance, and retain the results worth keeping. Together, these capabilities form an effective optimization flow.\n\nThe book develops a compiler foundation, domain-specific program analysis, a knowledge base, and GPU benchmarking and profiling. It then shows how agent workflows compose with this compiler harness to guide an optimization run.\n\nThe material grows out of real-world experience building and using agentic GPU\nprogramming systems. It is a companion book to\n[Modern GPU Programming for MLSys](https://mlc.ai/modern-gpu-programming-for-mlsys/),\nand we plan to integrate it into the\n[Machine Learning Systems](https://mlsyscourse.org/) course series at Carnegie\nMellon University.\n\nThis book is open source. Contributions, corrections, and examples are welcome\nthrough the [GitHub repository](https://github.com/mlc-ai/agentic-gpu-programming-for-mlsys).\n\n## How This Book Is Organized[#](#how-this-book-is-organized)\n\n- **Part I, Elements of Agentic GPU Programming.** This part introduces the\nfundamental elements of agentic GPU programming. It begins with an overview\nof how compiler infrastructure and agent workflows fit together in an\niterative kernel-development process. The following chapters develop the\ncore elements—domain-specific program analysis, reusable knowledge,\nbenchmarking, and profiling—through examples and interactive diagrams. The\npart concludes by comparing different agent workflows and showing how these\nelements can be composed into effective GPU optimization strategies.\n- **Part II, TIRx Harness Overview.** This part introduces TIRx, one concrete\ninstance of a compiler harness and the environment used for the examples in\nthe rest of the book. It walks through the TIRx compiler foundation and the\nKCoral benchmark server with its execution interface.\n- **Part III, Agentic GPU Programming in Action.** This part is a hands-on\ntutorial that puts the preceding chapters to work and carries agentic GPU\nprogramming out end to end. It looks closely at the feedback the harness\nreturns over the course of a run and at the interaction patterns between the\nagent and each element of the harness, and it closes with a few advanced\ntips. Grouped GEMM introduces the workflow; recorded Kimi Delta Attention\nruns supply the diagnostic and review examples.", "url": "https://wpnews.pro/news/agentic-gpu-programming-for-mlsys", "canonical_source": "https://mlc.ai/agentic-gpu-programming-for-mlsys/index.html", "published_at": "2026-09-30 20:54:19+00:00", "updated_at": "2026-09-30 21:20:12.396144+00:00", "lang": "en", "topics": ["machine-learning", "ai-agents", "ai-infrastructure", "mlops", "developer-tools"], "entities": ["Carnegie Mellon University", "Agentic GPU Programming for MLSys", "Modern GPU Programming for MLSys", "Machine Learning Systems course series", "TIRx", "KCoral", "Kimi Delta Attention", "mlc-ai"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/agentic-gpu-programming-for-mlsys", "markdown": "https://wpnews.pro/news/agentic-gpu-programming-for-mlsys.md", "text": "https://wpnews.pro/news/agentic-gpu-programming-for-mlsys.txt", "jsonld": "https://wpnews.pro/news/agentic-gpu-programming-for-mlsys.jsonld"}}