{"slug": "compiling-triton-kernels-without-the-triton-compiler", "title": "Compiling Triton kernels without the Triton compiler", "summary": "A September 29, 2026 arXiv paper reports that an LLM agent translating Triton kernels directly into NVIDIA PTX, a process the authors call \"AI lowering,\" achieved 0.83x to 3.34x the performance of autotuned Triton across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels from recent ML papers. The largest gains came from transformations Triton's lowering pipeline does not perform, including decoding packed binary weights directly into Tensor Core operands (3.34x on BitDelta), assigning each thread a complete softmax row in tensor memory (1.37x on FlashAttention), and reusing overlapping convolution windows (up to 2.23x). The work extends the Volta PTX verifier to support Blackwell's tcgen05 Tensor Core interface by modeling managed tensor memory, descriptor-based operand layouts, and asynchronous execution via commits, waits, memory barriers, and proxy fences.", "body_md": "# Computer Science > Artificial Intelligence\n\n  [Submitted on 29 Sep 2026]\n\n# Title:AI as a Compiler: Compiling Triton kernels without the Triton compiler\n\n[View PDF](https://arxiv.org/pdf/2609.36800)\n\n[HTML (experimental)](https://arxiv.org/html/2609.36800v1)\n\nAbstract:Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We study AI lowering from Triton to NVIDIA PTX: an LLM agent translates Triton kernels directly into PTX. We build an environment that evaluates candidate PTX, and an agentic harness in which an LLM translates Triton kernels into PTX. Across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels from recent ML papers, AI lowering achieves 0.83x-3.34x the performance of autotuned Triton. The largest gains come from transformations that Triton's lowering pipeline does not perform, such as decoding packed binary weights directly into Tensor Core operands (3.34x on BitDelta), assigning each thread a complete softmax row in tensor memory (1.37x on FlashAttention), and reusing overlapping convolution windows (up to 2.23x). These results rely on a robust evaluation harness with comprehensive verification support. We build on Volta, an existing PTX verifier, and substantially extend it to support modern GPU architectures by introducing support for Blackwell's tcgen05 Tensor Core interface. This requires modeling three architectural features: managed tensor memory, descriptor-based operand layouts, and asynchronous execution coordinated through commits, waits, memory barriers, and proxy fences. We discuss the challenges involved in formalizing them, as well as the current limitations. Our results suggest an emerging future in which AI compilers replace custom-written intermediate representations and checkers, reducing the time and engineering effort required to bring up software for new general-purpose and custom chips.\n    \n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/compiling-triton-kernels-without-the-triton-compiler", "canonical_source": "https://arxiv.org/abs/2609.36800", "published_at": "2026-09-30 05:48:53+00:00", "updated_at": "2026-09-30 06:17:59.839862+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-agents", "ai-chips"], "entities": ["Triton", "NVIDIA PTX", "Volta", "BitDelta", "FlashAttention", "Blackwell", "Hopper", "Ada"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/compiling-triton-kernels-without-the-triton-compiler", "markdown": "https://wpnews.pro/news/compiling-triton-kernels-without-the-triton-compiler.md", "text": "https://wpnews.pro/news/compiling-triton-kernels-without-the-triton-compiler.txt", "jsonld": "https://wpnews.pro/news/compiling-triton-kernels-without-the-triton-compiler.jsonld"}}