{"slug": "purlin-separating-orchestration-from-the-datapath-of-collectives", "title": "Purlin: Separating Orchestration from the Datapath of Collectives", "summary": "Researchers submitted Purlin, a scale-up GPU collective communication framework that separates orchestration from the datapath, to arXiv on 29 Sep 2026. Purlin specifies collectives as input/output layouts plus copy or reduce operations, coordinates them through a shared protocol called Stage, Notify, And Consume (SNAC), and runs on a hardware-specific datapath called Atom; evaluated on A100, H200, and B200 GPUs across seven collectives, it achieved latency speedups up to 5.14x and bandwidth improvements up to 4.50x over baselines. Integrated into SGLang, Purlin improved offline LLM serving throughput and interactivity by 1.13x on average and up to 1.37x, online LLM inference interactivity by 1.26x on average and up to 2.85x, and cut diffusion image generation end-to-end latency by up to 1.13x.", "body_md": "# Computer Science > Distributed, Parallel, and Cluster Computing\n\n  [Submitted on 29 Sep 2026]\n\n# Title:Purlin: Separating Orchestration from the Datapath of Collectives\n\n[View PDF](https://arxiv.org/pdf/2609.36954)\n\n[HTML (experimental)](https://arxiv.org/html/2609.36954v1)\n\nAbstract:Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing collective implementations often couple semantics, orchestration (where and when data moves), and the datapath (how data moves). This coupling makes it costly to adopt new hardware mechanisms and customize communication for applications. We present Purlin, a scale-up communication framework that separates these concerns. At the top of Purlin, we specify collectives as a naming of an input and output layout and a copy or reduction operation. In the middle, we introduce a shared orchestration protocol, Stage, Notify, And Consume (SNAC), which derives coordination from these specifications. Below SNAC sits a hardware-specific datapath we call Atom, which implements two key data movement primitives for collectives: copy and reduce. This separation lets us customize collectives and adopt new hardware mechanisms while reusing orchestration via SNAC. We evaluate Purlin on A100, H200, and B200 GPUs. Across seven collectives, Purlin achieves latency speedups of up to 5.14x and bandwidth improvements of up to 4.50x over baselines. Integrated into SGLang, Purlin improves offline LLM serving throughput and interactivity by 1.13x on average and up to 1.37x over baselines. For online LLM inference, Purlin improves interactivity by 1.26x on average and up to 2.85x, with the largest gain occurring under overload. For diffusion image generation, Purlin reduces end-to-end latency by up to 1.13x.\n    \n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/purlin-separating-orchestration-from-the-datapath-of-collectives", "canonical_source": "https://arxiv.org/abs/2609.36954", "published_at": "2026-09-30 20:49:35+00:00", "updated_at": "2026-09-30 21:20:18.696991+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-research", "mlops"], "entities": ["Purlin", "Stage, Notify, And Consume (SNAC)", "Atom", "SGLang", "A100", "H200", "B200", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/purlin-separating-orchestration-from-the-datapath-of-collectives", "markdown": "https://wpnews.pro/news/purlin-separating-orchestration-from-the-datapath-of-collectives.md", "text": "https://wpnews.pro/news/purlin-separating-orchestration-from-the-datapath-of-collectives.txt", "jsonld": "https://wpnews.pro/news/purlin-separating-orchestration-from-the-datapath-of-collectives.jsonld"}}