cd /news/ai-infrastructure/purlin-separating-orchestration-from… · home › topics › ai-infrastructure › article
[ARTICLE · art-142827] src=arxiv.org ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Purlin: Separating Orchestration from the Datapath of Collectives

Researchers submitted Purlin, a scale-up GPU collective communication framework that separates orchestration from the datapath, to arXiv on 29 Sep 2026. Purlin specifies collectives as input/output layouts plus copy or reduce operations, coordinates them through a shared protocol called Stage, Notify, And Consume (SNAC), and runs on a hardware-specific datapath called Atom; evaluated on A100, H200, and B200 GPUs across seven collectives, it achieved latency speedups up to 5.14x and bandwidth improvements up to 4.50x over baselines. Integrated into SGLang, Purlin improved offline LLM serving throughput and interactivity by 1.13x on average and up to 1.37x, online LLM inference interactivity by 1.26x on average and up to 2.85x, and cut diffusion image generation end-to-end latency by up to 1.13x.

read2 min views1 publishedSep 30, 2026
Purlin: Separating Orchestration from the Datapath of Collectives
Image: source
  [Submitted on 29 Sep 2026]


[View PDF](https://arxiv.org/pdf/2609.36954)

[HTML (experimental)](https://arxiv.org/html/2609.36954v1)

Abstract:Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing collective implementations often couple semantics, orchestration (where and when data moves), and the datapath (how data moves). This coupling makes it costly to adopt new hardware mechanisms and customize communication for applications. We present Purlin, a scale-up communication framework that separates these concerns. At the top of Purlin, we specify collectives as a naming of an input and output layout and a copy or reduction operation. In the middle, we introduce a shared orchestration protocol, Stage, Notify, And Consume (SNAC), which derives coordination from these specifications. Below SNAC sits a hardware-specific datapath we call Atom, which implements two key data movement primitives for collectives: copy and reduce. This separation lets us customize collectives and adopt new hardware mechanisms while reusing orchestration via SNAC. We evaluate Purlin on A100, H200, and B200 GPUs. Across seven collectives, Purlin achieves latency speedups of up to 5.14x and bandwidth improvements of up to 4.50x over baselines. Integrated into SGLang, Purlin improves offline LLM serving throughput and interactivity by 1.13x on average and up to 1.37x over baselines. For online LLM inference, Purlin improves interactivity by 1.26x on average and up to 2.85x, with the largest gain occurring under overload. For diffusion image generation, Purlin reduces end-to-end latency by up to 1.13x.

References & Citations

...

Bibliographic Explorer

(What is the Explorer?) Connected Papers

(What is Connected Papers?) Litmaps

(What is Litmaps?) scite Smart Citations

(What are Smart Citations?) alphaXiv

(What is alphaXiv?) CatalyzeX Code Finder for Papers

(What is CatalyzeX?) DagsHub

(What is DagsHub?) Gotit.pub

(What is GotitPub?) Hugging Face

(What is Huggingface?) ScienceCast

(What is ScienceCast?) Influence Flower

(What are Influence Flowers?) CORE Recommender

(What is CORE?) arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @purlin 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/purlin-separating-or…] indexed:0 read:2min 2026-09-30 · —