cd /news/ai-infrastructure/modal-launches-serverless-gpu-cluste… · home › topics › ai-infrastructure › article
[ARTICLE · art-143479] src=runtimewire.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Modal launches serverless GPU clusters with gang scheduling and RDMA

Modal made its multi-node GPU clusters generally available on October 1st, letting developers request coordinated GPU capacity with the Python decorator @modal.clustered(size=4, rdma=True) instead of assembling a cluster themselves. Modal says its new gang scheduler treats the whole cluster as a single scheduling unit, groups nodes by availability zone and network, and configures RDMA for PyTorch and NCCL, with clusters billed by the second from the same pooled capacity as Modal's other workloads. The launch extends the infrastructure bet of co-founder and CEO Erik Bernhardsson, whose company reported annualized revenue above $300 million after raising $355 million at a $4.65 billion post-money valuation in a May round led by General Catalyst and Redpoint Ventures.

read5 min views4 publishedOct 1, 2026
Modal launches serverless GPU clusters with gang scheduling and RDMA
Image: Runtimewire (auto-discovered)

Modal is extending Erik Bernhardsson's pooled compute platform to distributed training and inference, with per-second billing and a single Python decorator.

        By [RuntimeWire Staff](https://runtimewire.com/author/runtimewire-staff)
        · Published 

Primary source: [Modal Newsroom](https://modal.com/blog/modal-clusters-generally-available)

Why it matters #

Modal is moving its serverless model into jobs that traditionally require coordinated access to expensive, tightly networked hardware. The bet is that its scheduler, pooled capacity and isolation layer can make distributed AI workloads feel like ordinary cloud functions without removing the operational demands of keeping a cluster reliable.

Modal made its multi-node GPU clusters generally available on October 1st, giving developers a way to request coordinated GPU capacity with a Python decorator instead of arranging a cluster themselves. Co-founder and CEO Erik Bernhardsson (@bernhardsson) is extending the infrastructure bet he began with Modal: developers should be able to run compute-heavy code in the cloud without first building and maintaining the machinery around it.

Bernhardsson came to that bet after building data and recommendation systems at Spotify, where he helped develop the music recommendations behind products including Discover Weekly, and later leading the engineering organization at Better.com. In a 2022 account of starting Modal, he said he wanted to build software that could take code on a developer's computer and launch it in the cloud within a second. Clusters apply that same premise to workloads that need several tightly connected machines at once. Modal's other co-founder, Akshat Bubna, built the company with him.

The scheduler is the product

A developer requests a cluster through @modal.clustered(size=4, rdma=True). In Modal's example, the function requests four nodes, each equipped with eight B300 GPUs. The API gives each container its rank and the cluster's private IP addresses, the basic information needed to coordinate a distributed job. Modal says the feature works with its Volumes, Cloud Bucket Mounts and Queues, and configures RDMA for PyTorch and NCCL.

The consequential engineering work sits below the decorator. A conventional scheduler can assign jobs one machine at a time; a distributed training or inference task needs all the requested nodes placed together. Modal says its new gang scheduler treats the whole cluster as a single scheduling unit: it evaluates pending jobs across the fleet, groups nodes by availability zone and network, adds capacity when needed, then tells the available nodes to start. The scheduler draws from the same capacity pool as Modal's other workloads.

Modal sells that pooled capacity as the commercial proposition: Clusters can start within seconds and are billed by the second, rather than requiring customers to reserve hardware by the hour or keep a cluster running between jobs. Modal has spent years building the custom runtime, scheduler and storage layers that make this sort of abstraction possible. In a May financing announcement, Modal said it raised $355 million at a $4.65 billion post-money valuation in a round led by General Catalyst and Redpoint Ventures, with Menlo Ventures, Bain Capital Ventures and Accel also joining. Modal reported annualized revenue above $300 million. The round gives Modal substantial backing as it invests in infrastructure that must work across clouds and large GPU jobs.

RDMA inside a sandbox

Modal advertises InfiniBand verbs networking of up to 6.4 Tbps per node. Remote Direct Memory Access, or RDMA, lets machines exchange data without routing it through the usual operating-system software path. That matters when GPUs repeatedly synchronize model weights during training or transfer large caches during inference. Modal's announcement uses a GLM 4.7 example: it says moving roughly 717 GB of BF16 weights takes nearly two minutes over 50 Gbps TCP and less than two seconds over RDMA. Those timings are Modal's illustration of the technology, not an independently reported benchmark of Clusters.

Making that speed available inside Modal's secure, multi-tenant containers required a less obvious change. Modal says its gVisor runtime did not support the RDMA operations it needed, so it built a proxy into gVisor and upstreamed the changes. The merged gVisor change documents a proxy for RDMA verbs and GPU memory access. That work addresses a real tension in shared infrastructure: direct access to networking hardware can improve performance, while a cloud provider still needs isolation between customers.

Modal's stated customer examples show the product serving more than model training. Modal says Decagon fine-tuned open models with up to one trillion parameters using the Miles framework, and contributed LoRA support to it. Modal also says robotics company 1X uses B300 clusters for pretraining and runs single-node evaluation jobs through the same scheduler, while Runway uses Clusters for multi-node video inference. These are customer examples published by Modal; the announcement does not provide independent workload measurements for them.

Clusters extend Bernhardsson's original developer-experience thesis to workloads that need several machines at once. RuntimeWire reported in September that Modal had opened a London office to expand its AI cloud sales across Europe. Clusters add a technical product to that expansion: customers can pursue larger training and inference jobs through the same code-first platform, while Modal takes responsibility for the placement, networking and container setup.

Modal's operational pitch depends on whether its pooled fleet can consistently assemble the right machines, keep them connected and recover when a node or network path fails. Modal says it has been testing Clusters for 1.5 years, and that GPU health checks now extend to RDMA health. The product is available to all Modal workspaces, but cluster size depends on the GPU limits of each plan. Developers can request a cluster with a line of Python; Modal is taking on the reliability work that makes that line worth trusting.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @modal 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/modal-launches-serve…] indexed:0 read:5min 2026-10-01 · —