cd /news/ai-infrastructure/stanford-prof-is-beating-the-drum-fo… · home › topics › ai-infrastructure › article
[ARTICLE · art-143469] src=theregister.com ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

Stanford prof is beating the drum for a new protocol to replace TCP

Stanford University professor emeritus John Ousterhout is promoting Homa, a message-based data transport protocol he says cuts p99 latency for short messages to 92 microseconds, 13 times faster than TCP's 1.2 milliseconds on a 100 Gbps network at 80% utilization, and is twice as fast on the longest messages. Ousterhout, who calls propagating Homa his "life's mission," is drafting an IETF standardization document and working to upstream Homa into the Linux kernel; the protocol was backported to Red Hat Enterprise Linux 8 and 9.5 in March. Network architect Ivan Pepelnjak published a 2023 position paper questioning Ousterhout's TCP performance characterizations and calling Homa a solution looking for a problem.

by read5 min views1 publishedOct 1, 2026
Stanford prof is beating the drum for a new protocol to replace TCP
Image: The Register

TCP, the data transport protocol upon which pretty much the entire web and most of cloud computing is built, is ill-suited for emerging AI workloads, a retired Stanford University professor argues. To solve the problem, he's promoting a new protocol called Homa.

“TCP, for all the amazing things it has done, is not a good match for datacenters,” said John Ousterhout, a professor emeritus of computer science at Stanford University, in a recent talk at the AI Engineer World’s Fair.

Shedding TCP sounds like an immense task, given the global reliance on the protocol.  But adding Homa into a network is fairly simple, Ousterhout told The Register. Compile Homa from its GitHub source, then install the module into the Linux kernels on the clients and servers. No reboot required.

“Homa works side by side with TCP, so you can gradually move applications from TCP to Homa,” he wrote. Running Homa even makes the remaining TCP applications run faster.

Old wine, new skin

Homa is a clean-slate rethink of how networks should manage traffic congestion, Ousterhout said.

Work on the protocol began as a PhD dissertation first published in 2019 by Behnam Montazeri, now a Google staff engineer.  Now that Ousterhout has retired from teaching, he has taken on the task of propagating Homa as his “life’s mission,” he said.

What makes Homa different from TCP is primarily that it is message-based, rather than stream-based. Much like the remote procedure call (RPC), Homa message lengths are explicitly defined.

Unlike TCP, Homa designates the receiver to manage congestion control. The first packet the receiver gets has information on how much data is incoming. It can then explicitly schedule when packets should be sent. In doing so, it prioritizes shorter messages over longer ones using a shortest-remaining-processing-time (SRPT) algorithm.

This approach cuts the latency of shorter messages by an order of magnitude, Ousterhout said. The 99th percentile (p99) of latency for shorter messages is 92 microseconds for Homa, which is 13 times faster than the 1.2 milliseconds p99 for TCP (based on packets swimming through a 100 Gbps network at 80% utilization).

Even on the longest messages, Homa is better by a factor of two, Ousterhout said.

A long list of other tasks where TCP falls short

Currently, Ousterhout is drafting an IETF standardization document of the protocol, as well as working through the process of upstreaming Homa into the Linux kernel. In March, the protocol was backported to Red Hat Enterprise Linux versions 8 and 9.5.

He is also helping large companies investigate Homa’s applicability – he is currently working with one large financial services company on a prototype.

Not that everyone is on board with tossing TCP over a few latency issues.

Prominent network architect Ivan Pepelnjak wrote a scathing position paper in 2023 about Homa, questioning Ousterhout’s performance characterizations of TCP and critiquing Homa as a solution looking for a problem.

To be fair, the AI community is not the only ecosystem frustrated by the dowdy TCP.

In the high-performance database community, DPDK (Data Plane Development Kit) is being used to bypass the TCP stack for faster querying. Storage area networks turned to NVMe-oF (NVMe over Fabrics) to speed access to solid-state drives over network fabrics, using transports including RDMA, Fibre Channel, and TCP. For the Web, Google devised QUIC – which became the basis for HTTP/3 – to bypass TCP’s head-of-line blocking and enable browsers to download more assets simultaneously.

Also, the high-frequency trading and multi-player gaming communities have felt the pinch of TCP sluggishness.

Specialized RDMA fabrics and Amazon Web Services’ Scalable Reliable Datagram have also tackled the issue of TCP latency. Top-of-rack switches have also gotten smarter at managing congestion, setting queue thresholds and marking packets with early congestion notifications.

Control lag is real

Vint Cerf and his fellow beardies created TCP to bring order to the unruly hordes of message packets across networks, giving them proper flow control, guaranteed delivery, connection handshakes and congestion control. Congestion control means avoiding overload in the network path, including switches and routers, while separate flow-control mechanisms keep senders from overwhelming receivers.

TCP’s data model is built on byte streams, a continuous flow of data packets with no differentiation. Messages are serialized as a single byte stream with no priority. For the receiver, longer sets of byte streams are indistinguishable from shorter ones.

A server deluged with traffic can send alerts to senders to slow input, but it has limited visibility over how much traffic is still coming in. The sender itself, which regulates its output depending on the timeliness of acknowledgments sent back from the receiver, has to guess how much to slow its roll.

For regular internet traffic or large-scale transfers within a datacenter, modest latency increases may be tolerable. But for latency-sensitive AI workloads, even delays measured in milliseconds can sting.

Think of the GPUs

The frontier labs driving the development of large language models (LLMs) have always required top-notch network performance for chores such as weight gradients, model weights, KV cache entries, and checkpoints.

However, large data transfers must increasingly share bandwidth with short bursts of traffic from agents and control tasks such as metadata coordination and cache lookups.

“For these workloads, what really matters is latency,” Ousterhout said. Even a millisecond of latency causes an expensive GPU to go idle.

“Legacy protocols are poorly suited for this environment,” Ousterhout said.

So has Homa finally found its problem to solve? Or was it just ahead of its time all along? For now anyway, TCP remains the champ.®

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @john ousterhout 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/stanford-prof-is-bea…] indexed:0 read:5min 2026-10-01 · —