{"slug": "stanford-prof-is-beating-the-drum-for-a-new-protocol-to-replace-tcp", "title": "Stanford prof is beating the drum for a new protocol to replace TCP", "summary": "Stanford University professor emeritus John Ousterhout is promoting Homa, a message-based data transport protocol he says cuts p99 latency for short messages to 92 microseconds, 13 times faster than TCP's 1.2 milliseconds on a 100 Gbps network at 80% utilization, and is twice as fast on the longest messages. Ousterhout, who calls propagating Homa his \"life's mission,\" is drafting an IETF standardization document and working to upstream Homa into the Linux kernel; the protocol was backported to Red Hat Enterprise Linux 8 and 9.5 in March. Network architect Ivan Pepelnjak published a 2023 position paper questioning Ousterhout's TCP performance characterizations and calling Homa a solution looking for a problem.", "body_md": "TCP, the data transport protocol upon which pretty much the entire web and most of cloud computing is built, is ill-suited for emerging AI workloads, a retired Stanford University professor argues. To solve the problem, he's promoting a new protocol called Homa.  \n\n“TCP, for all the amazing things it has done, is not a good match for datacenters,” said [John Ousterhout](https://engineering.stanford.edu/people/john-ousterhout), a professor emeritus of computer science at Stanford University, in [a recent talk](https://www.youtube.com/watch?v=eZ8WWZzoaR0) at the AI Engineer World’s Fair.\n\nShedding TCP sounds like an immense task, given the global reliance on the protocol.  But adding Homa into a network is fairly simple, Ousterhout told The Register. Compile Homa from its [GitHub source](https://github.com/PlatformLab/HomaModule), then install the module into the Linux kernels on the clients and servers. No reboot required. \n\n“Homa works side by side with TCP, so you can gradually move applications from TCP to Homa,” he wrote. Running Homa even makes the remaining TCP applications run faster.\n\n### Old wine, new skin\n\nHoma is a clean-slate rethink of how networks should manage traffic congestion, Ousterhout said.\n\nWork on the protocol [began](https://www.theregister.com/on-prem/2022/07/27/there-is-a-path-to-replace-tcp-in-the-datacenter/828009) as a [PhD dissertation](https://arxiv.org/abs/2210.00714) first published in 2019 by [Behnam Montazeri](https://www.linkedin.com/in/behnam-montazeri/), now a Google staff engineer.  Now that Ousterhout has retired from teaching, he has taken on the task of propagating Homa as his “life’s mission,” he said. \n\nWhat makes Homa different from TCP is primarily that it is message-based, rather than stream-based. Much like the remote procedure call (RPC), Homa message lengths are explicitly defined. \n\nUnlike TCP, Homa designates the receiver to manage congestion control. The first packet the receiver gets has information on how much data is incoming. It can then explicitly schedule when packets should be sent. In doing so, it prioritizes shorter messages over longer ones using a shortest-remaining-processing-time (SRPT) algorithm.\n\nThis approach cuts the latency of shorter messages by an order of magnitude, Ousterhout said. The 99th percentile (p99) of latency for shorter messages is 92 microseconds for Homa, which is 13 times faster than the 1.2 milliseconds p99 for TCP (based on packets swimming through a 100 Gbps network at 80% utilization). \n\nEven on the longest messages, Homa is better by a factor of two, Ousterhout said.\n\n### A long list of other tasks where TCP falls short\n\nCurrently, Ousterhout is drafting an IETF standardization document of the protocol, as well as working through the process of [upstreaming](https://www.usenix.org/system/files/atc21-ousterhout.pdf) Homa into the Linux kernel. In March, the protocol was backported to Red Hat Enterprise Linux versions 8 and 9.5.\n\nHe is also helping large companies investigate Homa’s applicability – he is currently working with one large financial services company on a prototype. \n\nNot that everyone is on board with tossing TCP over a few latency issues. \n\nProminent network architect Ivan Pepelnjak wrote a [scathing position paper](https://blog.ipspace.net/2023/01/data-center-tcp-replacement/) in 2023 about Homa, questioning Ousterhout’s performance characterizations of TCP and critiquing Homa as a solution looking for a problem.\n\nTo be fair, the AI community is not the only ecosystem frustrated by the dowdy TCP. \n\nIn the high-performance database community, [DPDK](https://www.dpdk.org/) (Data Plane Development Kit) is being used to bypass the TCP stack for faster querying. Storage area networks turned to [NVMe-oF](https://www.everpuredata.com/knowledge/what-is-nvme-over-fabrics-nvme-of.html) (NVMe over Fabrics) to speed access to solid-state drives over network fabrics, using transports including RDMA, Fibre Channel, and TCP. For the Web, Google devised [QUIC](https://www.theregister.com/networks/2026/07/08/media-over-quic-can-scale-real-time-streaming-and-carry-the-worlds-vids/5268101) – which became the basis for HTTP/3 – to bypass TCP’s head-of-line blocking and enable browsers to download more assets simultaneously. \n\nAlso, the [high-frequency trading](https://github.com/luishsr/hft-kernel-bypass) and [multi-player gaming](https://nva.sikt.no/registration/0198eb029b61-4a319b6a-cd4a-4543-af2a-316087facb73) communities have felt the pinch of TCP sluggishness. \n\n[Specialized RDMA fabrics](https://users.cs.duke.edu/~alvy/papers/CloudMicro_RDMA.pdf) and [Amazon Web Services’ Scalable Reliable Datagram](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ena-express.html) have also tackled the issue of TCP latency. Top-of-rack switches have also gotten smarter at managing congestion, setting queue thresholds and marking packets with early congestion notifications.\n\n### Control lag is real\n\nVint Cerf and his fellow beardies created [TCP](https://datatracker.ietf.org/doc/html/rfc9293) to bring order to the unruly hordes of message packets across networks, giving them proper flow control, guaranteed delivery, connection handshakes and congestion control. Congestion control means avoiding overload in the network path, including switches and routers, while separate flow-control mechanisms keep senders from overwhelming receivers.\n\n[TCP’s data model](https://datatracker.ietf.org/doc/html/rfc9293) is built on byte streams, a continuous flow of data packets with no differentiation. Messages are serialized as a single byte stream with no priority. For the receiver, longer sets of byte streams are indistinguishable from shorter ones. \n\nA server deluged with traffic can send alerts to senders to slow input, but it has limited visibility over how much traffic is still coming in. The sender itself, which regulates its output depending on the timeliness of acknowledgments sent back from the receiver, has to guess how much to slow its roll. \n\nFor regular internet traffic or large-scale transfers within a datacenter, modest latency increases may be tolerable. But for latency-sensitive AI workloads, even delays measured in milliseconds can sting.\n\n### Think of the GPUs\n\nThe frontier labs driving the development of large language models (LLMs) have always required top-notch network performance for chores such as weight gradients, model weights, KV cache entries, and checkpoints.\n\nHowever, large data transfers must increasingly share bandwidth with short bursts of traffic from agents and control tasks such as metadata coordination and cache lookups.\n\n“For these workloads, what really matters is latency,” Ousterhout said. Even a millisecond of latency causes an expensive GPU to go idle. \n\n“Legacy protocols are poorly suited for this environment,” Ousterhout said.  \n\nSo has Homa finally found its problem to solve? Or was it just ahead of its time all along? For now anyway, TCP remains the champ.®", "url": "https://wpnews.pro/news/stanford-prof-is-beating-the-drum-for-a-new-protocol-to-replace-tcp", "canonical_source": "https://www.theregister.com/networks/2026/10/01/stanford-prof-is-beating-the-drum-for-a-new-protocol-to-replace-tcp/5300629", "published_at": "2026-10-01 20:00:45+00:00", "updated_at": "2026-10-01 21:47:11.881021+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-chips"], "entities": ["John Ousterhout", "Stanford University", "Homa", "TCP", "Behnam Montazeri", "Google", "Red Hat Enterprise Linux", "Ivan Pepelnjak"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/stanford-prof-is-beating-the-drum-for-a-new-protocol-to-replace-tcp", "markdown": "https://wpnews.pro/news/stanford-prof-is-beating-the-drum-for-a-new-protocol-to-replace-tcp.md", "text": "https://wpnews.pro/news/stanford-prof-is-beating-the-drum-for-a-new-protocol-to-replace-tcp.txt", "jsonld": "https://wpnews.pro/news/stanford-prof-is-beating-the-drum-for-a-new-protocol-to-replace-tcp.jsonld"}}