{"slug": "in-the-long-run-nvidia-nvswitch-is-the-infiniband-of-scale-up-ai-networks", "title": "In The Long Run, Nvidia NVSwitch Is The InfiniBand Of Scale Up AI Networks", "summary": "Nvidia's NVSwitch is poised to become the dominant scale-up interconnect for AI networks, mirroring the historical role of InfiniBand, according to The Next Platform. The article notes that Nvidia generated $25.74 billion from InfiniBand in the trailing twelve months ending July, nearly matching the $25.67 billion from Ethernet and NVSwitch combined, but suggests InfiniBand may be at its peak as the industry shifts toward Ultra Ethernet.", "body_md": "# In The Long Run, Nvidia NVSwitch Is The InfiniBand Of Scale Up AI Networks\n\nWay back in the dawn of time, there were two rival future I/O schemes for PCs and servers, and as the Dot Com boom roared along, they buried the hatchet and made a compromise that resulted in the InfiniBand ports and switched fabric as the I/O bus for all of computing.\n\nBut alas, the Dot Com bust came along, and everybody who agreed on this standard kicked it into the ditch and focused on Ethernet for point to point interconnect (like our PCs and servers have linking to the Internet or in-house networks) and for what we now call scale out networking for aggregating compute across many machine. Given the very tough economy, the industry also opted for an updated PCI-X and then PCI-Express bus for peripherals.\n\nInfiniBand as a high performance, low latency interconnect was pulled out of the ashes by Voltaire and Mellanox Technologies, with Mellanox supplying switch ASICs and NICs and Voltaire making switches aimed at HPC ModSim workloads. [Mellanox bought Voltaire for $218 million in November 2010](https://www.theregister.com/on-prem/2010/11/29/mellanox-gobbles-up-voltaire-for-218m/862708) to consolidate its InfiniBand switching position. InfiniCon systems, founded by a bunch of ex-Unisys techies, got into the game and after many permutations became QLogic, the only real rival to Mellanox in the InfiniBand space, [which was acquired by Intel for $125 million in January 2012](https://www.theregister.com/on-prem/2012/01/23/intel-upsets-apple-cart-snaps-up-qlogics-infiniband-biz/544013). (That QLogic team [was spun out of Intel to become Cornelis Networks in September 2020](https://www.nextplatform.com/connect/2020/09/30/its-back-to-the-future-for-omni-path-infiniband/1651716).) TopSpin was an InfiniBand switch player, and was bought by Cisco Systems in 2005. Sun Microsystems, under the guidance of co-founder Andy Bechtosheim, adopted InfiniBand as the clustering technology in its “Constellation” HPC clusters and made its own InfiniBand ASICs and NICs. Oracle continued this for a whole after [it bought Sun Microsystems in April 2009 for $5.6 billion](https://www.theregister.com/off-prem/2009/04/20/king-larry-launches-oracle-sun-combo-at-big-blue-cisco/648246), and Big Larry even bought InfiniBand switch maker Xsigo Systems to bolster its position and then decided that Ethernet was its scale up, scale out, and scale across networking platform.\n\nWith [the Ultra Ethernet Consortium founded in July 2023](https://www.nextplatform.com/connect/2023/07/20/ethernet-consortium-shoots-for-1-million-node-clusters-that-beat-infiniband/1645131), the idea is to marry the low latency and high bandwidth of InfiniBand and some of its quality of service and adaptive routing capabilities to the higher scalability, multitenancy, and quality of service features of Ethernet to make a scale out network that can beat InfiniBand, hands down, and AI clusters that aim to support 1 million endpoints. Ethernet is also making its way into the scale up domain of Nvidia’s NVSwitch interconnects for lashing GPU memories into one coherent, shared memory space for 72 accelerators with a means of expanding the network to support 576 devices if you don’t mind the latency hops.\n\nFor the record, I love InfiniBand, and I love NVSwitch. InfiniBand was the undisputed low latency leader for two and a half decades, and still has an advantage today, which is why I think Nvidia has made $25.74 billion in the trailing twelve fiscal months ending in July. This is almost the same as the $25.67 billion that I think Nvidia sold for Ethernet and NVSwitch interconnects combined. But this is probably peak InfiniBand for large clusters, with so many hyperscalers, cloud builders, and AI model builders moving to standardize on impending Ultra Ethernet technologies and Ethernet ASIC vendors like Broadcom, Cisco Systems, and even Nvidia itself driving down the port hop latency and the end to end latency in their Ethernet fabrics.\n\nMore than a decade ago, I wrote that [InfiniBand was too quick to be killed by Ethernet](https://www.nextplatform.com/cloud/2015/04/01/infiniband-too-quick-for-ethernet-to-kill/1646794), and I expect for InfiniBand to be around for a long time from Nvidia, which is pretty much the only supplier these days aside from some custom HPC interconnects in China that are based on InfiniBand – licensed and sanctioned by the formerly independent Mellanox but perhaps enhanced enough to not be called InfiniBand anymore. But Ethernet can definitely corral InfiniBand into a relative niche, given its scale advantages and wide compatibility with campus, edge, and datacenter networks. And looking ahead, NVswitch for scale up and Spectrum-X Ethernet for scale out are going to dominate Nvidia’s own sales. This is a customer pull, not a vendor push. Enterprises, for sure, are going to want Ethernet, and so do the AI model builders – even Meta Platforms, which has given InfiniBand its time in the limelight on earlier AI clusters.\n\nThe old adage in the networking space is that Ethernet always wins. And it does. And that is because it steals every good idea from any fabric that can be broadly applied across enterprise networks as well as nichey stuff in HPC and now AI. And what is true of the scale out network linking systems together into clusters and people to these systems will be equally true of the scale up network that is being used to lash together GPUs and XPUs in AI clusters. We are already seeing very fast Ethernet being used as a transport for other memory sharing approaches.\n\nThis includes UALink, which fired the first salvo at NVLink and NVSwitch [when it was founded back in May 2024](https://www.nextplatform.com/connect/2025/04/08/ualink-fires-first-gpu-interconnect-salvo-at-nvidia-nvswitch/1638483) by AMD, Broadcom, Cisco Systems, Google, Hewlett Packard Enterprise, Intel, Meta Platforms, and Microsoft; significantly, AMD is using Broadcom Tomahawk 6 Ultra Ethernet switches to run the UALink protocol for the scale up network inside of its “Helios” racks. UALink is just a memory atomic protocol for linking XPUs and GPUs and if you want CPUs and DPUs to share memory, nothing will stop you. So is the ESUN protocol espoused by Meta Platforms and Microsoft, who were joined by AMD, Arista Networks, Arm, Cisco, Hewlett Packard Enterprise, Marvell, Nvidia, OpenAI, and Oracle. Broadcom is part of the ESUN/SUE-T effort (two different layers in an Ethernet scale up stack, with SUE-T handling load balancing and other higher level functions), has walked away from UALink, and as far as we know is not interested in the NVLink/NVSwitch combo. But that may change if its CPU and XPU customers in its chip shepherding business decide they want to endorse NVSwitch as their rackscale fabric and use Nvidia-style racks in their datacenters for both their GPUs and their XPUs.\n\nIt is hard to say how this will all shake out. Today, Nvidia bought $3.5 billion in convertible bonds issued by Taiwanese SoC maker MediaTek, which Nvidia co-founder and chief executive officer Jensen Huang characterized on *Bloomberg* as the biggest and best maker of SoCs in the world. (AMD and Intel would bristle at that, and understandably so.) As part of the deal, MediaTek will gain access to the NVLink Fusion stack, which allows those who design their own chips to put NVLink ports on them so they can be part of the NVSwitch memory fabric. And, just a reminder: Nvidia is not selling NVSwitch ASICs and complexes separately. You either need a “Grace” or “Vera” CPU or an Nvidia “Blackwell” or “Rubin” GPU in the mix, and it assumes you are buying racks that have NVSwitch from Nvidia as the scale up network.\n\nI think there could be a time when Nvidia is happy to sell NVSwitch networking to customers who have NVLink ports on both CPUs and GPU and XPU accelerators. There is no doubt some software support that would have to go along with that. But thus far, Nvidia’s public statements are that this is not allowed.\n\nIn the long run, a couple of different things will happen for the scale up network in AI systems.\n\nFor now, with Nvidia accounting for maybe 95 percent of the GPU revenues and maybe 75 percent of the combined GPU/XPU revenues, it has dominate share of scale up networking with NVLink/NVSwitch. That is probably only going to change by a few points downward through calendar 2027.\n\nBut it bears reminding that the hyperscalers, cloud builders, AI model builders, and large enterprises prefer widely adopted standards that allow for interchangeability across hardware suppliers. The NVLink Fusion is an effort to put a spanner in the UALink and ESUN/SUE-T works, to stall the inevitable for another generation or two.\n\nBut the tech titans can do whatever they want – they have the money to create a new standard for their own devices, and there is every reason to believe that they will eventually come together much as they did with [25 Gb/sec signaling for 100 Gb/sec and faster Ethernet](https://www.theregister.com/on-prem/2015/04/01/ethernet-alliance-plots-16-terabit-per-second-future/680544) back in 2014 once Google, Microsoft, and Arista Networks had had enough of the slow moving IEEE, which wanted to stick with 10 Gb/sec signaling and have ten lanes. This was too much heat, and they told the IEEE to get out of the kitchen.\n\nUALink can meet NVLink/NVSwitch toe to toe, port to port and can – in theory – scale nearly twice as far as NVLink/NVSwitch. The UALink 1.0 spec, [launched in April 2025](https://www.nextplatform.com/connect/2025/04/08/ualink-fires-first-gpu-interconnect-salvo-at-nvidia-nvswitch/1638483), can scale to 1,024 XPUs or GPUs in a single level of UALink switching. We are pretty sure that it takes two tiers of NVSwitch to get to 576 GPUs sharing their memory.\n\nThere is a chance – someone might even say a probability – that the UALink and ESUN/SUE-T camps will bury the hatchet and create a standard that makes all parties happy, and it will probably be called UALink 2.0. (This is how we got the CXL memory coherency standard.) And this hypothetical UALink 2.0 standard is the one that will take on NVSwitch directly. InfiniBand, at 100 nanoseconds to 120 nanoseconds, has one third to one half the latency in port hops, mind you, of Ethernet switches, and the UALink spec says the port to port hop is expected to be around 100 nanoseconds. (The spec doesn’t require a specific latency; vendors will do what they can.) NVSwitch port hops are rumored to be more like Ethernet than InfiniBand, but no one has ever confirmed this. So that will be a big advantage for the UALink camp if those rumors turn out to be true and UALink vendors can get InfiniBand-like latencies. If they can do it atop Ethernet switches, so much the better.\n\nAnd guess what happens if UALink succeeds and becomes the *de facto* scale up network standard? Nvidia will just make the best damned UALink switches and ports that it can, and it will have made maybe $100 billion on NVLink/NVSwitch. Just like it will have made that much selling InfiniBand scale out networks in a world that only wants Ethernet.", "url": "https://wpnews.pro/news/in-the-long-run-nvidia-nvswitch-is-the-infiniband-of-scale-up-ai-networks", "canonical_source": "https://www.nextplatform.com/connect/2026/08/31/in-the-long-run-nvidia-nvswitch-is-the-infiniband-of-scale-up-ai-networks/5293474", "published_at": "2026-08-31 17:12:47+00:00", "updated_at": "2026-08-31 17:23:11.650980+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-chips"], "entities": ["Nvidia", "Mellanox Technologies", "Voltaire", "QLogic", "Intel", "Cornelis Networks", "Cisco Systems", "Ultra Ethernet Consortium"], "alternates": {"html": "https://wpnews.pro/news/in-the-long-run-nvidia-nvswitch-is-the-infiniband-of-scale-up-ai-networks", "markdown": "https://wpnews.pro/news/in-the-long-run-nvidia-nvswitch-is-the-infiniband-of-scale-up-ai-networks.md", "text": "https://wpnews.pro/news/in-the-long-run-nvidia-nvswitch-is-the-infiniband-of-scale-up-ai-networks.txt", "jsonld": "https://wpnews.pro/news/in-the-long-run-nvidia-nvswitch-is-the-infiniband-of-scale-up-ai-networks.jsonld"}}