{"slug": "unified-ai-networking-architecture-bridges-data-center-gaps", "title": "Unified AI Networking Architecture Bridges Data Center Gaps", "summary": "Networking startup Upscale introduced Token Fabric, a unified AI networking architecture that combines its proprietary SkyFabriX scale-up silicon with scale-out systems built on Nvidia Spectrum-X Ethernet, managed by a common software layer. The Santa Clara company, which has raised roughly $500 million including a June round with Nvidia as a strategic investor, plans general availability in early 2027 with early-access and joint-validation programs already running. CEO Barun Kar said compute capabilities are ramping up rapidly while networking infrastructure has lagged, creating inefficiencies in how AI models are trained and deployed.", "body_md": "AI clusters depend on two types of networking to function effectively for modern workloads. Scale-up networking connects accelerators inside a rack, while scale-out networking connects racks across the data center. A single AI job spans both, which is a challenge for many organizations that treat them as separate.\n\nNetworking startup Upscale introduced Token Fabric to bring these two critical components together. The architecture combines the proprietary SkyFabriX scale-up silicon with scale-out systems built on Nvidia Spectrum-X Ethernet technology. A common software layer manages operations across the entire environment. General availability for the platform is planned for early 2027, though early-access and joint-validation programs are currently running.\n\nThe announcement marks a significant step for the Santa Clara-based company. Upscale emerged from stealth in September 2025 with a 100 million dollar seed round. In June, it raised an additional 190 million dollars while outlining plans for SkyHammer, its custom scale-up switch chip. SkyFabriX is built on that specific SkyHammer architecture.\n\nNvidia joined the June funding round as a strategic investor and supplier of the Spectrum-X silicon. Upscale has now raised approximately 500 million dollars in total capital. The firm currently supports a workforce of more than 300 employees dedicated to solving bottlenecks in high-speed computing.\n\nCEO Barun Kar explained that while compute capabilities are ramping up at a rapid pace, networking infrastructure has lagged behind. This discrepancy creates inefficiencies in how AI models are trained and deployed. Token Fabric aims to close this gap by ensuring the network can keep pace with modern GPU demands.\n\nSkyFabriX silicon serves as the foundation for the scale-up switch hardware. It supports Ethernet for Scale-Up Networking and standard IP protocols. The design is intended to accommodate evolving standards like Ultra Accelerator Link. The current switch capacity reaches 115.2 terabits per second, with plans to reach multiple petabits in the future.\n\nThe system also includes specialized scale-up switch trays. These components come in both custom and standard rack form factors to suit various data center needs. They pool GPUs and other accelerators into large scale-up domains to maximize resource utilization.\n\nScale-out systems within the architecture utilize Nvidia Spectrum-X silicon. These systems operate at speeds ranging from 400G and 800G up to 1.6T. Through the use of SkyOS, they integrate with various accelerator clusters to ensure high-speed communication between different racks.\n\nSkyOS serves as the network operating system for both types of fabrics. It is based on the open-source SONiC platform but is optimized specifically for AI-related protocols. This ensures that the software can handle the unique traffic patterns associated with large-scale machine learning.\n\nSkyCMD provides the orchestration and observability layer for the entire stack. It offers a single management plane across scale-up and scale-out environments. This unified view simplifies the task of monitoring network health and performance across complex hardware configurations.\n\nCustomers have several ways to adopt this technology based on their specific needs. Some may choose only the silicon or software, while others might opt for the full stack. This flexibility targets different market segments, including enterprise teams and smaller cloud providers that lack massive internal engineering resources.\n\nToken Fabric is built from the protocol layer up to ensure compatibility and performance. Upscale is extending existing protocols rather than replacing them entirely. This approach allows for easier integration into existing data center environments while providing the specialized features needed for AI.\n\nScale-up traffic within the system uses Ethernet for Scale-Up Networking alongside standard IP. Scale-out traffic relies on standard Ethernet with RoCE. This configuration carries remote direct memory access over Ethernet to reduce latency and overhead.\n\nThe company is also an active participant in several open-source projects. These include contributions to SONiC and extensions of the Switch Abstraction Interface. They also work with the Ultra Ethernet Consortium and UALink groups to help shape the future of industry standards.\n\nSoftware acts as the primary link between the two distinct fabrics. SkyOS abstracts the underlying hardware to provide a consistent control plane across the cluster. This prevents the operational silos that often occur when scale-up and scale-out networks are managed by different tools.\n\nSkyCMD exposes this control plane through a single interface for the user. Multiple network elements sit behind this layer, allowing any compute platform to control the network. This design is particularly useful for heterogeneous clusters containing different types of hardware.\n\nKar noted that an operating system must be lean and fast to maintain flexibility and security. The orchestration layer on top is what truly enables heterogeneous compute environments to function as a single unit. Without this integration, managing large clusters becomes an overwhelming task for IT departments.\n\nThe focus on abstraction allows developers to interact with the network without needing deep knowledge of the underlying hardware. This shift is essential as organizations move away from custom-built silos toward more standardized AI infrastructure. By using open standards like SONiC, Upscale ensures that its customers are not locked into a completely proprietary ecosystem.\n\nTraditional network teams have focused on metrics like packets, throughput, and jitter. However, the rise of large language models has introduced a new unit of measurement. The question now is whether tokens should be the primary metric for network performance optimization.\n\nKar argued that tokens are indeed the correct unit for the modern era. Users and businesses currently pay for AI services based on token counts. Therefore, metrics like time to first token and tokens per watt are becoming the standard for evaluating data center efficiency.\n\nOptimizing these metrics requires a deep focus on networking, especially as clusters grow to include hundreds of thousands of accelerators. There is currently a gap in many networks that makes it difficult to reach peak token efficiency. This gap is often caused by latency issues and a lack of scale-up features in the silicon.\n\nPredictive analytics and telemetry are also vital on the software side. These tools keep machines running and help prevent downtime that could stall expensive training jobs. By focusing on the token as the final output, Upscale aligns its networking goals with the business goals of its clients.\n\nTo handle the complexity of modern clusters, Upscale has integrated agent-based operations directly into the platform. These agents use AI to monitor the network in real-time. They can detect potential failures before they occur and take corrective action.\n\nAn agent might identify a cable or optic that is beginning to fail. It can also detect congestion within the fabric and find alternative paths for data. This proactive approach reduces the manual workload for human operators and increases the overall reliability of the cluster.\n\nThese agents are constantly looking for ways to improve efficiency within the fabric. They can adjust configurations on the fly to respond to changing workloads. This level of automation is necessary when dealing with the massive scale of modern AI deployments.\n\nThe integration of AI into the management of AI networks creates a self-optimizing system. As the hardware becomes more powerful, the software must become more intelligent to manage it. This circular relationship is at the heart of the Token Fabric philosophy. Using agents ensures that the network is always tuned for the highest possible token throughput.", "url": "https://wpnews.pro/news/unified-ai-networking-architecture-bridges-data-center-gaps", "canonical_source": "https://dev.to/vpodk/unified-ai-networking-architecture-bridges-data-center-gaps-1kgd", "published_at": "2026-10-09 17:46:23+00:00", "updated_at": "2026-10-09 17:51:28.942929+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-chips", "ai-agents"], "entities": ["Upscale", "Token Fabric", "SkyFabriX", "Nvidia", "Spectrum-X", "SkyHammer", "SkyOS", "Barun Kar"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/unified-ai-networking-architecture-bridges-data-center-gaps", "markdown": "https://wpnews.pro/news/unified-ai-networking-architecture-bridges-data-center-gaps.md", "text": "https://wpnews.pro/news/unified-ai-networking-architecture-bridges-data-center-gaps.txt", "jsonld": "https://wpnews.pro/news/unified-ai-networking-architecture-bridges-data-center-gaps.jsonld"}}