cd /news/ai-infrastructure/unified-ai-networking-architecture-b… · home › topics › ai-infrastructure › article
[ARTICLE · art-148412] src=dev.to ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Unified AI Networking Architecture Bridges Data Center Gaps

Networking startup Upscale introduced Token Fabric, a unified AI networking architecture that combines its proprietary SkyFabriX scale-up silicon with scale-out systems built on Nvidia Spectrum-X Ethernet, managed by a common software layer. The Santa Clara company, which has raised roughly $500 million including a June round with Nvidia as a strategic investor, plans general availability in early 2027 with early-access and joint-validation programs already running. CEO Barun Kar said compute capabilities are ramping up rapidly while networking infrastructure has lagged, creating inefficiencies in how AI models are trained and deployed.

by read6 min views3 publishedOct 9, 2026

AI clusters depend on two types of networking to function effectively for modern workloads. Scale-up networking connects accelerators inside a rack, while scale-out networking connects racks across the data center. A single AI job spans both, which is a challenge for many organizations that treat them as separate.

Networking startup Upscale introduced Token Fabric to bring these two critical components together. The architecture combines the proprietary SkyFabriX scale-up silicon with scale-out systems built on Nvidia Spectrum-X Ethernet technology. A common software layer manages operations across the entire environment. General availability for the platform is planned for early 2027, though early-access and joint-validation programs are currently running.

The announcement marks a significant step for the Santa Clara-based company. Upscale emerged from stealth in September 2025 with a 100 million dollar seed round. In June, it raised an additional 190 million dollars while outlining plans for SkyHammer, its custom scale-up switch chip. SkyFabriX is built on that specific SkyHammer architecture.

Nvidia joined the June funding round as a strategic investor and supplier of the Spectrum-X silicon. Upscale has now raised approximately 500 million dollars in total capital. The firm currently supports a workforce of more than 300 employees dedicated to solving bottlenecks in high-speed computing.

CEO Barun Kar explained that while compute capabilities are ramping up at a rapid pace, networking infrastructure has lagged behind. This discrepancy creates inefficiencies in how AI models are trained and deployed. Token Fabric aims to close this gap by ensuring the network can keep pace with modern GPU demands.

SkyFabriX silicon serves as the foundation for the scale-up switch hardware. It supports Ethernet for Scale-Up Networking and standard IP protocols. The design is intended to accommodate evolving standards like Ultra Accelerator Link. The current switch capacity reaches 115.2 terabits per second, with plans to reach multiple petabits in the future.

The system also includes specialized scale-up switch trays. These components come in both custom and standard rack form factors to suit various data center needs. They pool GPUs and other accelerators into large scale-up domains to maximize resource utilization.

Scale-out systems within the architecture utilize Nvidia Spectrum-X silicon. These systems operate at speeds ranging from 400G and 800G up to 1.6T. Through the use of SkyOS, they integrate with various accelerator clusters to ensure high-speed communication between different racks.

SkyOS serves as the network operating system for both types of fabrics. It is based on the open-source SONiC platform but is optimized specifically for AI-related protocols. This ensures that the software can handle the unique traffic patterns associated with large-scale machine learning.

SkyCMD provides the orchestration and observability layer for the entire stack. It offers a single management plane across scale-up and scale-out environments. This unified view simplifies the task of monitoring network health and performance across complex hardware configurations.

Customers have several ways to adopt this technology based on their specific needs. Some may choose only the silicon or software, while others might opt for the full stack. This flexibility targets different market segments, including enterprise teams and smaller cloud providers that lack massive internal engineering resources.

Token Fabric is built from the protocol layer up to ensure compatibility and performance. Upscale is extending existing protocols rather than replacing them entirely. This approach allows for easier integration into existing data center environments while providing the specialized features needed for AI.

Scale-up traffic within the system uses Ethernet for Scale-Up Networking alongside standard IP. Scale-out traffic relies on standard Ethernet with RoCE. This configuration carries remote direct memory access over Ethernet to reduce latency and overhead.

The company is also an active participant in several open-source projects. These include contributions to SONiC and extensions of the Switch Abstraction Interface. They also work with the Ultra Ethernet Consortium and UALink groups to help shape the future of industry standards.

Software acts as the primary link between the two distinct fabrics. SkyOS abstracts the underlying hardware to provide a consistent control plane across the cluster. This prevents the operational silos that often occur when scale-up and scale-out networks are managed by different tools.

SkyCMD exposes this control plane through a single interface for the user. Multiple network elements sit behind this layer, allowing any compute platform to control the network. This design is particularly useful for heterogeneous clusters containing different types of hardware.

Kar noted that an operating system must be lean and fast to maintain flexibility and security. The orchestration layer on top is what truly enables heterogeneous compute environments to function as a single unit. Without this integration, managing large clusters becomes an overwhelming task for IT departments.

The focus on abstraction allows developers to interact with the network without needing deep knowledge of the underlying hardware. This shift is essential as organizations move away from custom-built silos toward more standardized AI infrastructure. By using open standards like SONiC, Upscale ensures that its customers are not locked into a completely proprietary ecosystem.

Traditional network teams have focused on metrics like packets, throughput, and jitter. However, the rise of large language models has introduced a new unit of measurement. The question now is whether tokens should be the primary metric for network performance optimization.

Kar argued that tokens are indeed the correct unit for the modern era. Users and businesses currently pay for AI services based on token counts. Therefore, metrics like time to first token and tokens per watt are becoming the standard for evaluating data center efficiency.

Optimizing these metrics requires a deep focus on networking, especially as clusters grow to include hundreds of thousands of accelerators. There is currently a gap in many networks that makes it difficult to reach peak token efficiency. This gap is often caused by latency issues and a lack of scale-up features in the silicon.

Predictive analytics and telemetry are also vital on the software side. These tools keep machines running and help prevent downtime that could stall expensive training jobs. By focusing on the token as the final output, Upscale aligns its networking goals with the business goals of its clients.

To handle the complexity of modern clusters, Upscale has integrated agent-based operations directly into the platform. These agents use AI to monitor the network in real-time. They can detect potential failures before they occur and take corrective action.

An agent might identify a cable or optic that is beginning to fail. It can also detect congestion within the fabric and find alternative paths for data. This proactive approach reduces the manual workload for human operators and increases the overall reliability of the cluster.

These agents are constantly looking for ways to improve efficiency within the fabric. They can adjust configurations on the fly to respond to changing workloads. This level of automation is necessary when dealing with the massive scale of modern AI deployments.

The integration of AI into the management of AI networks creates a self-optimizing system. As the hardware becomes more powerful, the software must become more intelligent to manage it. This circular relationship is at the heart of the Token Fabric philosophy. Using agents ensures that the network is always tuned for the highest possible token throughput.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @upscale 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/unified-ai-networkin…] indexed:0 read:6min 2026-10-09 · —