# IBM, Together AI Ink $240 Million Deal for Nvidia-Powered AI Inference Cluster

> Source: <https://insideai.news/news/ai-in-business/ibm-together-ai-ink-240-million-deal-for-nvidia-powered-ai-inference-cluster/7602/>
> Published: 2026-08-11 13:06:18+00:00

**August 11, 2026**, (Inside AI) — IBM and Together AI have signed a **$240 million** multi-year deal to build a massive AI inference cluster on IBM Cloud, powered by Nvidia's latest systems. The cluster will focus on running open-source AI models, a segment gaining momentum as enterprises seek cost control and security alternatives to proprietary models from Anthropic, OpenAI, and Meta.

The agreement, announced Tuesday, marks one of the largest dedicated inference infrastructure commitments to date. It directly addresses the surging demand for inference computing, the process where trained AI models generate responses. This has become a primary driver of infrastructure spending, with cloud providers and chipmakers investing billions.

The cluster will utilize Nvidia's **HGX B300** systems, which incorporate the company's newer **Blackwell** processors and **Spectrum-X** Ethernet networking. Nvidia has positioned Blackwell as optimized for AI inference workloads, promising significant performance gains over previous generations.

Together AI, a San Francisco-based startup last valued at **$8.3 billion** in July, operates a platform that allows companies to train and run AI workloads on open models such as **DeepSeek**, **MiniMax**, and **Kimi**. The company claims its approach delivers lower costs than closed systems, a key selling point as enterprises scrutinize AI expenditures.

"The cluster on IBM Cloud will use Nvidia's HGX B300 systems, which link the chipmaker's newer Blackwell processors, and its Spectrum-X Ethernet networking gear," the companies stated. This hardware configuration is designed to handle the intense throughput requirements of large-scale inference, where latency and throughput directly impact user experience.

Open-source models have seen a dramatic rise in enterprise adoption over the past year. Concerns over vendor lock-in, unpredictable pricing, and data privacy incidents have pushed many organizations toward self-hosted or community-driven alternatives. The IBM-Together AI cluster aims to provide the dedicated infrastructure needed to serve these models at scale without relying on public API endpoints.

This deal also reflects IBM's strategic push to reclaim relevance in the cloud AI market, where it has lagged behind hyperscalers like AWS, Microsoft Azure, and Google Cloud. By partnering with a high-growth startup and leveraging Nvidia's latest silicon, IBM is betting on inference as a differentiated workload rather than competing directly on training compute.

For Nvidia, the agreement reinforces the growing importance of inference as a revenue driver. While the company's H100 GPUs initially dominated training workloads, the shift toward deployment has made inference the larger long-term opportunity. Blackwell's architecture, with its focus on transformer engine optimizations and lower precision computing, is tailored for this transition.

Industry analysts note that dedicated inference clusters could reshape the economics of AI deployment. Rather than sharing general-purpose cloud instances, companies can access hardware specifically tuned for model serving, potentially reducing per-token costs by **30-50%** compared to traditional GPU instances.

The deal also highlights the growing infrastructure demands of open-source models. DeepSeek's latest models, for example, have demonstrated competitive performance against proprietary counterparts but require substantial memory and compute for real-time inference. The IBM Cloud cluster is expected to provide the necessary scale to serve these models to thousands of concurrent users.

Financial terms beyond the **$240 million** total were not disclosed. The multi-year structure suggests a long-term capacity reservation, a model increasingly common as AI compute supply remains constrained. This approach guarantees Together AI access to Nvidia's latest hardware, which has faced allocation challenges due to overwhelming demand.

Together AI's platform has attracted attention for its ability to run diverse model architectures efficiently. The company supports a range of open-weight models, enabling enterprises to switch between or combine models based on task requirements. This flexibility is a key differentiator from single-model API services.

The announcement comes amid heightened scrutiny of AI cybersecurity. Recent incidents involving model vulnerabilities and data leakage have accelerated interest in self-hosted solutions where organizations retain full control over model access and data flows. IBM's enterprise security credentials may provide additional assurance for regulated industries.

Looking ahead, the cluster's deployment timeline and geographic location remain undisclosed. However, the partnership signals a broader industry trend: the convergence of open-source AI, specialized hardware, and cloud infrastructure to create new deployment paradigms that challenge the dominance of proprietary model APIs.
