cd /news/ai-infrastructure/ibm-and-together-ai-put-240m-into-a-… · home topics ai-infrastructure article
[ARTICLE · art-91916] src=storagereview.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

IBM and Together AI Put $240M Into a Dedicated HGX B300 Inference Cluster on IBM Cloud

IBM and Together AI announced a multi-year, $240 million agreement to deploy a dedicated NVIDIA HGX B300-based AI inference cluster on IBM Cloud, marking IBM Cloud's first dedicated large-scale inference cluster built around NVIDIA HGX B300 systems. Together AI, whose inference platform serves 400 trillion tokens monthly, will use the environment to deliver inference services for open-source models, with the deployment designed to deliver 30x more AI factory output compared to prior generations. The collaboration aims to provide scalable, economical, enterprise-grade AI infrastructure for production inference workloads.

read2 min views1 publishedAug 11, 2026
IBM and Together AI Put $240M Into a Dedicated HGX B300 Inference Cluster on IBM Cloud
Image: Storagereview (auto-discovered)

IBM has announced a multi-year, $240 million agreement with Together AI to deploy a dedicated NVIDIA HGX B300-based AI inference cluster on IBM Cloud. Together AI plans to use the environment to deliver inference services for open-source models, expanding its AI Native Cloud platform for enterprise customers.

The deployment is positioned as IBM Cloud’s first dedicated large-scale inference cluster built around NVIDIA HGX B300 systems, following the Grace Blackwell capacity IBM added through CoreWeave last year. It will also use NVIDIA Spectrum-X Ethernet networking, creating an AI factory architecture intended to support high-throughput, low-latency inference workloads. IBM says the deployment is built to deliver 30x more AI factory output compared to prior generations. However, workload-level performance will depend on model architecture, precision, batch size, and serving configuration.

Together AI offers infrastructure and software services spanning inference, model training, fine-tuning, and agentic AI workflows. The company reports that its inference platform now serves 400 trillion tokens monthly. The new IBM Cloud deployment is intended to add GPU capacity for production inference while improving performance and token economics for organizations deploying open-weight and open-source models at scale.

“Together AI is proud to lead the way in bringing production inference to market with NVIDIA’s latest AI infrastructure on IBM Cloud,” said Vipul Ved Prakash, CEO at Together AI. “Working alongside IBM with NVIDIA accelerates our mission to make advanced AI broadly accessible through open source and to empower builders with the infrastructure and platform capabilities they need to build the future.”

The collaboration brings together IBM Cloud’s enterprise infrastructure, NVIDIA’s HGX B300 compute systems and Spectrum-X Ethernet fabric, and Together AI’s inference platform. The resulting stack is designed to support organizations that need to deploy and operate AI services across cloud and hybrid environments, particularly for workloads where throughput, response time, and infrastructure utilization directly affect operating costs.

“Enterprises are in a race to adopt agentic AI at scale to drive real business outcomes,” said Alan Peacock, general manager of IBM Cloud. “IBM and NVIDIA are delivering scalable, economical, enterprise-grade AI infrastructure that can help Together AI accelerate innovation for the next generation of AI infrastructure.”

Together AI selected IBM and NVIDIA based on their GPU infrastructure roadmaps and their ability to provide capacity at the pace needed for AI service expansion. The company raised an $800 million Series C at an $8.3 billion valuation in July to expand its AI Native Cloud platform.

IBM characterized the agreement as part of its broader work with NVIDIA across AI infrastructure and software. The companies have also referenced joint work around GPU-native data analytics, unstructured data extraction, hybrid infrastructure, and consulting services. For IBM Cloud customers, the Together AI deployment adds another route to production-grade inference infrastructure built on NVIDIA’s latest HGX platform and high-performance Ethernet networking.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @ibm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ibm-and-together-ai-…] indexed:0 read:2min 2026-08-11 ·