IBM is betting $240 million that the next AI infrastructure war gets fought over renting out inference capacity for open-source models, not building bigger training clusters.
IBM and Together AI announced a multiyear, $240 million deal on August 11, 2026, to build a dedicated Nvidia Blackwell inference cluster on IBM Cloud. The initial buildout uses roughly 2,000 Nvidia Blackwell 300 chips packed into HGX B300 systems, tied together with Nvidia's Spectrum-X networking. It's set to go live in the first quarter of 2027. That's according to IBM's own newsroom announcement.
Here's the part that matters most: this cluster won't train anything. Not one bit. It exists purely to run inference, the actual serving of answers, for open-source and open-weight models. That's a deliberate bet, and it's a different bet than the one Microsoft, Google and Amazon have mostly been making with their own Nvidia buildouts.
Together AI already runs a sizable open-source inference business, serving models like DeepSeek, MiniMax and Kimi to enterprise customers who want an alternative to closed systems from OpenAI, Anthropic and Meta. The company says its platform now processes 400 trillion tokens a month. Demand for the new IBM cluster is already ahead of supply. Together AI's chief revenue officer, Kai Mak, said the capacity will likely sell out two to three months before it's even ready for service, according to BNN Bloomberg's report on the deal.
That's the tell. Enterprises aren't just chasing the smartest model anymore. They're chasing the cheapest way to run a model that's good enough, and open-weight options let them avoid the pricing and lock-in risk that comes with a single closed-model vendor. IBM is positioning itself as the landlord for that shift, not the model builder.
IBM's re-entry into a market it mostly sat out #
IBM has not been a serious player in large-scale AI infrastructure for years. AWS, Microsoft Azure and Google Cloud built the hyperscale GPU fleets that trained and served the current generation of frontier models, while IBM focused on enterprise software, consulting and its own smaller Granite model line. This deal changes that calculus. At least at the margins. A dedicated 2,000-chip Blackwell cluster is not huge by hyperscaler standards. But it's a real, named commitment, with a real dollar figure and a real delivery date, and it gives IBM Cloud something concrete to sell against the big three.
Frankly, the timing tells its own story. Nvidia's Blackwell B300 chips are still fresh off the production line, and building an entire cluster around inference rather than training is a wager that the economics of serving models will matter more to enterprise buyers over the next year than the economics of building them. IBM doesn't need to out-spend Microsoft or Google on training silicon to win this bet. It just needs the capacity, priced right. That's the whole bet: enough capacity for companies that have already decided open-weight models are good enough for their workloads.
For founders and investors building on open-source models, this is a concrete signal, not a speculative one. That matters. Capacity dedicated to inference at this scale, with a named go-live date, gives builders a real data point on where pricing and availability for open-weight serving are headed. If Together AI's forecast holds and the cluster does sell out months before it powers on, that scarcity will show up in what enterprises pay to run open models at scale, well before Q1 2027 arrives. Nvidia doesn't lose either way. For its part, it keeps winning regardless of who runs the cluster or what it's used for. Every HGX B300 system sold, whether it trains a frontier model or serves inference for an open-weight one, is still a sale for Nvidia. The real contest, the one IBM just spent $240 million to enter, is over who controls the pipes those chips run through.
Also read: DeepSeek Ships V4 Pro, Rivaling Claude and GPT-5 for a Fraction of the Cost • Elon Musk Says SpaceX AI Revenue Will Overtake Rockets by September • India's Power Grid May Not Be Ready to Fuel the AI Data Center Boom