IBM signed a multi-year $240 million agreement with Together AI to deploy a dedicated NVIDIA HGX B300 inference cluster on IBM Cloud in Q1 2027. Sixteen months ago IBM rented Blackwell capacity from CoreWeave to train Granite. Futurum reads the deal as IBM’s entry into dedicated AI capacity, de-risked by a tenant that expects full offtake before the cluster powers on.
What Is Covered in This Article:
- The $240 million multi-year agreement between IBM and Together AI for NVIDIA HGX B300 systems with Spectrum-X networking, expected Q1 2027, the first dedicated large-scale inference cluster on IBM Cloud.
- Together AI’s position: FlashAttention-4 and the ATLAS-2 speculator, 400 trillion tokens served monthly, annual bookings above $1.15 billion, and an $8.3 billion valuation after an $800 million Series C.
- Futurum’s view that reserved capacity converts Together’s software gains into margin, and that IBM’s investment-grade cost of capital makes it a live rival to the AI clouds for dedicated capacity tenants.
The News: IBM (NYSE: IBM) announced on August 11 a multi-year $240 million agreement with Together AI to deploy a large cluster of NVIDIA HGX B300 systems with Spectrum-X Ethernet networking on IBM Cloud, available in Q1 2027. Together AI will use the cluster for open-source model inference; NVIDIA claims the platform delivers 30x more AI factory output than prior generations, a vendor figure. Together AI CRO Kai Mak told Reuters the cluster will hold about 2,000 Blackwell 300 chips in the US, adding, “We’ll have full offtake well before it’s ready for service.”
“Enterprises want the performance of the best frontier models without the closed-model price tag,” said Vipul Ved Prakash, CEO at Together AI, which reports serving 400 trillion tokens monthly.
IBM and Together AI: Did IBM Cloud Just Become a Neocloud?
Analyst Take: IBM and Together AI framed the deal as an infrastructure collaboration yet the structure is the story. In April 2025, IBM rented one of CoreWeave’s first GB200 deployments to train its Granite models. Sixteen months later it is the landlord, deploying roughly 2,000 Blackwell Ultra GPUs for a tenant that expects the capacity sold out before it powers on. Futurum views the agreement as IBM’s entry into dedicated AI capacity under the safest structure available, and as Together AI’s clearest move from a developer platform toward enterprise accounts.
IBM Crosses From GPU Tenant to Landlord With the Demand Risk Removed
This is the first dedicated large-scale inference cluster on IBM Cloud, and IBM built nothing until a tenant signed: Mak expects the cluster sold out 2 to 3 months before service. The arithmetic is legible. A $240 million commitment against roughly 2,000 GPUs works out to about $120,000 of contracted revenue per GPU across the undisclosed term, Futurum’s derivation, and it sits comfortably above the hardware cost of a B300 accelerator. An inference-only cluster also monetizes from the first token, in line with the inference-dominated era Futurum has documented across its 2026 silicon coverage. The risks are timing and scale: Q1 2027 lands inside NVIDIA’s Rubin transition, so B300 economics must clear a falling price umbrella, and 2,000 GPUs is a rounding error against hyperscaler fleets. The deal matters as a template rather than as capacity.
Cheap Capital Makes IBM a Rival to the AI Clouds
IBM’s ability to compete in the market depends onits cost of capital. CoreWeave, Nebius, and their peers finance GPU fleets with debt priced near 10% unsecured and that spread passes into the rental rate a tenant pays. IBM funds capacity from an investment-grade balance sheet and facilities it already owns, so it can price a dedicated cluster inside a neocloud quote at the same return, with compliance-audited data centers attached. Fleet operations at high utilization is a discipline the neoclouds have practiced for years while IBM is standing up its first cluster. If the template holds, financing spread becomes the deciding variable in the landlord market.
Reserved Capacity Turns Together’s Software Into Margin
Together sells inference at usage-based prices and buys compute through long-term commitments, earning the spread between a fixed capacity cost and the tokens its software extracts from it. Every gain from FlashAttention-4, tuned for Blackwell by Chief Scientist Tri Dao’s team, and the ATLAS-2 adaptive speculator multiplies sellable output at near-zero marginal cost. Together tells customers it can cut inference costs by up to 60x versus closed-model pricing with capacity bought at committed rates and run hot, which 400 trillion tokens of monthly traffic makes plausible. The reservation also fixes B300 supply pricing through the Rubin transition. The commitment cuts the other way if demand softens. $240 million is a fixed obligation against falling token prices, a leveraged bet that bookings above $1.15 billion keep growing into the capacity.
Two Open-Source Inference Stacks Now Share One Cloud
The open-source alignment separates this from a colocation deal. Red Hat AI Inference on IBM Cloud, powered by vLLM and llm-d with a Granite-anchored catalog, went generally available in May 2026; Together’s proprietary engine now runs beside it. Both bets pay off if open-weight models take enterprise share from closed APIs, and IBM monetizes that outcome twice: as software margin on its own service and as rent from the fastest independent engine. The distribution logic favors IBM too. Futurum’s decision maker research finds 63.9% of enterprises deploy AI through managed cloud services, and IBM’s regulated industry client base, governance stack, and OpenShift hybrid path reach accounts Together’s developer motion does not. The bear case is that AWS and Azure serve open models at subsidized prices and Groq and Fireworks compete for the same inference workloads. A Together listing inside watsonx or Red Hat AI would signal joint selling rather than cohabitation.
What to Watch:
- Whether contracted capacity disclosures in late 2026 will test Together AI’s sellout claim.
- A Together AI listing inside watsonx or Red Hat AI Inference within 12 months would resolve the vLLM service and the tenant competing for the same workloads.
- Additional inference platforms signing with IBM or other investment-grade clouds, and the neocloud rate response, will show whether cost of capital decides the landlord market.
- Named regulated industry customers on the cluster, and Together’s bookings trajectory past $1.15 billion through 2027, are the demand-side proof points.
The full announcement is available on the IBM newsroom.
Declaration of generative AI and AI-assisted technologies in the writing process: This content has been generated with the support of artificial intelligence technologies. Due to the fast pace of content creation and the continuous evolution of data and information, The Futurum Group and its analysts strive to ensure the accuracy and factual integrity of the information presented. However, the opinions and interpretations expressed in this content reflect those of the individual author/analyst. The Futurum Group makes no guarantees regarding the completeness, accuracy, or reliability of any information contained herein. Readers are encouraged to verify facts independently and consult relevant sources for further clarification.
Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.
Read the full Futurum Group Disclosure.
Other Insights From Futurum:
The Orchestrator: Who’s Conducting Your Enterprise?
Building A Digital-Native Airline: How IBM Powered Riyadh Air’s Launch
Orchestrating IT at Global Scale How IBM Consulting Enabled Nestlé’s AI-Driven Delivery Model
Author Information
Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers.
Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.
Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.

