IBM and Together AI: Did IBM Cloud Just Become a Neocloud?

IBM and Together AI: Did IBM Cloud Just Become a Neocloud?

IBM signed a multi-year $240 million agreement with Together AI to deploy a dedicated NVIDIA HGX B300 inference cluster on IBM Cloud in Q1 2027. Sixteen months ago IBM rented Blackwell capacity from CoreWeave to train Granite. Futurum reads the deal as IBM’s entry into dedicated AI capacity, de-risked by a tenant that expects full offtake before the cluster powers on.

What Is Covered in This Article:

  • The $240 million multi-year agreement between IBM and Together AI for NVIDIA HGX B300 systems with Spectrum-X networking, expected Q1 2027, the first dedicated large-scale inference cluster on IBM Cloud.
  • Together AI’s position: FlashAttention-4 and the ATLAS-2 speculator, 400 trillion tokens served monthly, annual bookings above $1.15 billion, and an $8.3 billion valuation after an $800 million Series C.
  • Futurum’s view that reserved capacity converts Together’s software gains into margin, and that IBM’s investment-grade cost of capital makes it a live rival to the AI clouds for dedicated capacity tenants.

The News: IBM (NYSE: IBM) announced on August 11 a multi-year $240 million agreement with Together AI to deploy a large cluster of NVIDIA HGX B300 systems with Spectrum-X Ethernet networking on IBM Cloud, available in Q1 2027. Together AI will use the cluster for open-source model inference; NVIDIA claims the platform delivers 30x more AI factory output than prior generations, a vendor figure. Together AI CRO Kai Mak told Reuters the cluster will hold about 2,000 Blackwell 300 chips in the US, adding, “We’ll have full offtake well before it’s ready for service.”

“Enterprises want the performance of the best frontier models without the closed-model price tag,” said Vipul Ved Prakash, CEO at Together AI, which reports serving 400 trillion tokens monthly.

IBM and Together AI: Did IBM Cloud Just Become a Neocloud?

Analyst Take: IBM and Together AI framed the deal as an infrastructure collaboration yet the structure is the story. In April 2025, IBM rented one of CoreWeave’s first GB200 deployments to train its Granite models. Sixteen months later it is the landlord, deploying roughly 2,000 Blackwell Ultra GPUs for a tenant that expects the capacity sold out before it powers on. Futurum views the agreement as IBM’s entry into dedicated AI capacity under the safest structure available, and as Together AI’s clearest move from a developer platform toward enterprise accounts.

IBM Crosses From GPU Tenant to Landlord With the Demand Risk Removed

This is the first dedicated large-scale inference cluster on IBM Cloud, and IBM built nothing until a tenant signed: Mak expects the cluster sold out 2 to 3 months before service. The arithmetic is legible. A $240 million commitment against roughly 2,000 GPUs works out to about $120,000 of contracted revenue per GPU across the undisclosed term, Futurum’s derivation, and it sits comfortably above the hardware cost of a B300 accelerator. An inference-only cluster also monetizes from the first token, in line with the inference-dominated era Futurum has documented across its 2026 silicon coverage. The risks are timing and scale: Q1 2027 lands inside NVIDIA’s Rubin transition, so B300 economics must clear a falling price umbrella, and 2,000 GPUs is a rounding error against hyperscaler fleets. The deal matters as a template rather than as capacity.

Cheap Capital Makes IBM a Rival to the AI Clouds

IBM’s ability to compete in the market depends onits cost of capital. CoreWeave, Nebius, and their peers finance GPU fleets with debt priced near 10% unsecured and that spread passes into the rental rate a tenant pays. IBM funds capacity from an investment-grade balance sheet and facilities it already owns, so it can price a dedicated cluster inside a neocloud quote at the same return, with compliance-audited data centers attached. Fleet operations at high utilization is a discipline the neoclouds have practiced for years while IBM is standing up its first cluster. If the template holds, financing spread becomes the deciding variable in the landlord market.

Reserved Capacity Turns Together’s Software Into Margin

Together sells inference at usage-based prices and buys compute through long-term commitments, earning the spread between a fixed capacity cost and the tokens its software extracts from it. Every gain from FlashAttention-4, tuned for Blackwell by Chief Scientist Tri Dao’s team, and the ATLAS-2 adaptive speculator multiplies sellable output at near-zero marginal cost. Together tells customers it can cut inference costs by up to 60x versus closed-model pricing with capacity bought at committed rates and run hot, which 400 trillion tokens of monthly traffic makes plausible. The reservation also fixes B300 supply pricing through the Rubin transition. The commitment cuts the other way if demand softens. $240 million is a fixed obligation against falling token prices, a leveraged bet that bookings above $1.15 billion keep growing into the capacity.

Two Open-Source Inference Stacks Now Share One Cloud

The open-source alignment separates this from a colocation deal. Red Hat AI Inference on IBM Cloud, powered by vLLM and llm-d with a Granite-anchored catalog, went generally available in May 2026; Together’s proprietary engine now runs beside it. Both bets pay off if open-weight models take enterprise share from closed APIs, and IBM monetizes that outcome twice: as software margin on its own service and as rent from the fastest independent engine. The distribution logic favors IBM too. Futurum’s decision maker research finds 63.9% of enterprises deploy AI through managed cloud services, and IBM’s regulated industry client base, governance stack, and OpenShift hybrid path reach accounts Together’s developer motion does not. The bear case is that AWS and Azure serve open models at subsidized prices and Groq and Fireworks compete for the same inference workloads. A Together listing inside watsonx or Red Hat AI would signal joint selling rather than cohabitation.

What to Watch:

  • Whether contracted capacity disclosures in late 2026 will test Together AI’s sellout claim.
  • A Together AI listing inside watsonx or Red Hat AI Inference within 12 months would resolve the vLLM service and the tenant competing for the same workloads.
  • Additional inference platforms signing with IBM or other investment-grade clouds, and the neocloud rate response, will show whether cost of capital decides the landlord market.
  • Named regulated industry customers on the cluster, and Together’s bookings trajectory past $1.15 billion through 2027, are the demand-side proof points.

The full announcement is available on the IBM newsroom.


Declaration of generative AI and AI-assisted technologies in the writing process: This content has been generated with the support of artificial intelligence technologies. Due to the fast pace of content creation and the continuous evolution of data and information, The Futurum Group and its analysts strive to ensure the accuracy and factual integrity of the information presented. However, the opinions and interpretations expressed in this content reflect those of the individual author/analyst. The Futurum Group makes no guarantees regarding the completeness, accuracy, or reliability of any information contained herein. Readers are encouraged to verify facts independently and consult relevant sources for further clarification.
Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.
Read the full Futurum Group Disclosure.

Other Insights From Futurum:

The Orchestrator: Who’s Conducting Your Enterprise?

Building A Digital-Native Airline: How IBM Powered Riyadh Air’s Launch

Orchestrating IT at Global Scale How IBM Consulting Enabled Nestlé’s AI-Driven Delivery Model

Author Information

Brendan Burke, Research Director

Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers. 

Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.

Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.

Related Insights
SCSK Reshuffles Leadership to Drive Strategic Initiatives
August 17, 2026

SCSK Reshuffles Leadership to Drive Strategic Initiatives

SCSK Corporation's personnel restructuring signals a strategic pivot toward AI software and consulting, capitalizing on channel partners' overwhelming demand for AI-driven solutions in 2026....
Nagarro's Strategic Shift: What Persistent's Takeover Offer Means for the Market
August 15, 2026

Nagarro’s Strategic Shift: What Persistent’s Takeover Offer Means for the Market

Persistent Systems' takeover offer for Nagarro creates a scaled AI services provider positioned to capture surging channel partner demand for AI consulting and custom application development in a market forecast...
Is the Rise of Agentic AI Threatening Cybersecurity Readiness?
August 15, 2026

Is the Rise of Agentic AI Threatening Cybersecurity Readiness?

Taiwan confirmed an autonomous AI cyber attack in July 2026. Tenable tracked seven incidents across three threat actors, signaling enterprises must urgently build defenses against offensive AI....
Unisys Positions Itself as a Leader in AI and Cloud with Upcoming Investor Conferences
August 15, 2026

Unisys Positions Itself as a Leader in AI and Cloud with Upcoming Investor Conferences

Unisys is positioning itself as a leader in AI and cloud transformation as the Software Lifecycle Engineering market reaches $344B by 2028, driven by accelerating enterprise AI adoption....
Coherent Q4 FY 2026 Earnings 1.6T Transceivers Ramp, CPO Revenue Approaches
August 14, 2026

Coherent Q4 FY 2026 Earnings: 1.6T Transceivers Ramp, CPO Revenue Approaches

Brendan Burke, Research Director at Futurum, analyzes Coherent’s Q4 FY 2026 earnings, focusing on AI datacenter optics, indium phosphide capacity, CPO, NPO, and optical circuit switching....
Cisco Q4 FY 2026 Earnings Point to Broader AI Infrastructure Demand
August 14, 2026

Cisco Q4 FY 2026 Earnings Point to Broader AI Infrastructure Demand

Futurum Research analyzes Cisco’s Q4 FY 2026 earnings, focusing on AI infrastructure orders, networking demand, security traction, and FY 2027 guidance....

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.