IBM and Together AI: Did IBM Cloud Just Become a Neocloud?

IBM and Together AI: Did IBM Cloud Just Become a Neocloud?

IBM signed a multi-year $240 million agreement with Together AI to deploy a dedicated NVIDIA HGX B300 inference cluster on IBM Cloud in Q1 2027. Sixteen months ago IBM rented Blackwell capacity from CoreWeave to train Granite. Futurum reads the deal as IBM’s entry into dedicated AI capacity, de-risked by a tenant that expects full offtake before the cluster powers on.

What Is Covered in This Article:

  • The $240 million multi-year agreement between IBM and Together AI for NVIDIA HGX B300 systems with Spectrum-X networking, expected Q1 2027, the first dedicated large-scale inference cluster on IBM Cloud.
  • Together AI’s position: FlashAttention-4 and the ATLAS-2 speculator, 400 trillion tokens served monthly, annual bookings above $1.15 billion, and an $8.3 billion valuation after an $800 million Series C.
  • Futurum’s view that reserved capacity converts Together’s software gains into margin, and that IBM’s investment-grade cost of capital makes it a live rival to the AI clouds for dedicated capacity tenants.

The News: IBM (NYSE: IBM) announced on August 11 a multi-year $240 million agreement with Together AI to deploy a large cluster of NVIDIA HGX B300 systems with Spectrum-X Ethernet networking on IBM Cloud, available in Q1 2027. Together AI will use the cluster for open-source model inference; NVIDIA claims the platform delivers 30x more AI factory output than prior generations, a vendor figure. Together AI CRO Kai Mak told Reuters the cluster will hold about 2,000 Blackwell 300 chips in the US, adding, “We’ll have full offtake well before it’s ready for service.”

“Enterprises want the performance of the best frontier models without the closed-model price tag,” said Vipul Ved Prakash, CEO at Together AI, which reports serving 400 trillion tokens monthly.

IBM and Together AI: Did IBM Cloud Just Become a Neocloud?

Analyst Take: IBM and Together AI framed the deal as an infrastructure collaboration yet the structure is the story. In April 2025, IBM rented one of CoreWeave’s first GB200 deployments to train its Granite models. Sixteen months later it is the landlord, deploying roughly 2,000 Blackwell Ultra GPUs for a tenant that expects the capacity sold out before it powers on. Futurum views the agreement as IBM’s entry into dedicated AI capacity under the safest structure available, and as Together AI’s clearest move from a developer platform toward enterprise accounts.

IBM Crosses From GPU Tenant to Landlord With the Demand Risk Removed

This is the first dedicated large-scale inference cluster on IBM Cloud, and IBM built nothing until a tenant signed: Mak expects the cluster sold out 2 to 3 months before service. The arithmetic is legible. A $240 million commitment against roughly 2,000 GPUs works out to about $120,000 of contracted revenue per GPU across the undisclosed term, Futurum’s derivation, and it sits comfortably above the hardware cost of a B300 accelerator. An inference-only cluster also monetizes from the first token, in line with the inference-dominated era Futurum has documented across its 2026 silicon coverage. The risks are timing and scale: Q1 2027 lands inside NVIDIA’s Rubin transition, so B300 economics must clear a falling price umbrella, and 2,000 GPUs is a rounding error against hyperscaler fleets. The deal matters as a template rather than as capacity.

Cheap Capital Makes IBM a Rival to the AI Clouds

IBM’s ability to compete in the market depends onits cost of capital. CoreWeave, Nebius, and their peers finance GPU fleets with debt priced near 10% unsecured and that spread passes into the rental rate a tenant pays. IBM funds capacity from an investment-grade balance sheet and facilities it already owns, so it can price a dedicated cluster inside a neocloud quote at the same return, with compliance-audited data centers attached. Fleet operations at high utilization is a discipline the neoclouds have practiced for years while IBM is standing up its first cluster. If the template holds, financing spread becomes the deciding variable in the landlord market.

Reserved Capacity Turns Together’s Software Into Margin

Together sells inference at usage-based prices and buys compute through long-term commitments, earning the spread between a fixed capacity cost and the tokens its software extracts from it. Every gain from FlashAttention-4, tuned for Blackwell by Chief Scientist Tri Dao’s team, and the ATLAS-2 adaptive speculator multiplies sellable output at near-zero marginal cost. Together tells customers it can cut inference costs by up to 60x versus closed-model pricing with capacity bought at committed rates and run hot, which 400 trillion tokens of monthly traffic makes plausible. The reservation also fixes B300 supply pricing through the Rubin transition. The commitment cuts the other way if demand softens. $240 million is a fixed obligation against falling token prices, a leveraged bet that bookings above $1.15 billion keep growing into the capacity.

Two Open-Source Inference Stacks Now Share One Cloud

The open-source alignment separates this from a colocation deal. Red Hat AI Inference on IBM Cloud, powered by vLLM and llm-d with a Granite-anchored catalog, went generally available in May 2026; Together’s proprietary engine now runs beside it. Both bets pay off if open-weight models take enterprise share from closed APIs, and IBM monetizes that outcome twice: as software margin on its own service and as rent from the fastest independent engine. The distribution logic favors IBM too. Futurum’s decision maker research finds 63.9% of enterprises deploy AI through managed cloud services, and IBM’s regulated industry client base, governance stack, and OpenShift hybrid path reach accounts Together’s developer motion does not. The bear case is that AWS and Azure serve open models at subsidized prices and Groq and Fireworks compete for the same inference workloads. A Together listing inside watsonx or Red Hat AI would signal joint selling rather than cohabitation.

What to Watch:

  • Whether contracted capacity disclosures in late 2026 will test Together AI’s sellout claim.
  • A Together AI listing inside watsonx or Red Hat AI Inference within 12 months would resolve the vLLM service and the tenant competing for the same workloads.
  • Additional inference platforms signing with IBM or other investment-grade clouds, and the neocloud rate response, will show whether cost of capital decides the landlord market.
  • Named regulated industry customers on the cluster, and Together’s bookings trajectory past $1.15 billion through 2027, are the demand-side proof points.

The full announcement is available on the IBM newsroom.


Declaration of generative AI and AI-assisted technologies in the writing process: This content has been generated with the support of artificial intelligence technologies. Due to the fast pace of content creation and the continuous evolution of data and information, The Futurum Group and its analysts strive to ensure the accuracy and factual integrity of the information presented. However, the opinions and interpretations expressed in this content reflect those of the individual author/analyst. The Futurum Group makes no guarantees regarding the completeness, accuracy, or reliability of any information contained herein. Readers are encouraged to verify facts independently and consult relevant sources for further clarification.
Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.
Read the full Futurum Group Disclosure.

Other Insights From Futurum:

The Orchestrator: Who’s Conducting Your Enterprise?

Building A Digital-Native Airline: How IBM Powered Riyadh Air’s Launch

Orchestrating IT at Global Scale How IBM Consulting Enabled Nestlé’s AI-Driven Delivery Model

Author Information

Brendan Burke, Research Director

Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers. 

Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.

Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.

Related Insights
OPSWAT at GISEC 2026: Can You Secure What You Can't Detect?
September 7, 2026

OPSWAT at GISEC 2026: Can You Secure What You Can’t Detect?

OPSWAT debuted its AI Content Inspector at GISEC Global 2026, introducing a new defense layer against semantic fraud and AI-generated content that bypasses traditional malware scans in enterprise environments....
HGC and Macroview Bet on AI SecOps Education to Win Enterprise Trust
September 7, 2026

HGC and Macroview Bet on AI SecOps Education to Win Enterprise Trust

HGC and Macroview Telecom co-launch an AI ASOC Workshop Series in October-November 2026, positioning themselves as trusted advisors helping enterprises modernize network and security operations amid critical SOC capacity gaps....
GPT-6 Astra Sharpens Cross-File Bug Detection at a 2.5× Price
September 5, 2026

GPT-6 Astra Sharpens Cross-File Bug Detection at a 2.5× Price

CodeRabbit's evaluation reveals GPT-6 Astra catches 20% more bugs than GPT-5.6 Sol on cross-file reviews, demonstrating superior multi-system reasoning despite premium pricing at $10/M input tokens....
Guidewire FY2026: AI Demand Accelerates the Cloud Transition
September 5, 2026

Guidewire FY2026: AI Demand Accelerates the Cloud Transition

Guidewire closed FY2026 with $1.24B ARR (19% growth) and $1.48B total revenue (23% growth), with AI emerging as the primary catalyst for insurance customers' cloud transition and platform modernization....
OpenAI’s GPT-6 Astra: Benchmarks, Cyber Risks, and Market Impact
September 4, 2026

OpenAI’s GPT-6 Astra: Benchmarks, Cyber Risks, and Market Impact

Nick Patience, VP and Practice Lead, AI Platforms at Futurum, shares his insights on GPT-6 Astra and what its cyber threshold and monitorability trade-offs mean for Anthropic and Google....
Adobe's CEO Succession Bets on Agentic AI and CX Dominance
September 4, 2026

Adobe’s CEO Succession Bets on Agentic AI and CX Dominance

Keith Kirkpatrick, Vice President & Research Director at Futurum, analyzes how Adobe's CEO succession positions the company to capitalize on surging enterprise demand for agentic AI and customer experience orchestration....

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.