Research

The Off Ramp From Per-Token Pricing

The Off Ramp From Per-Token Pricing

In Partnership with:

Agentic AI is multiplying token consumption per task by up to 100x, and that growth lands directly on the bill for organizations still running production inference on per-token serverless APIs. Per-token pricing is often the fastest path to experimentation and early production, but sustained, high-volume inference turns that same pricing model into a cost control problem that compounds exponentially rather than linearly as usage scales.

AI workload deployment has already gone hybrid: 41% of workloads run in public cloud, 36% in organizations’ own data centers, 13% in colocation, and 6% on bare metal or HPC providers. Reserved and owned compute together account for 66% of AI compute consumption, while on-demand sits at just 19%. Infrastructure tier selection, not workload design, is the primary lever organizations control to keep AI costs predictable as agentic workloads move toward production.

In our latest thought leadership report, The Off Ramp From Per-Token Pricing: How Enterprises Regain AI Cost Control With Reserved Bare Metal, completed in partnership with QumulusAI, Futurum Research examines how organizations progress from token-metered experimentation to reserved, hybrid AI infrastructure as their AI applications mature, and how that progression helps them balance cost, model control, privacy, and resource flexibility.

In this report, you will learn:

  • How AI workload deployment has already gone hybrid, and why reserved and owned compute account for 66% of AI compute consumption
  • The four-stage progression organizations follow, from serverless inference APIs to hybrid multi-tier infrastructure, as workloads move from experimentation to production
  • Why AI-first cloud, the tier spanning bare metal GPU providers, specialized AI clouds, and GPU marketplaces, is forecast to grow faster than any other deployment tier
  • A decision framework for determining which workloads belong on reserved bare metal versus hyperscaler infrastructure
  • How compute leaders at Qubrid AI, Runpod, and Amberd.ai describe their customers’ migration from token-metered APIs to dedicated, reserved infrastructure

If you are interested in learning more, be sure to download your copy of The Off Ramp From Per-Token Pricing: How Enterprises Regain AI Cost Control With Reserved Bare Metal today.

Author Information

Brendan Burke, Research Director

Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers. 

Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.

Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.