Enterprise agents have started calling other agents, and the token volume behind that traffic looks nothing like the earlier wave of employees typing into a chatbot. Vendors priced tokens below cost for the past two years to build share, and those subsidies are ending as bills catch up to real usage at scale. The habit of pointing every task at the largest available model, regardless of cost, is losing ground as finance teams open the invoice and ask what they can actually afford.
Finance already tracks cost per transaction on the cloud bill, and AI spending needs that same discipline applied to tokens. Only 36% of AI Platform decision-makers track cost per request today, and just 15% track token efficiency, according to Futurum Research’s 1H2026 AI Platforms Decision Maker survey. Cutting costs before switching hardware starts with four moves: caching the prompt prefix, routing routine calls to a cheaper model, batching non-urgent work, and capping retries on failed runs.
In our latest thought leadership brief, The Token Cost Reckoning Arrives in the CFO’s Office, completed in partnership with Google Cloud, Futurum Research examines the token economics reshaping enterprise AI budgets and lays out the architecture questions that determine what an agent workload actually costs to run.
In this report, you will learn:
- How to calculate cost per successful run (CPS), the metric that connects agent spend to what customers actually accept
- Four practical moves that cut token costs before a hardware switch: prompt caching, model routing, batching, and retry caps
- Why compute silicon choice, not just model choice, drives the blended cost of every agent transaction
- How software portability lets workloads move across accelerators without paying twice for idle capacity
- What questions finance should bring to the CIO at the next budget review
If you are interested in learning more, be sure to download your copy of The Token Cost Reckoning Arrives in the CFO’s Office today.
Author Information
Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers.
Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.
Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.
Daniel is the CEO of The Futurum Group. Living his life at the intersection of people and technology, Daniel works with the world’s largest technology brands exploring Digital Transformation and how it is influencing the enterprise.
From the leading edge of AI to global technology policy, Daniel makes the connections between business, people and tech that are required for companies to benefit most from their technology investments. Daniel is a top 5 globally ranked industry analyst and his ideas are regularly cited or shared in television appearances by CNBC, Bloomberg, Wall Street Journal and hundreds of other sites around the world.
A 7x Best-Selling Author including his most recent book “Human/Machine.” Daniel is also a Forbes and MarketWatch (Dow Jones) contributor.
An MBA and Former Graduate Adjunct Faculty, Daniel is an Austin Texas transplant after 40 years in Chicago. His speaking takes him around the world each year as he shares his vision of the role technology will play in our future.
