We have officially collided with the physical friction points of the AI buildout. As enterprise demand scales into multi-trillion token agentic loops, the bottleneck has fundamentally shifted from software capabilities to the raw engineering maturity of the data center floor. The resulting token price paradox means raw compute is no longer a deflationary commodity but a tightly rationed premium utility. In this environment, the ultimate competitive weapon is token price relief.”

Brendan Burke

Research Director, Semiconductors, Supply Chain & Emerging Tech

Predictions:

Friction at the Infrastructure Frontier Will Drive Token Prices Higher

In the back half of this year, the AI infrastructure market will collide with a structural paradox. While enterprise token demand has reached a multi-trillion-dollar utility scale, physical system integration barriers will temporarily break the historical trend of deflationary token pricing. The extreme complexity and structural immaturity of next-generation, liquid-cooled rack-scale architectures, combined with rigid enterprise capital constraints, will cause an intermediate re-rating of token prices higher as cutting-edge reasoning models debut before next-gen token factories can fully scale.

  • Massive Token Production Outpaces Lab Projections: Enterprise token volume is no longer a forward-looking metric—it has matured into an active, heavy data center load. 54.7% of decision-makers plan to produce 51 trillion or more in tokens annually. To put this in perspective, when Meta disclosed a burn rate of roughly 60T tokens in a single month (~720T annualized), it was viewed as an outlier. Today, a significant portion of corporate IT buyers occupy that exact same bracket, transforming token generation into a core infrastructure utility bill.
  • Infrastructure Bottleneck & Capex Walls: Rising demand is colliding with physical and financial realities. 23.3% of infrastructure decision-makers now cite CapEx limits as their primary barrier to scaling AI compute. This is compounded by a severe deployment bottleneck, where 61.8% of leaders report token production lead times exceeding 4 months due to the complexity of liquid-cooled rack architectures. Simultaneously, an acute NAND flash shortage is inflating hardware bills of materials, driving up system integration costs, and squeezing deployment margins.
  • The Monolithic Pivot to NVIDIA Vera Rubin: Faced with CapEx limits, enterprises are consolidating around monolithic, co-designed architectures to maximize throughput. A definitive 74% of compute decision-makers plan to adopt NVIDIA’s Vera Rubin NVL72 platform for its unified networking fabric and promise of 35x token per megawatt gains. 
  • Leveraging Custom Silicon for Token Price Relief: Deploying application-specific custom silicon allows infrastructure vendors to bypass the 2026 token price re-rating, transforming architectural optimization into an aggressive pricing weapon. Utilizing specialized internal silicon shifts the vendor competitive landscape away from a pure arms race of raw processing speed and into a targeted price-performance war, effectively stealing high-volume enterprise contracts from competitors tethered to premium commercial GPU margins.
  • Capacity Planning For Cooling and Networking: For infrastructure and capacity planners, the primary challenge is navigating a highly volatile hardware pipeline while preparing data centers for aggressive networking and cooling overhauls. With long hardware lead times and a storage shortage, planners must map out power allocation and liquid cooling retrofits long before hardware arrives. This friction is compounded by a near-universal cycle of networking switch upgrades, encouraging capacity teams to synchronize server arrivals with high-speed network availability to prevent next-generation clusters from sitting idle. Vendors that bring GPU-accelerated software simulation into data center design can stand out in system performance.
  • Early Adoption of Rack-scale Token Factories: Securing early system validation and shipment allocation pipelines for next-generation platforms—specifically NVIDIA’s Vera Rubin NVL72 and AMD’s open-standard Helios architecture—instantly positions a vendor as a tier 1 provider capable of commanding pricing premiums from enterprise buyers trapped in multi-month supply queues. By establishing early operational mastery over the complex systems, first movers entrench their primary multi-trillion token production factories within the specific hardware and software fabrics of these pioneer deployments.

Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers. 

Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.

Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.

Recent Insights, News & Research

Synopsys, Cadence, and Siemens Take Agentic Chip Design Autonomous at DAC
July 29, 2026

Synopsys, Cadence, and Siemens Take Agentic Chip Design Autonomous at DAC

Brendan Burke, Research Director at Futurum, examines how NVIDIA and AMD are driving agentic chip design at DAC 2026 through collaborations with Synopsys, Cadence, and Siemens spanning models, physics libraries,...
Cadence Q2 FY 2026 Earnings Climb on Agentic AI and Record Backlog
July 29, 2026

Cadence Q2 FY 2026 Earnings Climb on Agentic AI and Record Backlog

Brendan Burke, Research Director at Futurum, reviews Cadence's Q2 FY 2026 earnings, agentic AI demand, record backlog, and raised full-year outlook....
AMD Advancing AI 2026: Does AMD Now Build the World’s Best CPUs and GPUs?
July 28, 2026

AMD Advancing AI 2026: Does AMD Now Build the World’s Best CPUs and GPUs?

Brendan Burke, Research Director at Futurum, shares his insights on AMD Advancing AI 2026, where Lisa Su declared Helios the world’s best AI rack and best data center CPU and...
STMicroelectronics Q2 FY 2026 Earnings Climb on Cloud AI and Satellite Strength
July 27, 2026

STMicroelectronics Q2 FY 2026 Earnings Climb on Cloud AI and Satellite Strength

Brendan Burke, Research Director at Futurum, reviews STMicroelectronics' Q2 FY 2026 earnings, the broad demand recovery, and the raised AI data center and satellite revenue ambitions....
Texas Instruments Q2 FY 2026 Earnings Climb on Broad-based Analog Growth
July 24, 2026

Texas Instruments Q2 FY 2026 Earnings Climb on Broad-based Analog Growth

Brendan Burke, Research Director at Futurum, reviews Texas Instruments' Q2 FY 2026 earnings, the broad industrial recovery, doubling data center revenue, and automotive re-acceleration....
AMD Helios Reaches Parity with Vera Rubin NVL72: Can Open Standards Outflank NVIDIA?
July 24, 2026

AMD Helios Reaches Parity with Vera Rubin NVL72: Can Open Standards Outflank NVIDIA?

Brendan Burke, Research Director at Futurum, examines AMD's Helios launch and whether spec advantages over Vera Rubin NVL72 and a bet on open standards can convert into data center share....

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.