AI Infra Summit 2026 – Can Optics Break the AI Power Wall by 2028?

AI Infra Summit 2026 – Can Optics Break the AI Power Wall by 2028?

Analyst(s): Brendan Burke
Publication Date: September 22, 2026

With power-capped data centers as the binding constraint, NVIDIA, AMD, and Intel raced on agentic tokens per megawatt, challengers from Cerebras to Qualcomm attacked memory placement, and an optics ecosystem from Ayar Labs to Scintil Photonics set 2028 as its volume deadline.

What Is Covered in This Article:

  • AI Infra Summit 2026 in Santa Clara on September 15-18 with keynotes from NVIDIA, AMD, AWS, Meta, and Google
  • NVIDIA’s claim of 30x AI factory throughput per megawatt for Vera Rubin on agentic workloads
  • Challenger silicon from Cerebras, d-Matrix, SambaNova, Qualcomm, and AWS Trainium is converging on memory placement over peak compute
  • Co-packaged optics roadmaps from Ayar Labs, Lumentum, iPronics, Salience Labs, Opticore, and Scintil Photonics targeting 2028 volume
  • Manufacturing hurdles of 99.9% yield thresholds, laser supply, thermal coupling, and missing interoperability standards
  • EDA orchestration of AI-designed silicon and physical AI from Cadence, Siemens, and Synopsys

The Event—Major Themes & Vendor Moves: AI Infra Summit 2026, convened September 15-17 at the Santa Clara Convention Center, now in its ninth year, drew 8,000 attendees who build the infrastructure AI runs on. Keynotes included Ian Buck of NVIDIA, Forrest Norrod of AMD, Peter DeSantis of AWS, Santosh Janardhan of Meta, and Google’s David Patterson, with Pat Gelsinger framing the optics agenda at an Ayar Labs ecosystem briefing and the OCI multi-source agreement holding its first public panel with Meta, Broadcom, and AWS. Futurum attended in person.

The program organized itself around one constraint. U.S. power capacity is growing roughly 3% per year, while Gelsinger argued the world needs on the order of 10,000x more AI than it can economically deploy today, so every vendor pitched its architecture as a route to more revenue per megawatt rather than more peak compute. Three answers structured the week. Incumbents re-architected the full stack for agentic workloads whose average input lengths now reach 142,000 tokens, challengers moved memory closer to compute to cut the energy of data movement, and the optical ecosystem declared 2028 the year co-packaged optics must reach volume manufacturing. The EDA vendors positioned themselves as the layer that orchestrates all of it, from AI-accelerated chip design to the physics simulation behind physical AI.

AI Infra Summit 2026 – Can Optics Break the AI Power Wall by 2028?

Analyst Take: AI Infra Summit 2026 made tokens per megawatt the scoreboard of the AI buildout, and 2028 the deadline for the optical technologies that can move it. “Energy capacity in the digital AI world is economic capacity,” said Pat Gelsinger, General Partner at Playground Global and former CEO of Intel, and the week validated the frame from three directions. NVIDIA presented Vera Rubin as a per-megawatt machine rather than a faster chip; the challenger field attacked the power spent moving data between memory and compute, and the photonics supply chain committed to specific yield, laser, and standards milestones that must land within roughly two years.

The bull case is coherence, with compute vendors, memory architects, and optics suppliers optimizing the same metric for the first time. The bear case is sequencing, because the optical transition depends on 99.9% yields, multi-source standards, and laser volumes that do not yet exist, and a slip past 2028 would leave scale-up domains stuck at the reach of copper while agentic demand compounds. AI Infra Summit 2026 positioned optical connectivity and physical AI as the industry’s answers to power constraints, yet left claims open that still require production validation.

Agentic Workloads Turn Tokens per Megawatt Into the Scoreboard for NVIDIA, AMD, and Intel

NVIDIA quantified how far the workload has moved from the chatbot era it was benchmarked on. Ian Buck reported that agentic workloads are roughly 100x more demanding than the 2023 chat baseline, with average input sequence lengths of 142,000 tokens against roughly 1,000 three years ago, KV caches that must be preserved across thousands of sub-agents, and turn counts no longer paced by human reading speed. His conclusion was that the number that matters is total token throughput per megawatt. NVIDIA reported Vera Rubin delivering 30x better AI factory throughput than Grace Blackwell on agentic workloads, up to 60x at some points on the Pareto curve, results published the weekend before the show.

AI Infra Summit 2026 – Can Optics Break the AI Power Wall by 2028
Source: Futurum

The CPU claim carried third-party support from Signal65, which ran the CPU side of Terminal-Bench on a Vera C2 server and measured agents completing tool calls 1.6x faster, up to 2x on code compilation and validation tasks. NVIDIA now delivers the agent’s critical path end-to-end, from the monolithic Vera die with 40% lower memory latency to smart storage managing KV cache across the data center.

AI Infra Summit 2026 – Can Optics Break the AI Power Wall by 2028
Source: Signal65

AMD answered with adaptability rather than a counter-rack. Forrest Norrod argued the biggest risk is not choosing the wrong technology but building infrastructure that cannot adapt, and positioned Helios, now in production with the MI455, as an open-standards rack built on roughly two dozen OCP specifications with derivative designs coming from other vendors with different fabrics and accelerators. AMD’s CPU numbers targeted NVIDIA directly. Sixth-generation EPYC Venice delivers 20% higher performance per core than NVIDIA’s 88-core Vera part at similar core counts, 1.3x per socket, and up to 2.5x the throughput for agentic sandbox workloads.

AI Infra Summit 2026 – Can Optics Break the AI Power Wall by 2028
Source: Futurum

Intel made its case through its CEO rather than a product keynote, and the message validated the CPU revival from the supply side. “The CPU right now is in high demand,” said Lip-Bu Tan, CEO of Intel, pointing to reinforcement learning, agentic AI, the control plane, and orchestration, and adding that Intel can serve only a percentage of customer demand while CEOs call him asking for more. The execution evidence backed the confidence. Tan said 18A is in volume production, the 14A 0.9 PDK will be done in October with the 1.0 PDK in Q1 2027, and Intel’s $23 billion raise came in 5.6x oversubscribed. Customer proof arrived from the show floor, where SambaNova credited its Xeon 6 adoption with unlocking regulated banking deployments in which the trusted CPU security stack is the admission ticket for brownfield enterprise data centers. The summit showed all three vendors now compete for that growth on agentic tool-call latency rather than general-purpose benchmarks.

AI Infra Summit 2026 – Can Optics Break the AI Power Wall by 2028
Source: Futurum

Challengers Attack the Memory Wall Rather Than NVIDIA’s Compute Lead

The challenger field no longer argues it can outperform NVIDIA on token throughput and instead attacks where the power goes, in the movement of data between memory and compute. Cerebras claimed an order-of-magnitude inference speed lead from its wafer-scale SRAM design and previewed a next-generation disaggregated system that uses other vendors’ chips for prefill while its wafers decode, with its named scaling constraints now energy grids and concrete rather than silicon. d-Matrix is building its next-generation Raptor XPU on NVLink Fusion, connecting to Vera CPUs and NVLink switch trays, with customer announcements teased. AWS disclosed that Project Rainier has scaled past 1 million Trainium chips with commitments for 5 additional gigawatts including 2 gigawatts with OpenAI, that next-generation Trainium will adopt NVLink Fusion, and that its supply chain now targets two weeks from finished chip to revenue-generating tokens by pre-building racks before the silicon arrives.

AWS Trainium3 UltraServer (Liquid-cooled)

AI Infra Summit 2026 – Can Optics Break the AI Power Wall by 2028
Source: Futurum

Qualcomm gave the memory argument its most complete architectural form. “Tokens per watt has become the new key metric in this AI war,” said Tony Pialis, who leads Qualcomm’s data center business, projecting that worldwide AI power consumption approaches roughly 25% of U.S. energy demand by the end of the decade, while a single agent call now spawns 50 to 100 inference calls and more than 1 million tokens against a chatbot’s 500. Qualcomm’s answer of “high bandwidth compute,” or HBC, bonds a DRAM stack directly to the XPU that Pialis says delivers more than 6x bandwidth per watt versus HBM and more than 200x memory capacity per watt versus SRAM-based designs, with a first generation releasing in 2027 at an 18x bandwidth improvement over legacy solutions and a second generation at 54x on an annual cadence. The Dragonfly product line wraps that memory play in bespoke prefill and decode processors, a C1000 data center CPU claiming more than 2x performance per watt against the x86 and Arm field, and the modular software stack Qualcomm opened to all hardware at its ModCon event.

The pattern comes with a dependency. Three of the five challengers now route their scale-up story through NVIDIA’s own interconnect, since NVLink Fusion and the NVHBM base die program, which NVIDIA says frees up to 30-40% of XPU compute area, puts the incumbent’s IP inside the alternatives to it. Heterogeneity is real, and SambaNova’s demonstration of disaggregated inference across SambaNova, Intel, and NVIDIA silicon in a single cloud shows enterprises will buy it, but the tollbooth on the scale-up network increasingly belongs to NVIDIA either way.

Benchmarks remain the other gap. SambaNova argued buyers should dismiss first-party benchmarks and treat SemiAnalysis’s AgentX, with its real agentic traces and normalized comparisons, as the most representative test, and both SambaNova and the Trainium team say they are working toward submissions. Until challenger silicon appears on a common agentic harness, the 2x-to-10x efficiency claims made across the week remain these rather than procurement mandates.

Co-Packaged Optics Has a 2028 Deadline and a Yield Problem

The optics case begins with physics that no roadmap escapes. Passive copper reach collapses with data rate, from roughly 3 meters at 100G signaling to about 1 meter at 200G and under a quarter meter at 400G, and Marvell’s scale-up architects put the practical limit at about 2 meters as rates move from 224G to 448G, which confines copper scale-up domains to a single rack just as KV cache growth pushes memory demand beyond 144 interconnected GPUs. “The network is the AI,” Gelsinger argued, describing the most expensive compute in history idling while it waits for KV cache and context memory.

The structural answer arrived through standards. The OCI MSA, with Meta, Broadcom, and AWS as founding members, deliberately standardizes only what travels on the fiber, decoupling optical from electrical line rates through multi-wavelength micro-ring architectures, and its members sketched a five-generation path from Gen 1’s 4x50G toward 3,200G that Broadcom expects to make a substantial share of new SerDes optically enabled by 2029-2030. The challenger ecosystem is converging on that window. Ayar Labs anchored its high-volume path on its TSMC COUPE partnership, Lumentum said its laser capacity is ready for the foreseeable future while flagging wide error bars five years out, iPronics pitched silicon photonics optical circuit switching that reconfigures roughly 1,000x faster than MEMS on a path from 32×32 toward 300×300 radix in a market it sizes at $10 billion by 2028, and Scintil Photonics presented multi-wavelength laser integration on a foundry model targeting the same 2028 ramp.

The manufacturing disclosures are where the 2028 target faces its hardest tests. MediaTek’s ASIC team stated the requirement plainly. Once an optical engine is attached to a package carrying tens of thousands of dollars of HBM and compute, a failed attach scraps working silicon, so component and attach yields must reach 99.9% before the economics close, and XPU endpoints, with their tens-of-degrees instantaneous thermal swings, will adopt optics last, after the thermally stable switch. Wiwynn reported that assembling a single CPO rack currently takes about four hours and requires new cleanroom-like environments, operator retraining, and automation before cluster-scale volumes are possible. Lumentum’s own framing, that a niche compound semiconductor laser industry must now match CMOS-scale volume and readiness, is the honest summary. The missing piece is interoperability. Panelists across the day called for a hyperscaler to anoint a standard among the four announced at OFC and for an optical plugfest equivalent to PCI compliance testing, because single-sourced optics is a non-starter for every large buyer. 2028 may be achievable for switch-adjacent optics, while endpoint CPO at the XPU likely arrives later. The winners will be decided by yield curves and MSA consolidation rather than by link demonstrations.

EDA Vendors Become the Orchestration Layer for AI-Designed Silicon and Physical AI

The quiet enablers of the week were the design tool vendors, present in nearly every startup’s speed story. Etched attributed its decision to tape out a full-reticle chip on its first attempt to large-scale emulation with Synopsys, running more than a half dozen models end-to-end before committing to silicon, and now runs on-premises servers partly to feed AI agents doing kernel, PPA, and telemetry work. Meta’s accelerator team reported design tasks that took 2 weeks, now completing in a day with AI assistance, and put fully automated “dark factory” chip design 2 to 3 years out, with verification as the remaining human bottleneck. Cadence’s Tensilica group pitched the premise directly, arguing that every chip, from startup to trillion-dollar vendor, is now a collection of IP, which makes the EDA and IP layer the coordination point for the proliferating heterogeneous designs the rest of the show described. Synopsys extended the same orchestration argument into physical AI, presenting agentic AI as the force lowering the barrier to entry for simulation and CFD tools that historically demanded specialist licenses, with 100x to 1,000x acceleration claims and digital twins carrying demand into vehicle, aerospace, and energy verticals. Siemens EDA occupies the same position in 3D IC design, where Futurum has covered its packaging flow as the third leg of this orchestration layer.

AI Infra Summit 2026 – Can Optics Break the AI Power Wall by 2028
Source: Futurum

David Patterson’s session with Ian Cutress supplied the architectural frame that ties the EDA story to the memory and optics arguments. He described a memory-centric era in which architects must now start from memory rather than logic, endorsed processing-in-memory and high-bandwidth flash as the credible new entrants precisely because they repackage proven cells rather than introduce new ones, and noted that optical interconnect, economical at chip-to-chip reach, could finally make small, standardized chiplets viable. His economics cut through the buildout anxiety. Google Cloud’s leadership says a GPU server pays for itself in 2 years and a TPU in 1 year, while 6-year-old TPUs remain fully utilized, evidence that demand still outruns depreciation. The design layer, not the fab, is where the industry’s 10,000x efficiency gap will be met or missed, because every path presented at the summit, from near-memory silicon to co-packaged optics to physics-informed world models, requires co-design across boundaries that today’s tools and standards are only beginning to span.

AI Infra Summit 2026 – Can Optics Break the AI Power Wall by 2028
Source: Futurum

What to Watch:

  • Whether the OCI MSA Gen 2 specification arrives with multi-vendor interoperability demonstrations, the optical plugfest MediaTek called for
  • Whether a hyperscaler names a single CPO standard for scale-up deployment, the consolidation step, panelists said, must precede volume
  • Whether the optical engine component and attach yields reach the 99.9% threshold ahead of the 2028 ramp decisions
  • Whether Cerebras, SambaNova, Qualcomm, and Trainium submit to AgentX alongside NVIDIA
  • Whether AMD’s Venice per-core claims against NVIDIA’s Vera survive independent agentic benchmarking

You can read more about the event and its program on the AI Infra Summit website.


Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.

Other Insights From Futurum:

NVIDIA Engineers the Agentic Data Center with DSX Software Control and Vera CPU Acceleration

Cerebras’ $95B First Day Valuation Sizes Up the 2028 XPU Opportunity

Author Information

Brendan Burke, Research Director

Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers. 

Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.

Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.

Related Insights
AI Fuels Record Cyberattacks in Italy, and a Channel Opportunity
September 22, 2026

AI Fuels Record Cyberattacks in Italy, and a Channel Opportunity

Exprivia's Q2 2026 threat report reveals record AI-driven cyberattacks on Italian organizations, as cybersecurity becomes a top revenue driver for global channel partners....
Lititz Mutual Bets on Embedded AI to Solve the Agent Knowledge Gap
September 22, 2026

Lititz Mutual Bets on Embedded AI to Solve the Agent Knowledge Gap

Lititz Mutual Insurance uses Guidewire ProNavigator's conversational AI to enhance underwriting and claims workflows, helping agents access institutional knowledge while tackling onboarding and retention challenges in P&C insurance....
Wiz Bets on MCP to Make WIN the AI Security Integration Layer
September 22, 2026

Wiz Bets on MCP to Make WIN the AI Security Integration Layer

Wiz's MCP-powered agent integrations let partner AI agents access live security context during investigations, positioning the company at the forefront of AI-driven security innovation in a market forecast to reach...
Fastly Bets the Edge on AI Governance as Machine Traffic Tops 50%
September 22, 2026

Fastly Bets the Edge on AI Governance as Machine Traffic Tops 50%

Fastly's new AI Runtime Control, AI Firewall, and API Security capabilities address enterprise security urgency as machine-generated traffic crosses 50% on its network, with AI traffic growing 6.5x faster than...
HubSpot Bets the Platform on Growth Context
September 21, 2026

HubSpot Bets the Platform on Growth Context

Keith Kirkpatrick, Vice President, Research, at Futurum, HubSpot's Fall 2026 Spotlight release introduces Growth Context, an AI-native framework positioning the company as a CRM leader for SMB and mid-market teams....
Lattice Mach-N2 Turns the Post-Quantum Procurement Gate Into an FPGA Moat
September 21, 2026

Lattice Mach-N2 Turns the Post-Quantum Procurement Gate Into an FPGA Moat

Brendan Burke, Research Director at Futurum, shares his insights on why Lattice Mach-N2's full CNSA 2.0 compliance and integrated flash extend the secure control FPGA lead as the post-quantum procurement...

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.