Deciding When to Use Intel Xeon CPUs for AI Inference, AI Field Day

Deciding When to Use Intel Xeon CPUs for AI Inference, AI Field Day

Introduction

Intel presented the capabilities of Intel Xeon CPUs for AI inference at AI Field Day, filling out a complete day with a series of Intel partner presentations following the same theme. Intel has been building workload-specific acceleration into CPU designs for over a decade. The 5th Generation Xeon Scalable CPUs added an AI-specific accelerator (AMX) alongside a few new built-in accelerators. This is part of the evidence that Intel is dedicated to allowing customers to run AI on their CPUs rather than requiring add-in card accelerators for every AI use.

Ronak Shah presented this continuing vision at AI Field Day 4 where delegates wanted to understand the decision points for using older Xeon CPUs, 5th Generation Xeon Scalable or adding an off-CPU accelerator such as an NVIDIA GPU. Ronak was very clear that not all AI use cases suit Intel Xeon CPUs for AI inference and that the decision is not clear-cut. The rule of thumb seems to be that large language models (LLMs) with over 20 billion parameters will seldom deliver acceptable performance on CPUs. Smaller models and non-LLM-based AI can often use Intel Xeon CPUs for AI inference and deliver the required latency.

The AI Pipeline CPU-GPU Sandwich

Ronak outlined Intel’s view of an AI pipeline, starting with training data preparation, a CPU-dominated task that mostly involved moving data and extract-transform-load (ETL) tasks. After data preparation, the next phase is model training, which is almost always a GPU-dominated task where the massive parallelization of a GPU can be continuously loaded. The third stage is inference, deploying the AI model to do its job. Ronak sees many production uses of Intel Xeon CPUs for AI inference. Mainly, when the AI is a part of a complete business application, this use of CPU for data prep, GPU for training, and CPU for inference is what I’m calling the AI pipeline CPU-GPU sandwich.

One of the big benefits of Intel Xeon CPUs for AI inference is that you already have them. There is no need to build a specialized infrastructure just for AI. The AI application can live alongside other applications on your shared computing platform. It is essential to recognize that generative AI is not the only player in the game; most production use of AI uses much smaller models. These smaller models are ideally suited to CPUs. Notably, the AMX accelerator speeds machine vision use cases up to two orders of magnitude compared with 4th Generation Xeon Scalable. In many production use cases, using Intel Xeon CPUs for AI inference makes sense.

Disclosure: The Futurum Group is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.

Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of The Futurum Group as a whole.

Other Insights from The Futurum Group:

Intel’s AI Everywhere Event Unveils Strategic Moves in the Era of AI

Intel Developer Cloud: Driving AI Chip Design, Filling AI Workload Gap

Intel 5th Gen Xeon Scalable Processors Make Breakthroughs

Author Information

Alastair has made a twenty-year career out of helping people understand complex IT infrastructure and how to build solutions that fulfil business needs. Much of his career has included teaching official training courses for vendors, including HPE, VMware, and AWS. Alastair has written hundreds of analyst articles and papers exploring products and topics around on-premises infrastructure and virtualization and getting the most out of public cloud and hybrid infrastructure. Alastair has also been involved in community-driven, practitioner-led education through the vBrownBag podcast and the vBrownBag TechTalks.

Related Insights
Coherent Q4 FY 2026 Earnings 1.6T Transceivers Ramp, CPO Revenue Approaches
August 14, 2026

Coherent Q4 FY 2026 Earnings: 1.6T Transceivers Ramp, CPO Revenue Approaches

Brendan Burke, Research Director at Futurum, analyzes Coherent’s Q4 FY 2026 earnings, focusing on AI datacenter optics, indium phosphide capacity, CPO, NPO, and optical circuit switching....
Cisco Q4 FY 2026 Earnings Point to Broader AI Infrastructure Demand
August 14, 2026

Cisco Q4 FY 2026 Earnings Point to Broader AI Infrastructure Demand

Futurum Research analyzes Cisco’s Q4 FY 2026 earnings, focusing on AI infrastructure orders, networking demand, security traction, and FY 2027 guidance....
Can Google's Pixel 11 Series Redefine the Smartphone Experience?
August 14, 2026

Can Google’s Pixel 11 Series Redefine the Smartphone Experience?

Google's Pixel 11 series features the Tensor G6 chip, delivering advanced AI and enhanced photography. Starting at $899, these devices redefine smartphones through personalized AI and superior camera performance....
ASUS Q2 FY 2026 Earnings Hit Record Revenue on AI Servers
August 14, 2026

ASUS Q2 FY 2026 Earnings Hit Record Revenue on AI Servers

Olivier Blanchard, Research Director & Practice Lead, Intelligent Devices with Futurum, analyzes ASUS Q2 FY 2026 earnings, focusing on AI server revenue doubling, and how agentic AI compute positions the...
Microchip Q1 FY 2027 Earnings Beat as Data Center Exposure Nears $1 Billion
August 14, 2026

Microchip Q1 FY 2027 Earnings Beat as Data Center Exposure Nears $1 Billion

Brendan Burke, Research Director at Futurum, reviews Microchip's Q1 FY 2027 earnings, its data center exposure nearing $1 billion, and margin gains above its long-term model....
SK hynix's 54 Trillion Won Investment: A Strategic Move for AI Memory Dominance
August 14, 2026

SK hynix’s 54 Trillion Won Investment: A Strategic Move for AI Memory Dominance

SK Hynix invested 54 trillion won in two fabrication facilities to support Enterprise AI growth, as memory infrastructure becomes the foundational constraint for data intelligence....

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.