Qualcomm NPU: A Key to Unlocking On-Device Generative AI?

Qualcomm NPU: A Key to Unlocking On-Device Generative AI?

The News: On February 8, Qualcomm Senior Vice President and General Manager of Technology Planning & Edge Solutions Durga Malladi published a blog post called “What is an NPU? And why is it key to unlocking on-device generative AI?” The post is a brief summary of a deeper whitepaper the company published called “Unlocking on-device generative AI with an NPU and heterogeneous computing.” The whitepaper is an in-depth look at Qualcomm’s latest on-device computing architecture which has been designed to enable generative AI applications. The paper also provides a glimpse of pragmatic on-device generative AI use cases.

Here are the key details:

  • Qualcomm has significant experience in building compute to enable on-device applications. They have been designing on-device compute for AI since 2015. Experience in these areas helped them to quickly understand generative AI compute demand and potential use cases.
  • The company’s latest Neural Processing Units (NPU) have been designed from the ground up for generative AI.
  • Because of the diverse requirements and computational demands of generative AI, different processors are needed. A heterogeneous computing architecture with processing diversity gives the opportunity to use each processor’s strengths, namely an AI-centric custom-designed NPU, along with the CPU and GPU, each excelling in different task domains.
  • Integrating processors into SoCs is key to on-device AI. This integration in chip design provides many benefits, including improvements in peak performance, power efficiency, performance per area, chip size, and cost.
  • The CPU and GPU are general-purpose processors. Designed for flexibility, they are very programmable and have ‘day jobs’ running the operating system, games, and other applications, which limits their available capacity for AI workloads at any point in time. The NPU is built specifically for AI — AI is its day job. It trades off some ease of programmability for peak performance, power efficiency, and area efficiency to run the large number of multiplications, additions, and other operations required in machine learning.
  • Applying a system approach to this heterogeneous computing solution is essential since heterogeneous computing encompasses the entire SoC, which has three layers — the diverse processors, the system architecture, and the software. The holistic view allows Qualcomm architects to evaluate constraints, requirements, and dependencies between each of these layers and then make the most appropriate choices for the SoC and end-product usage, such as designing the shared memory subsystem or deciding what data types each processor should support.
    • Since Qualcomm custom designs the entire system, they can make the appropriate design tradeoffs and use that insight to deliver a more synergistic solution.
  • Qualcomm sees on-device generative AI use cases in three categories:
    • On-demand use cases are triggered by a user, require an immediate response, and include photo/video capture, image generation/editing, code generation, audio recording transcription/summarization, and text (email, document, etc.) creation/summarization. This includes creating a custom image while texting on your phone, generating a meeting summary on your PC, or using voice to locate the nearest gas station while driving your car.
    • Sustained use cases run for a longer period and include speech recognition, gaming and video super resolution, video call audio/video processing, and real-time translation. This includes using your phone as a real-time conversation interpreter while on a business travel overseas and running super resolution every frame while gaming on your PC.
    • Pervasive use cases constantly run in the background and include always-on predictive AI assistants, AI personalization based on contextual awareness, and advanced text auto-complete. This includes your phone suggesting a meeting with a colleague based on your conversation, or your tutor assistant on your PC adjusting study material based on your answers to questions.
  • These AI use cases have two key challenges in common. First, their demanding and diverse computational requirements are difficult to meet in power- and thermally-constrained devices using general-purpose CPUs or GPUs, which serve multiple needs on the platform. Second, they are constantly evolving, so implementing them in purely fixed-function hardware can be impractical. As a result, a heterogeneous computing architecture with processing diversity gives the opportunity to use each processor’s strengths, namely an AI-centric custom-designed NPU, along with the CPU and GPU.

Read the blog post, “What is an NPU? And why is it key to unlocking on-device generative AI?” here.

You can download the whitepaper from the blog post.

Qualcomm NPU: A Key to Unlocking On-Device Gen AI?

Analyst Take: One of the biggest stories of the generative AI era to date has been the massive compute workloads generative AI requires and the scarcity of data center compute resources to meet generative AI demand. With that in mind, compute on-device processing or use cases for on-device generative AI seemed challenging. But then in October Qualcomm introduced powerful new AI-focused SoCs – Snapdragon X Elite and Snapdragon 8 Gen 3 – and became the first devices chip maker to show generative AI possibilities and share a bit about the potential. Now, Qualcomm is sharing more detail about how on device AI compute can work, and the best types of on-device generative AI use cases. Here are my thoughts on the findings.

Experience Matters

Since the beginning, processors for mobile devices have had to fit a limiting form factor. Experienced mobile chip makers like Qualcomm have eternally worked to get the most compute processing using the least power consumption as a matter of necessity. As such, they have found themselves in a unique position to lead in addressing the challenge presented by on-device compute for generative AI. Qualcomm and other mobile device processor makers have been building not only CPUs, but GPUs and NPUs for some time, and have placed those processors in a SoC for some time. Their explanation of the need for heterogeneous computing makes logical sense because of their heritage in addressing similar challenges.

Parsing Duties Within an SoC

From the whitepaper: Applying a system approach to this heterogeneous computing solution is essential since heterogeneous computing encompasses the entire SoC, which has three layers — the diverse processors, the system architecture, and the software. The holistic view allows Qualcomm architects to evaluate constraints, requirements, and dependencies between each of these layers and then make the most appropriate choices for the SoC and end-product usage, such as designing the shared memory subsystem or deciding what data types each processor should support.

Diversity of compute capabilities being built into mobile devices (smartphones, laptops, tablets) is the key to efficient on-device generative AI. It doesn’t exist (in such a manner) for cloud-based Generative AI.

Segmented Use Cases Draw a Clearer Picture

As developers think about on-device generative AI, it will help for them to think of applications in terms of the segments Qualcomm outlined – on-demand, sustained and pervasive. The segmentation also showed an understanding of the types of tasks users have come to expect from their mobile devices. In the segment descriptions, there is a strong sense that most of the use cases lean toward the utility of particular devices — the always with us nature of smartphones (photo/video capture, interpretation/translation, speech recognition) and the ubiquitous work tools that are laptop/tablets (audio recording transcription/summarization, text summarization, super video resolution, video call audio/video processing, real time translation). Many span both, including AI assistants.

Conclusion

The Qualcomm whitepaper goes into a lot more detail about the reasons the company’s SoC processors and approach are great choices. Regardless, the company does a great job of explaining how on-device generative AI can become a reality through solving the computational challenges and sheds some light on the types of on-device generative AI applications that might make sense.

Disclosure: The Futurum Group is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.

Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of The Futurum Group as a whole.

Other Insights from The Futurum Group:

With Snapdragon, Qualcomm Sets the Pace for On-Device AI

On-Device AI – The AI Moment, Episode 4

On-Device AI, Part 2 | The AI Moment, Episode 6

Author Information

Based in Tampa, Florida, Mark is a veteran market research analyst with 25 years of experience interpreting technology business and holds a Bachelor of Science from the University of Florida.

Related Insights
RingCX Goes Agentic: Can AIR Pro Win Enterprise CCaaS?
August 26, 2026

RingCX Goes Agentic: Can AIR Pro Win Enterprise CCaaS?

RingCentral unveiled AIR Pro, an agentic AI suite designed for enterprise contact centers, featuring vertical readiness for healthcare, unified workforce engagement, and built-in compliance analytics at Customer Contact Week 2026....
Google's Vertical AI Bet Governance Matters More Than Models
August 26, 2026

Google’s Vertical AI Bet: Governance Matters More Than Models

Nick Patience, VP & Practice Lead for AI Platforms at Futurum, examines Google Cloud's new vertical AI platforms for legal and financial services, and asks whether governed connectors can finally...
nCino's Q2 FY2027: Agentic AI Banking Thesis Meets Margin Reality
August 26, 2026

nCino’s Q2 FY2027: Agentic AI Banking Thesis Meets Margin Reality

nCino's Q2 FY2027 results showcase agentic AI's banking impact: subscription revenues grew 10% to $143.5M, GAAP operating income swung to $13.6M profit, aligning with 86.6% of tech leaders prioritizing agentic...
FPT IS Bets on Vietnam's Insurance Gap With Atomi Digital MOU
August 26, 2026

FPT IS Bets on Vietnam’s Insurance Gap With Atomi Digital MOU

FPT IS and Atomi Digital partner to develop end-to-end insurance technology solutions targeting Vietnam's Insurance Gap. The MOU aims to capitalize on government initiatives to grow the sector from 1.8%...
Ransomware Hits 2026 Peak: Is Your Channel Ready for AI-Driven Attacks?
August 26, 2026

Ransomware Hits 2026 Peak: Is Your Channel Ready for AI-Driven Attacks?

NCC Group's July 2026 Threat Intelligence Report reveals 894 ransomware cases—a 22% surge and 2026 peak. The emergence of JADEPUFFER, the first fully autonomous AI attack agent, signals a critical...
AI Maps Cancer's Hidden States to Predict Winning Drug Combos
August 26, 2026

AI Maps Cancer’s Hidden States to Predict Winning Drug Combos

AI algorithms identified ultraconserved cancer cell states across patients and predicted synergistic drug combinations with ~90% accuracy, challenging assumptions about tumor heterogeneity....

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.