NVIDIA’s New Rubin CPX Targets Future of Large-Scale Inference

NVIDIA’s New Rubin CPX Targets Future of Large-Scale Inference

Analyst(s): Ray Wang
Publication Date: September 18, 2025

NVIDIA has introduced Rubin CPX, a GPU purpose-built for massive-context inference. In the new Vera Rubin NVL144 CPX racks, the company highlights configurations of 8 EF NVFP4 compute, 100 TB of high-speed memory, and 1.7 PB/s bandwidth. We believe CPX represents a game-changing architecture that pressures competitors to re-architect around prefill-optimized inference.

What is Covered in this Article:

  • NVIDIA Rubin CPX announcement and Vera Rubin NVL144 CPX specs and claims.
  • CPX’s role in disaggregated serving (prefill vs. decode) and attention acceleration.
  • Analysis on competitive impact (AMD and custom silicon) and rack-scale options.
  • Practical design changes: GDDR7, PCIe Gen6, cableless modules, liquid cooling.
  • A $5 billion token revenue per $100 million invested and million-token use cases.

The News: NVIDIA has introduced Rubin CPX, a new GPU class built for massive-context processing, powering million-token coding and generative video. The Vera Rubin NVL144 CPX platform delivers eight exaflops of NVFP4 AI compute, 100TB of high-speed memory, and 1.7 PB/s bandwidth per rack – 7.5x more AI performance than GB300 NVL72 systems. A dedicated CPX compute tray is also available for existing setups.

NVIDIA positions Rubin CPX as the top option for long-context processing, combining video encode/decode with large-context inference on a single chip. It integrates into the full NVIDIA AI stack (Dynamo, Nemotron, NIM, CUDA-X), with interest from Cursor, Runway, and Magic. General availability is expected at the end of 2026, with NVIDIA projecting $5 billion in token revenue per $100 million invested.

NVIDIA’s New Rubin CPX Targets Future of Large-Scale Inference

Analyst Take: Rubin CPX is designed for the compute-heavy prefill phase, using a monolithic die optimized for NVFP4 and GDDR7 memory, while standard Rubin GPUs focus on the bandwidth-heavy decode phase. The Vera Rubin NVL144 CPX rack combines 144 Rubin GPUs, 144 Rubin CPX GPUs, and 36 Vera CPUs to hit 8 EF NVFP4, 100TB fast memory, and 1.7 PB/s bandwidth – a 7.5x jump over GB300 NVL72. Each Rubin CPX GPU delivers up to 30 PFLOPS (NVFP4), 3x faster attention, and 128GB GDDR7 for context-driven workloads.

We think CPX could be a game changer, widening NVIDIA’s lead in rack-scale design and pressuring rivals to rework their silicon strategies. In effect, Rubin CPX shifts inference economics toward disaggregated serving, making it a key term for investors and industry players to watch. Zooming out, we believe CPX positions NVIDIA even more strongly against rivals targeting the inference market—a compute segment expanding faster than training. We expect this segment to accelerate as video and coding applications powered by generative AI mature, driving increasingly complex inference workloads in the coming years.

Prefill Specialization and System Design

Rubin CPX is built for massive-context inference, handling million-token coding and long-form video within a single chip that merges video codecs with context processing. Its design emphasizes compute power (NVFP4) over bandwidth, aligning with the FLOPS-heavy prefill phase while avoiding costly HBM. Attention throughput is central, with CPX delivering 3x faster attention than GB300 NVL72 to sustain long sequences at high speed. In the Vera Rubin NVL144 CPX rack, Rubin and Rubin CPX GPUs complement each other – one for prefill, the other for decode – enabling efficient disaggregated serving across phases. This pairing makes large-scale inference faster, cheaper, and more balanced across phases.

Rack-Scale Options and Engineering Choices

The lineup includes three rack designs: VR200 NVL144, VR200 NVL144 CPX, and a Vera Rubin CPX Dual Rack pairing a VR NVL144 with a CPX rack. Engineering updates feature cableless, modular trays, liquid cooling on the CPX modules (~370 kW per NVL144 CPX rack), and daughter cards integrating CX-9 NICs, OSFP cages, NVMe, and Rubin CPX. Signal routing now uses Paladin board-to-board connectors and a PCB midplane, with NIC placement optimized for shorter high-speed paths, enabling PCIe Gen6 over PCB. These design moves focus on density, serviceability, and scale-out networking.

Competitive Implications for Rivals

We believe the emergence of Rubin CPX could compel AMD and custom silicon vendors to reassess their roadmaps, as the growing relevance of prefill-optimized hardware intersects with NVIDIA’s widening rack-scale advantage. By leveraging GDDR7, PCIe Gen6, and a lower-cost profile relative to HBM-centric designs, CPX delivers strong efficiency for prefill workloads. System-level gains beyond the chip level, such as higher aggregate memory bandwidth and an integrated rack architecture, further reinforce NVIDIA’s lead.

Monetization Claims and Early Ecosystem Signals

NVIDIA projects Vera Rubin NVL144 CPX to deliver $5 billion in token revenue per $100 million invested, tying its pitch directly to ROI in long-context inference. Early partners such as Cursor, Runway, and Magic are testing CPX for coding assistants, cinematic content, and agent-driven software engineering with massive context windows. Full-stack support (Dynamo, Nemotron, NIM, CUDA-X, AI Enterprise) ensures smooth deployment across cloud and data center environments.

With launch slated for late 2026, we see the new product as a potential incremental driver of NVIDIA’s quarterly performance from late 2026 through 2027. In light of ongoing advances in inference, we believe CPX strengthens NVIDIA’s competitive edge by providing a more cost-efficient yet high-performance solution, positioning the company to capture greater market share in inference-driven workloads.

What to Watch:

  • Practical benefits of 3x faster attention on real-world long-context workloads.
  • Adoption of Vera Rubin NVL144 CPX vs. Dual Rack, where power/cooling differ (~370 kW vs. ~190 kW for NVL144).
  • How Cursor, Runway, and Magic productize million-token contexts and generative video.
  • Competitive responses: prefill-specialized chips from AMD and custom silicon providers.
  • Deployment timelines toward end-2026 availability and integration with NVIDIA’s AI stack.

See the complete press announcement and event details on the NVIDIA Rubin CPX introduction on the NVIDIA website.

Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.

Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.

Other insights from Futurum:

Could NVIDIA’s Collaboration with MediaTek Trigger a $73 Billion Acquisition Bid?

NVIDIA Q2 FY 2026 Earnings: Networking Steals the Spotlight and Q3 Ramp Will Be Key To Watch

Is NVIDIA’s Jetson Thor the New Brain for General Robotics?

Image Credit: NVIDIA

Author Information

Ray Wang is the Research Director for Semiconductors, Supply Chain, and Emerging Technology at Futurum. His coverage focuses on the global semiconductor industry and frontier technologies. He also advises clients on global compute distribution, deployment, and supply chain. In addition to his main coverage and expertise, Wang also specializes in global technology policy, supply chain dynamics, and U.S.-China relations.

He has been quoted or interviewed regularly by leading media outlets across the globe, including CNBC, CNN, MarketWatch, Nikkei Asia, South China Morning Post, Business Insider, Science, Al Jazeera, Fast Company, and TaiwanPlus.

Prior to joining Futurum, Wang worked as an independent semiconductor and technology analyst, advising technology firms and institutional investors on industry development, regulations, and geopolitics. He also held positions at leading consulting firms and think tanks in Washington, D.C., including DGA–Albright Stonebridge Group, the Center for Strategic and International Studies (CSIS), and the Carnegie Endowment for International Peace.

Related Insights
SiFive BigSky Ships the First RISC-V Server. Is the GPU Head Node the Prize
August 28, 2026

SiFive BigSky Ships the First RISC-V Server. Is the GPU Head Node the Prize?

Brendan Burke, Research Director at Futurum, shares his insights on SiFive's BigSky, the first enterprise-grade RISC-V server, and why CUDA head node duty for NVIDIA GPUs is the socket that...
Okta Q2 FY 2027 Earnings Beat and Raise on Core Identity Strength
August 28, 2026

Okta Q2 FY 2027 Earnings Beat and Raise on Core Identity Strength

Mitch Ashley, VP and Practice Lead, CIO & Technology Buyers at The Futurum Group, reviews Okta's Q2 FY 2027 earnings, where core identity strength and new products drove a beat...
NVIDIA Nears $12.9B Deal for Hugging Face, Escalating AI Ecosystem Strategy
August 28, 2026

NVIDIA Nears $12.9B Deal for Hugging Face, Escalating AI Ecosystem Strategy

Nick Patience, VP & Practice Lead of AI Platforms at Futurum, shares his insights on NVIDIA's reported $12.9 billion bid for Hugging Face and what it would mean for the...
QumulusAI Q2 FY 2026 118% Revenue Growth for Hyperspeed AI Compute Deployment
August 28, 2026

QumulusAI Q2 FY 2026: 118% Revenue Growth for Hyperspeed AI Compute Deployment

Brendan Burke, Research Director at Futurum, analyzes QumulusAI’s Q2 FY 2026 earnings, focusing on direct AI compute demand, GPU fleet expansion, and capacity execution....
NXP Tech Days Can Physical AI Reference Designs Solidify the Neural Axis
August 28, 2026

NXP Tech Days: Can Physical AI Reference Designs Solidify the Neural Axis?

Brendan Burke and Olivier Blanchard, Research Directors at Futurum, share their insights from NXP Tech Days Silicon Valley, where reference designs from humanoid hands to a four-board MCX A5 networking...
MTG-I2 Launch Reveals Thales's Critical Infrastructure Security Depth
August 28, 2026

MTG-I2 Launch Reveals Thales’s Critical Infrastructure Security Depth

Thales Alenia Space's MTG-I2 satellite completes the Meteosat Third Generation constellation, positioning Thales as a critical infrastructure security provider for European meteorological data in the expanding cybersecurity market....

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.