NVIDIA Aims Vera Rubin at Agentic Post-Training With Proven CoreWeave Results

NVIDIA Aims Vera Rubin at Agentic Post-Training With Proven CoreWeave Results

Analyst(s): Brendan Burke
Publication Date: July 21, 2026

NVIDIA disclosed the first measured customer results for the Vera Rubin platform and the Vera CPU: CoreWeave recorded a 10x performance-per-watt improvement over Blackwell on DeepSeek-R1, and the Vera CPU posted 1.8x on SPEC 2026 agentic workloads versus a 128-core AMD EPYC Turin, alongside reinforcement learning sandbox benchmarks from Perplexity and Prime Intellect. NVIDIA’s accompanying blog supplies the system value frame, declaring agentic post-training the central workload of the agentic era and “intelligence per dollar” its governing metric.

What Is Covered in This Article:

  • CoreWeave measured a 10x performance-per-watt improvement for Vera Rubin NVL72 over Blackwell on DeepSeek-R1 across the full Pareto frontier — NVIDIA’s first customer-measured Vera Rubin result, notable because it matches the launch claim on hardware whose firmware and kernels are far less mature than Blackwell’s.
  • The Vera CPU posted 1.8x more performance on the newly released SPEC 2026 agentic suite versus a 128-core AMD EPYC Turin and 1.6x on EDA workloads.
  • Agentic post-training as the central workload of the agentic era, governed by a new metric: “intelligence per dollar.”
  • NVIDIA claims Blackwell cuts a representative 20 billion token post-training run from ~$84,000 to ~$2,400 and raises intelligence per $1,000 from 0.85 to 29.
  • Futurum expects definitions of agents to evolve toward the RL-based paradigm, elevating sandbox startup time and single-threaded tool-calling efficiency — including sub-agent tool execution — into paramount procurement metrics.

The News: NVIDIA disclosed the first measured customer results for the Vera Rubin platform and the Vera CPU. CoreWeave — among the first providers to stand up, validate, and test Vera Rubin NVL72 racks — measured a 10x performance-per-watt improvement over Blackwell on DeepSeek-R1, running across the full Pareto frontier of throughput and interactivity. The 10x figure NVIDIA offered at launch was internal modeling, and this is its first customer-measured confirmation.

On the CPU side, the newly released SPEC 2026 suite shows the Vera CPU delivering 1.8x more performance on agentic workloads versus a 128-core AMD EPYC Turin and 1.6x on EDA workloads, while Perplexity measured 1.5x faster sandbox job completion and 1.9x faster concurrent sandbox startup versus its x86 production environment, Prime Intellect reported 30% greater RL sandbox throughput per CPU versus alternative x86 architectures, and DeepInfra completed internal testing ahead of published benchmarks.

NVIDIA framed the system value of these results in a blog published July 17, 2026, “NVIDIA Vera Rubin Maximizes Intelligence per Dollar for Post-Training Workloads,” which argues that post-training is no longer a one-time finishing step but a continuous reinforcement learning cycle, defines intelligence per dollar as intelligence gained with post-training divided by the frequency of runs times the cost per run, and presents NVIDIA Nemotron 3 Ultra — an open-weight, 550 billion-parameter mixture-of-experts model scoring 71.7% on SWE-bench Verified — as the reference workload. Vera and Vera Rubin hit production floors in the second half of 2026.

NVIDIA Aims Vera Rubin at Agentic Post-Training with Proven CoreWeave Results

Analyst Take: NVIDIA’s first measured Vera Rubin and Vera CPU results declare that the hill AI hardware climbs has moved. For a decade, the summit was pre-training floating-point throughput; over the past two years, it shifted to inference token economics; the benchmarks NVIDIA disclosed this week — CoreWeave’s 10x performance per watt on Vera Rubin and the Vera CPU’s 1.8x SPEC 2026 agentic result against AMD’s EPYC Turin — assert the next summit is the reinforcement learning loop that turns fluent models into working agents.

First Measured Performance of NVIDIA Vera Rubin NVL72
Source: CoreWeave, NVIDIA

The intelligence-per-dollar blog is the system value argument that makes those numbers cohere, and NVIDIA has enlisted DeepInfra, Perplexity, and Prime Intellect, three of the AI infrastructure innovators closest to frontier workloads, as witnesses. The case deserves to be taken seriously because it is falsifiable: NVIDIA published a reference model (Nemotron 3 Ultra), a disclosed recipe on NeMo RL, a cost table, and named customers running the workload.

It also deserves scrutiny, because nearly every number in it is vendor-supplied or partner-run, the “intelligence” in intelligence per dollar is currently defined by a single coding benchmark, and the cost-per-run comparison rests on illustrative assumptions of 20 billion rollout tokens. Agentic post-training is now the declared target for hardware hill climbing — but the procurement mandate will follow only if independent validation reproduces what NVIDIA’s partners measured.

Agentic Post-Training Becomes the Target for Hardware Hill Climbing

NVIDIA responds to customer demand by defining agentic post-training as a different compute pattern, not a larger version of an old one. An RL run orchestrates thousands of environments, generating rollouts in parallel, each rollout an agent trajectory of plan, tool call, and observation, scored against a reward and fed back into weight updates. The forward pass of that loop is inference, which means every improvement in cost per token flows directly into the cost of building intelligence — the two metrics are nested, not competing.

NVIDIA’s numbers make the economic case vivid: at $4.20 per million tokens on Hopper, a representative post-training run costs ~$84,000 and gets rationed; at $0.12 on GB300 NVL72, it costs ~$2,400 and can run continuously, with Vera Rubin claimed to extend the trajectory by training the largest models with one-fourth the GPUs of the Blackwell generation. The compute footprint grows not because any single run is larger but because the runs never stop. These figures are NVIDIA’s own, based on illustrative rollout-token assumptions, and should be treated as a thesis rather than a settled benchmark — but the direction of travel matches what the frontier labs themselves are building.

Sub-Agent Tool Execution Makes Sandbox Startup a Procurement Metric

The most consequential detail in this launch is not on the GPU at all. Every rollout in an RL run executes its tool calls — cloning repositories, running test suites, validating code, executing Python — on the CPU, inside an ephemeral sandbox. The agent loop is sequential, so you cannot throw more cores at it; you need a more performant core. Sub-agent architectures compound the demand, because each spawned sub-agent claims its own sandbox and its own single-threaded tool-execution budget, multiplying sandbox creation events per rollout.

That is why the benchmarks the partners chose to publish are so telling: Perplexity measured concurrent sandbox startup time — 1.9x faster on Vera than its x86 production environment, alongside 1.5x faster test-suite completion — and Prime Intellect measured RL sandbox throughput per CPU, up 30% on average versus alternative x86 architectures. These are not traditional CPU benchmarks; they are the operational metrics of an RL factory.

NVIDIA Vera Delivers Nearly 2x Faster Performance
Source: NVIDIA

Vera’s architecture maps directly onto them:

  • The custom Olympus core claims 2x per-core performance,
  • The second-generation Scalable Coherency Fabric on a monolithic die claims 3x core-to-core bandwidth with zero “chiplet tax”
  • The LPDDR5X subsystem claims 40% lower loaded memory latency — supported by an early SPEC 2026 result of 1.8x on agentic workloads versus a 128-core AMD EPYC Turin — nearly 2x on an industry-standard benchmark, though SPEC 2026’s agentic suite is itself new and thinly populated with submissions.
NVIDIA Vera-40% Lower Latency Under Load
Source: NVIDIA

The pattern extends beyond agents: Los Alamos National Laboratory measured 7x better performance on its Ursa agentic workload and 3x on radiation-transport and algebraic multigrid simulation, and the New York Stock Exchange’s evaluation of Redpanda, a C++ data streaming engine, delivered 6x lower P99 latency on Vera versus x86 — a data-streaming result, not a general-purpose latency claim, but a demonstration that latency under load holds outside the agent loop. DeepInfra’s forthcoming benchmark blog will be the next data point to test the pattern.

Measured Factory Results Move the 10x Claim From Model to Evidence

The platform-level proof point arrived from CoreWeave, one of the first providers to stand up, validate, and test Vera Rubin NVL72 racks: a measured 10x performance-per-watt improvement over Blackwell on DeepSeek-R1, run across the full Pareto frontier of throughput and interactivity rather than at a single cherry-picked operating point. When NVIDIA launched Vera Rubin, the 10x figure was internal modeling; this is customer-measured performance on a real MoE model, achieved on firmware and kernels far less mature than GB200’s heavily optimized stack, which suggests headroom rather than a ceiling.

The networking layer carries a large share of the gain. The sixth-generation NVLink scale-up fabric with its purpose-built switch delivers 2x more performance on AI applications versus off-the-shelf Ethernet, with 3x lower latency, 10x higher packet rates, and 130 teraflops of in-network compute. Spectrum-X Ethernet photonics demonstrates roughly 1.6x better RDMA bandwidth with reduced jitter, and co-packaged optics is claimed to improve energy efficiency by about 5x and resiliency by about 10x. In scale-across, Spectrum-XGS extends the model across sites at a claimed 1.9x versus standard Ethernet.

Agent Definitions Will Converge on the Reinforcement Learning Paradigm

We expect the definition of an agent itself to evolve. Today, “agent” mostly describes prompt-orchestration frameworks: scaffolding that chains model calls at inference time. The paradigm NVIDIA and its partners are building toward defines an agent as a policy trained by reinforcement learning in sandboxed environments against verifiable rewards, with examples including:

  • Prime Intellect continuously post-training frontier open models with NeMo Gym environments and Dynamo orchestration.
  • Perplexity syncing trillion-parameter weights between training and inference nodes in under 2 seconds to keep the loop hot.

Once that redefinition takes hold, the metrics that matter are precisely the ones surfacing in this launch: sandbox startup time, single-threaded tool-calling efficiency, weight-transfer latency, and rollouts per dollar. This reframing also restores the CPU to the center of AI infrastructure economics, a thesis Futurum has advanced through this cycle. NVIDIA’s claim of an incremental $200 billion in CPU TAM aligns with our view of a 1.9x ratio of CPUs to GPUs by 2030, inclusive of both host CPUs and standalone worker CPUs.

The Competitive Field Has Not Yet Agreed to Fight on This Hill

NVIDIA is defining the contest before its competitors have entered it. AMD’s EPYC roadmap through Venice and Intel’s Clearwater Forest continue to optimize the chiplet formula of core counts, throughput, and cost per vCPU — the economics that won the cloud era — and both will contest the “chiplet tax” framing with legitimate counterarguments on yield, cost, and manufacturability. Arm-based hyperscaler silicon, such as AWS Graviton, owns the deployment base that x86 alternatives are measured against, and AWS illustrates the split hand incumbents will play: it has announced Vera Rubin cluster deployments and is building Trainium v4 within NVLink Fusion, while retaining EFA networking and Nitro rather than adopting NVIDIA’s full pod architecture.

The bear case is straightforward: single-threaded performance under load is a real differentiator only if RL post-training becomes the dominant compute pattern, and if agent economics disappoint — or if enterprises continue consuming agents through managed services that abstract the silicon away — Vera’s premium ASP must win on TCO grounds, NVIDIA has asserted but not yet published. The counter-signal to watch is whether AMD and Intel begin publishing sandbox startup and agentic SPEC 2026 numbers of their own. The moment they do, NVIDIA will have won the framing even if it loses individual benchmarks.

What to Watch:

  • Whether Perplexity’s 1.9x sandbox startup and Prime Intellect’s 30% throughput results can be reproduced by enterprises adapting to agentic post-training techniques.
  • Whether hyperscalers beyond CoreWeave adopt the Vera CPU and full pod architecture or cherry-pick components.
  • Whether AMD or Intel publishes sandbox startup or single-threaded-under-load results rather than core-count and throughput claims.
  • Whether sandbox startup time and RL rollout throughput appear in Neocloud pricing and SLAs.
  • Whether the definition of “agent” in enterprise RFPs shifts from prompt-orchestration frameworks toward RL-trained policies with sandboxed tool execution, validating the paradigm shift this launch presumes.

Read the complete blog on Vera Rubin’s intelligence per dollar for post-training workloads on the NVIDIA website.

Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.

Other Insights From Futurum:

NVIDIA DSX Promises More Revenue per Gigawatt. Who Actually Captures It?

NVIDIA Engineers the Agentic Data Center with DSX Software Control and Vera CPU Acceleration

COMPUTEX 2026: Are Agentic CPUs Rivals or Complements?

Featured Image: NVIDIA

Author Information

Brendan Burke, Research Director

Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers. 

Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.

Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.

Related Insights
Azure's AMD Partnership Expands: Is Reinforcement Learning the Hardware Bottleneck?
July 21, 2026

Azure’s AMD Partnership Expands: Is Reinforcement Learning the Hardware Bottleneck?

Brendan Burke, Research Director at Futurum, examines how Azure's AMD Partnership expands across GPUs, CPUs, and networking to address frontier model inference and reinforcement learning workloads, signaling evolution of AI...
Fortinet's AI Controls Join the Field. Can Integration Set Them Apart?
July 21, 2026

Fortinet’s AI Controls Join the Field. Can Integration Set Them Apart?

Fernando Montenegro, VP at Futurum, examines why FortiEndpoint's consolidated AI controls are a real buyer win, while platform and SASE integration, not the individual features, will decide whether Fortinet stands...
ASUS ROG Gjallar: Is AI-Enhanced Audio the Next Gaming Peripheral Battleground?
July 21, 2026

ASUS ROG Gjallar: Is AI-Enhanced Audio the Next Gaming Peripheral Battleground?

ASUS launches ROG Gjallar, a gaming soundbar with Dolby Atmos and AI-powered echo cancellation, signaling that intelligent audio endpoints are the next gaming frontier....
Bain & Company Elevates AI Strategy as OpenAI Elite Partner
July 21, 2026

Bain & Company Elevates AI Strategy as OpenAI Elite Partner

Bain & Company achieved OpenAI Elite Partner status, positioning it to capitalize on surging AI consulting demand as 83.9% of channel partners expect AI to drive business growth in 2026....
Salesforce and Databricks Forge New AI Partnership: What's at Stake?
July 20, 2026

Salesforce and Databricks Forge New AI Partnership: What’s at Stake?

Keith Kirkpatrick, Vice President & Research Director, Enterprise Software & Di at Futurum, The Salesforce Databricks partnership enables enterprises to deploy AI agents securely while maintaining data privacy and compliance...
QumulusAI Lists on Nasdaq: The Speed Advantage of a Distributed Path to Multi-Gigawatt Scale
July 20, 2026

QumulusAI Lists on Nasdaq: The Speed Advantage of a Distributed Path to Multi-Gigawatt Scale

Brendan Burke, Research Director at Futurum, examines QumulusAI's Nasdaq listing targeting 2.5 gigawatts of GPU capacity by 2027, with distributed inference accelerating infrastructure deployment....

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.