NVIDIA AI Storage Goes Open at FMS 2026. Is Open Source the New Moat?

NVIDIA AI Storage

NVIDIA arrived at FMS 2026 arguing that the next leap in AI depends as much on the infrastructure feeding the GPU as on the GPU itself. The company open-sourced its cuFile APIs with Google, Intel, and Meta as co-maintainers, aligned more than 40 vendors behind its Storage-Next initiative, and pressed BlueField-4 STX and the CMX context tier as the template for AI-native data infrastructure. Futurum examines whether the open-source turn broadens the market or extends NVIDIA’s platform control into storage.

What is Covered in this Article

  • NVIDIA moved the APIs at the heart of GPUDirect Storage to a neutral GitHub organization, with Google, Intel, Meta, and NVIDIA as inaugural maintainers.
  • The Storage-Next coalition, uniting more than 40 storage and flash vendors, including DDN, KIOXIA, and Micron, is aligning GPU-driven storage behavior into interoperable standards.
  • NVIDIA cited an internal test in which the Vera CPU inside BlueField-4 STX delivered up to 3.21x higher throughput than an x86 CPU in a two-stage compression and encryption pipeline.
  • NVIDIA’s Context Memory Storage platform targets KV cache retention for long-context agentic inference, a hotly contested category
  • The competitive stakes against AMD’s Pensando Salina DPU, hyperscaler offload architectures, and the position of the 12 storage providers codesigning STX systems.

The News: NVIDIA used the Future of Memory and Storage (FMS) conference, held August 4-6 in Santa Clara, California, to announce it is open-sourcing its cuFile application programming interfaces, the component of NVIDIA GPUDirect Storage that lets GPUs read from and write to storage directly, with access latencies measured in microseconds. The code moves to a new GitHub organization with Google, Intel, Meta, and NVIDIA as inaugural maintainers. NVIDIA also detailed Storage-Next, an industry initiative of more than 40 storage and flash vendors, including DDN, KIOXIA, and Micron, to align GPU-driven storage behavior into interoperable, open standards, and it presented its SCADA framework, which DDN is integrating with its Infinia data platform. The company cited an internal benchmark in which the Vera CPU inside NVIDIA Vera BlueField-4 STX delivered up to 3.21x higher throughput than an x86 CPU in a two-stage compression and encryption pipeline.

NVIDIA AI Storage Goes Open at FMS 2026. Is Open Source the New Moat?

Analyst Take: The NVIDIA AI storage strategy reached an inflection at FMS 2026. The company open-sourced cuFile, seated Google, Intel, and Meta as co-maintainers, and put more than 40 vendors behind Storage-Next, its effort to standardize how storage systems behave when GPUs issue thousands of concurrent requests. The underlying problem is that storage systems buckle under the parallel access patterns of a modern GPU cluster. NVIDIA’s answer routes the entire data path through its own silicon via Vera CPUs, BlueField-4 storage processors, Spectrum-X Ethernet, and the DOCA software stack. NVIDIA has concluded that the next constraint on AI economics sits between the GPU and the data, and it intends to own the architecture that removes that constraint, from the API definition down to the rack.

Open Sourcing cuFile Extends the CUDA Playbook to NVIDIA AI Storage

The cuFile playbook resembles CUDA’s early years: publish the interface, cultivate an ecosystem around it, and let the performance advantage accrue to the silicon underneath. Placing cuFile in a neutral GitHub organization lowers the perceived risk of building on GPUDirect Storage and answers procurement teams that flag single-vendor APIs as a dependency. The co-maintainer roster gives the move weight. Google and Meta operate two of the largest inference estates in the world, and Intel maintains a market share lead in server CPUs.

The open source license covers the interface. The performance that gives the interface commercial value still runs through NVIDIA hardware, which reaches the 3.21x compression-and-encryption figure is an NVIDIA internal measurement against an undisclosed x86 baseline, and the platform-level numbers from the March GTC launch of BlueField-4 STX, up to 5x token throughput, 4x energy efficiency, and 2x faster data ingestion versus traditional CPU-based storage, likewise lack a published test configuration. Until a customer deployment or MLPerf Storage validates them, these figures function as marketing exhibits rather than procurement inputs.

The KV Cache Is Becoming a Storage Product Category

Futurum has argued throughout 2026 that agentic inference is restructuring infrastructure demand, and CMX Context Memory Storage is the clearest storage market expression of that shift to date. A KV cache holds the intermediate attention calculations a model has already produced. Recomputing it on every inference step wastes GPU cycles, and evicting it to conventional storage adds latency a user can feel in a live session. As agents run longer multi-turn sessions across tools and data sources, that cache becomes a persistent working set that needs its own tier between HBM and bulk storage, and CMX, built on the STX platform, is NVIDIA’s bid to define it.

The category is already contested. WEKA’s Augmented Memory Grid claims up to 10x higher token throughput from the same GPU footprint, and DDN’s Infinia integration of SCADA pursues the same outcome through NVIDIA’s own framework. The economics explain the crowding. Futurum research finds GPUs can sit idle for more than 50% of their runtime during inference, and Futurum’s 1H 2026 Data Intelligence Decision Maker Survey finds 56.7% of enterprises already use quantization or distillation to contain inference costs. A context tier that recovers idle accelerator time attacks the same cost line as those teams are optimizing by hand, which makes the NVIDIA AI storage pitch fit budgets that already exist.

Storage Vendors Gain a Standard and Concede the Architecture

The 12 storage providers codesigning STX systems, among them Dell Technologies, HPE, IBM, Hitachi Vantara, NetApp, VAST Data, and WEKA, face a familiar bargain. Adopting the reference architecture buys immediate relevance in AI procurement cycles, with STX-based platforms due from AIC, Supermicro, and Quanta Cloud Technology in H2 2026. Adoption also migrates differentiation into NVIDIA’s controller, DOCA software, and networking, a path that can reduce an array vendor to a qualified enclosure supplier over successive product generations.

Rivals with alternatives are keeping pace. AMD paired its Pensando Salina DPU and Vulcano AI NIC with the Helios rack architecture at Advancing AI 2026, giving OEMs a second data-path stack to design against. Hyperscalers remain the hardest territory for NVIDIA, since AWS runs its own Nitro offload architecture and Google co-designed Intel’s IPU line. Storage-Next standards will matter most in enterprise and neocloud deployments, where CoreWeave, Nebius, Oracle Cloud Infrastructure, and Vultr have signed on as early adopters.

Read the blog post covering the announcements on the NVIDIA website.

What to Watch

  • Whether AIC, Supermicro, and Quanta Cloud Technology systems reach production customers on schedule, and whether OEM pricing treats STX as a premium platform or a commodity reference design.
  • Commit activity in the new cuFile GitHub organization through year-end, particularly whether Google, Intel, or Meta land backends for non-NVIDIA hardware.
  • MLPerf Storage results or customer-published benchmarks that test the 3.21x Vera claim and the 5x token throughput figure against named baselines.
  • Whether Helios rack deployments in 2027 pair the Salina DPU and Vulcano NIC with a competing context memory tier.
  • How quickly CMX-class systems displace capacity economics with token economics in storage procurement.

Sources

1. Marketo Forms 2 Cross Domain request proxy frame, Nvidia, August 2026


Declaration of generative AI and AI-assisted technologies in the writing process: This content has been generated with the support of artificial intelligence technologies. Due to the fast pace of content creation and the continuous evolution of data and information, The Futurum Group and its analysts strive to ensure the accuracy and factual integrity of the information presented. However, the opinions and interpretations expressed in this content reflect those of the individual author/analyst. The Futurum Group makes no guarantees regarding the completeness, accuracy, or reliability of any information contained herein. Readers are encouraged to verify facts independently and consult relevant sources for further clarification.
Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.
Read the full Futurum Group Disclosure.

Other Insights from Futurum:

Can ADI Hot Swap Controllers De-Risk NVIDIA’s 800 VDC Transition?

Will Adobe and NVIDIA’s RTX Spark Partnership Redefine Creative AI Workflows?

NVIDIA Cosmos 3 and Open Agent Tools: Is Physical AI About to Leave the Lab?

Author Information

Brendan Burke, Research Director

Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers. 

Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.

Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.

Related Insights
Samsung zHBM
August 7, 2026

Samsung zHBM Stacks Memory on the GPU. Has HBM Run Out of Beachfront?

Brendan Burke, Research Director at Futurum, shares his insights on Samsung zHBM, the FMS 2026 concept that stacks memory directly on AI accelerators, and on whether V10 BV-NAND and HBM4E...
AMD Acquires Taalas to Advance AI Workload Optimization
August 7, 2026

AMD Acquires Taalas to Advance AI Workload Optimization

Brendan Burke, Research Director at Futurum, shares his insights on AMD's Taalas acquisition and why the two-month model-to-silicon design flow, arriving alongside agentic EDA, is the asset that matters....
Adobe's ChatGPT Plugin Bets the Creative Suite on Conversational AI
August 7, 2026

Adobe’s ChatGPT Plugin Bets the Creative Suite on Conversational AI

Keith Kirkpatrick, Vice President & Research Director, Enterprise Software & Di at Futurum, examines how Adobe's ChatGPT plugin integrates 70+ creative tools into conversational AI, opening new distribution channels while...
Active Storage Takes Over AWS DynamoDB Adds Native Vector Search for Agentic AI
August 7, 2026

Active Storage Takes Over: AWS DynamoDB Adds Native Vector Search for Agentic AI

Brad Shimmin, VP at Futurum, analyzes the launch of native vector search in AWS DynamoDB. By embedding semantic retrieval directly into its serverless operational database, AWS eliminates fragile AI data...
IonQ Q2 FY 2026 Tempo Quantum Computers Drive Growth Ahead of SkyWater Integration
August 7, 2026

IonQ Q2 FY 2026: Tempo Quantum Computers Drive Growth Ahead of SkyWater Integration

Brendan Burke, Research Director at Futurum, analyzes IonQ’s Q2 FY 2026 earnings, focusing on quantum platform growth, SkyWater integration, security demand, and raised FY 2026 guidance....
Lattice Semiconductor Q2 FY 2026 AMI Acquisition Pulls $1 Billion Revenue Rate Up to Q3 2026
August 7, 2026

Lattice Semiconductor Q2 FY 2026: AMI Acquisition Pulls $1 Billion Revenue Rate Up to Q3 2026

Brendan Burke, Research Director at Futurum, analyzes Lattice Semiconductor’s Q2 FY 2026 earnings, focusing on AI server demand, FPGA attach, AMI, and guidance....

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.