NVIDIA arrived at FMS 2026 arguing that the next leap in AI depends as much on the infrastructure feeding the GPU as on the GPU itself. The company open-sourced its cuFile APIs with Google, Intel, and Meta as co-maintainers, aligned more than 40 vendors behind its Storage-Next initiative, and pressed BlueField-4 STX and the CMX context tier as the template for AI-native data infrastructure. Futurum examines whether the open-source turn broadens the market or extends NVIDIA’s platform control into storage.
What is Covered in this Article
- NVIDIA moved the APIs at the heart of GPUDirect Storage to a neutral GitHub organization, with Google, Intel, Meta, and NVIDIA as inaugural maintainers.
- The Storage-Next coalition, uniting more than 40 storage and flash vendors, including DDN, KIOXIA, and Micron, is aligning GPU-driven storage behavior into interoperable standards.
- NVIDIA cited an internal test in which the Vera CPU inside BlueField-4 STX delivered up to 3.21x higher throughput than an x86 CPU in a two-stage compression and encryption pipeline.
- NVIDIA’s Context Memory Storage platform targets KV cache retention for long-context agentic inference, a hotly contested category
- The competitive stakes against AMD’s Pensando Salina DPU, hyperscaler offload architectures, and the position of the 12 storage providers codesigning STX systems.
The News: NVIDIA used the Future of Memory and Storage (FMS) conference, held August 4-6 in Santa Clara, California, to announce it is open-sourcing its cuFile application programming interfaces, the component of NVIDIA GPUDirect Storage that lets GPUs read from and write to storage directly, with access latencies measured in microseconds. The code moves to a new GitHub organization with Google, Intel, Meta, and NVIDIA as inaugural maintainers. NVIDIA also detailed Storage-Next, an industry initiative of more than 40 storage and flash vendors, including DDN, KIOXIA, and Micron, to align GPU-driven storage behavior into interoperable, open standards, and it presented its SCADA framework, which DDN is integrating with its Infinia data platform. The company cited an internal benchmark in which the Vera CPU inside NVIDIA Vera BlueField-4 STX delivered up to 3.21x higher throughput than an x86 CPU in a two-stage compression and encryption pipeline.
NVIDIA AI Storage Goes Open at FMS 2026. Is Open Source the New Moat?
Analyst Take: The NVIDIA AI storage strategy reached an inflection at FMS 2026. The company open-sourced cuFile, seated Google, Intel, and Meta as co-maintainers, and put more than 40 vendors behind Storage-Next, its effort to standardize how storage systems behave when GPUs issue thousands of concurrent requests. The underlying problem is that storage systems buckle under the parallel access patterns of a modern GPU cluster. NVIDIA’s answer routes the entire data path through its own silicon via Vera CPUs, BlueField-4 storage processors, Spectrum-X Ethernet, and the DOCA software stack. NVIDIA has concluded that the next constraint on AI economics sits between the GPU and the data, and it intends to own the architecture that removes that constraint, from the API definition down to the rack.
Open Sourcing cuFile Extends the CUDA Playbook to NVIDIA AI Storage
The cuFile playbook resembles CUDA’s early years: publish the interface, cultivate an ecosystem around it, and let the performance advantage accrue to the silicon underneath. Placing cuFile in a neutral GitHub organization lowers the perceived risk of building on GPUDirect Storage and answers procurement teams that flag single-vendor APIs as a dependency. The co-maintainer roster gives the move weight. Google and Meta operate two of the largest inference estates in the world, and Intel maintains a market share lead in server CPUs.
The open source license covers the interface. The performance that gives the interface commercial value still runs through NVIDIA hardware, which reaches the 3.21x compression-and-encryption figure is an NVIDIA internal measurement against an undisclosed x86 baseline, and the platform-level numbers from the March GTC launch of BlueField-4 STX, up to 5x token throughput, 4x energy efficiency, and 2x faster data ingestion versus traditional CPU-based storage, likewise lack a published test configuration. Until a customer deployment or MLPerf Storage validates them, these figures function as marketing exhibits rather than procurement inputs.
The KV Cache Is Becoming a Storage Product Category
Futurum has argued throughout 2026 that agentic inference is restructuring infrastructure demand, and CMX Context Memory Storage is the clearest storage market expression of that shift to date. A KV cache holds the intermediate attention calculations a model has already produced. Recomputing it on every inference step wastes GPU cycles, and evicting it to conventional storage adds latency a user can feel in a live session. As agents run longer multi-turn sessions across tools and data sources, that cache becomes a persistent working set that needs its own tier between HBM and bulk storage, and CMX, built on the STX platform, is NVIDIA’s bid to define it.
The category is already contested. WEKA’s Augmented Memory Grid claims up to 10x higher token throughput from the same GPU footprint, and DDN’s Infinia integration of SCADA pursues the same outcome through NVIDIA’s own framework. The economics explain the crowding. Futurum research finds GPUs can sit idle for more than 50% of their runtime during inference, and Futurum’s 1H 2026 Data Intelligence Decision Maker Survey finds 56.7% of enterprises already use quantization or distillation to contain inference costs. A context tier that recovers idle accelerator time attacks the same cost line as those teams are optimizing by hand, which makes the NVIDIA AI storage pitch fit budgets that already exist.
Storage Vendors Gain a Standard and Concede the Architecture
The 12 storage providers codesigning STX systems, among them Dell Technologies, HPE, IBM, Hitachi Vantara, NetApp, VAST Data, and WEKA, face a familiar bargain. Adopting the reference architecture buys immediate relevance in AI procurement cycles, with STX-based platforms due from AIC, Supermicro, and Quanta Cloud Technology in H2 2026. Adoption also migrates differentiation into NVIDIA’s controller, DOCA software, and networking, a path that can reduce an array vendor to a qualified enclosure supplier over successive product generations.
Rivals with alternatives are keeping pace. AMD paired its Pensando Salina DPU and Vulcano AI NIC with the Helios rack architecture at Advancing AI 2026, giving OEMs a second data-path stack to design against. Hyperscalers remain the hardest territory for NVIDIA, since AWS runs its own Nitro offload architecture and Google co-designed Intel’s IPU line. Storage-Next standards will matter most in enterprise and neocloud deployments, where CoreWeave, Nebius, Oracle Cloud Infrastructure, and Vultr have signed on as early adopters.
Read the blog post covering the announcements on the NVIDIA website.
What to Watch
- Whether AIC, Supermicro, and Quanta Cloud Technology systems reach production customers on schedule, and whether OEM pricing treats STX as a premium platform or a commodity reference design.
- Commit activity in the new cuFile GitHub organization through year-end, particularly whether Google, Intel, or Meta land backends for non-NVIDIA hardware.
- MLPerf Storage results or customer-published benchmarks that test the 3.21x Vera claim and the 5x token throughput figure against named baselines.
- Whether Helios rack deployments in 2027 pair the Salina DPU and Vulcano NIC with a competing context memory tier.
- How quickly CMX-class systems displace capacity economics with token economics in storage procurement.
Sources
1. Marketo Forms 2 Cross Domain request proxy frame, Nvidia, August 2026
Declaration of generative AI and AI-assisted technologies in the writing process: This content has been generated with the support of artificial intelligence technologies. Due to the fast pace of content creation and the continuous evolution of data and information, The Futurum Group and its analysts strive to ensure the accuracy and factual integrity of the information presented. However, the opinions and interpretations expressed in this content reflect those of the individual author/analyst. The Futurum Group makes no guarantees regarding the completeness, accuracy, or reliability of any information contained herein. Readers are encouraged to verify facts independently and consult relevant sources for further clarification.
Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.
Read the full Futurum Group Disclosure.
Other Insights from Futurum:
Can ADI Hot Swap Controllers De-Risk NVIDIA’s 800 VDC Transition?
Will Adobe and NVIDIA’s RTX Spark Partnership Redefine Creative AI Workflows?
NVIDIA Cosmos 3 and Open Agent Tools: Is Physical AI About to Leave the Lab?
Author Information
Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers.
Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.
Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.

