d-Matrix Joins NVLink Fusion. Is It NVIDIA’s Hedge on Groq?

d-Matrix Joins NVLink Fusion. Is It NVIDIA's Hedge on Groq?

d-Matrix will integrate its Raptor inference XPUs into NVIDIA MGX racks through NVLink Fusion under a multi-year collaboration announced September 10. The deal follows Corsair deployments at Gimlet Labs and Parasail that proved out decode-phase acceleration alongside NVIDIA GPUs. Futurum examines whether Raptor’s 3D-DRAM architecture gives the NVIDIA ecosystem a frontier-scale hedge on the SRAM-based Groq 3 LPX.

What Is Covered in This Article:

  • A multi-year d-Matrix and NVIDIA collaboration bringing Raptor XPUs into MGX racks through NVLink Fusion, alongside Vera CPUs, BlueField-4 DPUs, ConnectX-9 SuperNICs, and Spectrum-X networking.
  • A 144-XPU Raptor rack with 2.3 TB of 3D-DRAM capacity and 7.2 PB/s of memory bandwidth, projected at roughly 1,000 tokens per second per user on 3-trillion-parameter-class models at 1M context.
  • d-Matrix Corsair validation through Gimlet Labs speculative decoding results and Parasail’s heterogeneous prefill-decode deployment with NVIDIA GPUs
  • Raptor as a DRAM-based hedge on the SRAM-based Groq 3 LPX inside the NVIDIA inference platform.
  • ISCA 2026 early silicon data showing 4.71x higher throughput per card than HBM designs, with initial MGX availability expected Q4 2027.

The News: d-Matrix announced a collaboration with NVIDIA that includes a multi-year product roadmap and brings d-Matrix XPUs into the NVIDIA AI factory ecosystem. As an NVLink Fusion partner, d-Matrix will integrate its next-generation Raptor inference XPUs into the latest NVIDIA MGX rack reference architecture, which features NVIDIA Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs, and Spectrum-X Ethernet networking. Astera Labs joins the effort with custom connectivity solutions. The rack targets ultra-low latency inference for AI labs, hyperscalers, and neoclouds offering premium token services, using heterogeneous disaggregation to split workloads between GPUs handling compute-intensive prefill and d-Matrix XPUs accelerating latency-sensitive decode. Raptor, the follow-on to the Corsair platform now in production, stacks a DRAM memory die and an SRAM compute die into a single package, is expected to tape out before the end of 2026, and is being evaluated at AI hyperscalers and frontier labs. Initial availability of Raptor XPUs in the NVIDIA MGX rack is expected in Q4 2027.

‘Being integrated into NVIDIA’s latest MGX rack-scale infrastructure with NVLink Fusion means our customers can deploy our inference XPUs alongside the broadly available NVIDIA AI factory platform,’ said Sid Sheth, Founder and CEO of d-Matrix.

d-Matrix Joins NVLink Fusion. Is It NVIDIA’s Hedge on Groq?

Analyst Take: The d-Matrix NVLink Fusion collaboration is the clearest statement yet of how NVIDIA manages the inference challengers it takes seriously: invite them into the rack and monetize everything around them. d-Matrix gains entry into the most widely deployed AI infrastructure ecosystem on the market, with a mature MGX supply chain and modular cable-free trays replacing the PCIe card form factor that has bounded Corsair deployments. NVIDIA gains a second answer for the ‘premium token economy,’ a decode accelerator whose 3D-DRAM capacity reaches model sizes the SRAM-based Groq 3 LPX was never designed for. Jensen framed NVLink Fusion as ‘expanding accelerator choice for customers building the next generation of AI factories,’ and the economics behind that vision are visible in the NVIDIA parts list needed to support the MGX deployment. The distance between announcement and revenue is equally visible. Raptor has yet to tape out, initial MGX availability arrives in Q4 2027, and the Groq 3 LPX entered full production this year. d-Matrix has traded a measure of independence for entry into the ecosystem it once positioned against and accepted a two-year wait to complete the trade.

Gimlet and Parasail Deployments Validate the Decode Specialist Before Raptor Arrives

The partnership rests on commercial evidence that heterogeneous inference works with today’s silicon. In March, Gimlet Labs integrated Corsair into Gimlet Cloud alongside GPUs, with software that routes each segment of an agentic workload to the hardware that serves it best. Gimlet’s published benchmark ran gpt-oss-120b with a 1.6-billion-parameter speculative decoder and measured a 2-5x end-to-end request speedup on configurations tuned for interactivity, rising to 10x on energy-optimized configurations against the same speculative decoder running on GPUs at equivalent power. In July, Parasail deployed Corsair alongside NVIDIA Hopper and Blackwell GPUs, assigning prefill to the GPUs and decode to Corsair, and claimed up to 10x faster interactive inference with up to 3x better energy efficiency. The architectural pattern they demonstrate is the same one NVIDIA now endorses at rack scale. Parasail proved the model on cards wired together over PCIe. NVLink Fusion moves the same split onto a scale-up fabric whose switch trays cut latency 3x by NVIDIA’s figures, which converts a clever integration into a reference design.

d-Matrix NVLink Fusion Integration Hedges the Groq Bet at Frontier Scale

NVIDIA already owns a decode specialist. The Groq licensing deal and the resulting Groq 3 LPX rack, unveiled at Hot Chips 2026 and now in full production with Nebius as first customer, packs 256 LP30 LPUs with 128 GB of aggregate SRAM at 40 PB/s of bandwidth, demonstrating 11,000 tokens per second per user on a 31B parameter coding model. That speed comes from SRAM, and SRAM capacity confines the LPX to compact models: NVIDIA’s own Rubin-LPX disaggregation modeling topped out at a 2T parameter target.

The Raptor rack attacks the other end of the curve. d-Matrix’s design specifies 18 compute trays with 144 Raptor XPUs, 2.3 TB of 3D-DRAM capacity, and 7.2 PB/s of memory bandwidth, with projected performance of roughly 1,000 tokens per second per user on 3T parameter-class models at 1M context. Those projections are pre-silicon simulations from a chip that tapes out late this year and remain targets. The capacity arithmetic, however, is design fact: 2.3 TB against 128 GB is an 18x gap that no SRAM roadmap closes. NVIDIA has assembled a decode portfolio segmented by model size, LPX for fast small-model serving and Raptor for frontier-scale contexts.

Futurum’s 1H 2026 AI chipset forecast sizes the prize: agent- and reasoning-first inference silicon grows from $35.9 billion in 2025 to $546.0 billion by 2030, passing pre-training as the largest workload segment in 2027, while the XPU sub-market d-Matrix competes in expands from $37.4 billion to $237.2 billion over the same window. The field outside the ecosystem faces a harder position. Cerebras projects up to 5,000 tokens per second per user on future wafer-scale systems, AMD acquired Taalas for hardcoded inference silicon, and OpenAI’s Jalapeño serves internal demand, yet none of these plug into the MGX supply chain that hyperscaler procurement already qualifies.

Hot Chips and ISCA Data Position Raptor to Aggregate the Entire Inference Pipeline

The most consequential detail sits in the technical record rather than the press release. d-Matrix CTO Sudeep Bhoja previewed the 3D DRAM technology at Hot Chips 2026, where the memory track treated stacked DRAM as the insurgent path to decode bandwidth, and the company’s ISCA 2026 paper on early Raptor silicon quantifies the claim. Raptor sustains roughly 100 TB/s of memory bandwidth per card at 0.45 pJ/bit of I/O energy, about 6x below HBM3, and across Llama-3.1 70B, DeepSeek-V3, Kimi K2, GPT-OSS. For Whisper and Canary speech models it delivers 4.71x higher throughput per card than HBM-based designs and 2.44x higher than SRAM-based designs, with 9.96x lower time per output token than HBM. Those comparisons cover whole-model serving, and the paper adds that 3D-DRAM shows less sensitivity to network latency and bandwidth than SRAM architectures because it needs fewer cards per model. Read alongside the d-Matrix roadmap, which scales from Corsair at 100B class through Raptor at 3T class to Lightning at 20T class with multi-stacked DRAM, the data describes silicon capable of serving the full inference pipeline on its own racks. d-Matrix says as much, describing an architecture that works ‘independently or in partnership with GPUs.’ The NVLink Fusion announcement routes that ambition into the decode tray of an NVIDIA rack and may contain the vision in the long run. Prefill remains compute-bound territory where GPU FLOPs dominate, the 422W package runs junction temperatures up to 105°C on a first-of-its-kind stacking process, the 3D DRAM supply chain has no named manufacturing partner. Execution through those risks decides whether Raptor stays a decode complement or becomes the aggregation play its architecture permits.

What to Watch:

  • Whether Raptor tapes out before the end of 2026 and reaches Q4 2027 MGX rack availability on schedule
  • Whether hyperscaler and frontier lab evaluations of Raptor convert into named deployments
  • Whether NVIDIA segments LPX and Raptor by model size or lets them compete for the same decode sockets
  • Whether Gimlet and Parasail publish production case studies quantifying Corsair decode economics
  • Whether Groq LPX deployments at Nebius scale into a lead Raptor cannot recover by 2028

See the complete details about the collaboration on NVIDIA’s website.


Declaration of generative AI and AI-assisted technologies in the writing process: This content has been generated with the support of artificial intelligence technologies. Due to the fast pace of content creation and the continuous evolution of data and information, The Futurum Group and its analysts strive to ensure the accuracy and factual integrity of the information presented. However, the opinions and interpretations expressed in this content reflect those of the individual author/analyst. The Futurum Group makes no guarantees regarding the completeness, accuracy, or reliability of any information contained herein. Readers are encouraged to verify facts independently and consult relevant sources for further clarification.
Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.
Read the full Futurum Group Disclosure.

Other Insights From Futurum:

NVIDIA Q2 FY 2027: AI Infrastructure Demand Extends Into FY 2028

NVIDIA Nears $12.9B Deal for Hugging Face, Escalating AI Ecosystem Strategy

NVIDIA’s Credit Support Buys Exclusivity at OpenAI’s Ohio Data Center

Author Information

Brendan Burke, Research Director

Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers. 

Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.

Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.

Related Insights
Can the IBMArm Dual Architecture Processor Align the Mainframe With the Agentic CPU Market
September 14, 2026

Can the IBM/Arm Dual Architecture Processor Align the Mainframe With the Agentic CPU Market?

Brendan Burke, Research Director at Futurum, shares insights on IBM’s native Arm mainframe design and the execution tests that remain before deployment....
Adobe Q3 FY 2026 AI Momentum Builds Amid Leadership Transition
September 14, 2026

Adobe Q3 FY 2026: AI Momentum Builds Amid Leadership Transition

Futurum Research analyzes Adobe’s Q3 FY 2026 earnings, including AI-first product adoption, freemium user growth, leadership changes, and the outlook for monetization....
Oracle Q1 FY 2027 AI Infrastructure Contracts Convert Into Growth
September 14, 2026

Oracle Q1 FY 2027: AI Infrastructure Contracts Convert Into Growth

Futurum Research analyzes Oracle’s Q1 FY 2027 earnings, including OCI growth, AI contract conversion, data center spending, and agentic enterprise products....
NVIDIA Groq 3 LPX’s Promise of World’s Fastest Inference Enters Full Production
September 14, 2026

NVIDIA Groq 3 LPX’s Promise of World’s Fastest Inference Enters Full Production

Brendan Burke, Research Director at Futurum, shares his insights on how NVIDIA Groq 3 LPX strengthens Vera Rubin and what cloud providers must prove before faster tokens support premium pricing....
NVIDIA and SpaceXAI Link Grok Expansion With Orbital Computing
September 14, 2026

NVIDIA and SpaceXAI Link Grok Expansion With Orbital Computing

Brendan Burke, Research Director at Futurum, shares insights on NVIDIA Vera CPU adoption across Grok and Starmind and the evidence required to validate orbital AI computing....
FPT IS Builds Vietnam's Court KPI Platform in 60 Days
September 14, 2026

FPT IS Builds Vietnam’s Court KPI Platform in 60 Days

Vietnam's Supreme People's Court launched an AI-integrated KPI Platform in 60 days, unifying document management, task tracking, and staff evaluation to accelerate public-sector digital transformation....

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.