Marvell Scales AI Memory to 48TB Behind a Single CXL Switch at FMS 2026

Marvell Scales AI Memory to 48TB Behind a Single CXL Switch at FMS 2026

Analyst(s): Brendan Burke
Publication Date: August 5, 2026

At FMS 2026, Marvell announced additions to its AI memory infrastructure portfolio spanning SSD controllers, CXL memory expansion and pooling, and optical shared memory. The launches target the KV cache bottleneck in agentic AI inference and signal that CXL is moving from hyperscaler evaluation into deployment.

What Is Covered in This Article:

  • Marvell’s FMS 2026 announcement expands its AI memory infrastructure portfolio across the Bravera SC6 PCIe 6.0 SSD controller, Structera CXL memory expansion and pooling, and Photonic Fabric optical shared memory.
  • The Bravera SC6 doubles the performance of the PCIe 5.0 Bravera SC5, supports NAND from multiple suppliers, and is expected to begin sampling in Q4 2026.
  • Marvell-reported Structera benchmarks include up to 4.8x higher inference throughput and an 82.7% reduction in time-to-first-token in GPU memory-pooling configurations.

The News: Marvell Technology announced new products across its AI memory infrastructure portfolio at FMS 2026 on August 4, positioning the lineup around three tiers: server-level AI storage with the Bravera SC6 PCIe 6.0 SSD controller, rack-level CXL memory expansion and pooling with the Structera family, and pod-level optical shared memory with Photonic Fabric memory modules, NICs, and chiplets. The company frames the portfolio as a response to growing KV cache demands in agentic AI inference, where memory capacity and bandwidth increasingly gate GPU utilization and token output.

The Bravera SC6 controller supports PCIe Gen 6 and NVMe 2.2, doubles the performance of the widely deployed Bravera SC5, and runs NAND from multiple suppliers at up to 3,600 MT/s across 16 channels. It integrates 15 embedded Arm cores, including 12 Cortex-R82 cores, and is expected to begin sampling in Q4 2026. Photonic Fabric creates a shared memory tier across XPUs and racks up to 50 meters apart, enabling up to 32TB of warm KV cache offload and, per Marvell, 2 to 3x higher token throughput within existing power envelopes.

“As AI scales, memory must scale more independently of compute so resources can be deployed where they deliver the greatest value,” said Will Chu, Executive Vice President and General Manager, Custom Cloud Solutions, at Marvell. Companion blogs detail the Structera CXL portfolio: Structera X expanders shipping to hyperscalers, the Structera A 2504 near-memory accelerator with 16 Arm Neoverse V2 cores, and the Structera S 30260 CXL 3.x switch connecting 16 or 32 hosts to as much as 48TB of pooled memory.

Marvell Scales AI Memory Infrastructure to 48TB Behind a Single CXL Switch

Analyst Take: Futurum’s recent silicon coverage has tracked a single structural shift: inference is now the dominant workload, and it is memory-bound. Marvell’s FMS 2026 portfolio is the most complete silicon expression of that thesis to date. The company is betting that AI memory infrastructure becomes a distinct procurement category with its own controllers, switches, and optical tiers, and that the vendor spanning all of them wins disproportionate content per rack.

Hyperscaler Demand Moves CXL From Evaluation Into Deployment

Marvell states that Structera X 2404 and 2504 platforms are already shipping to hyperscalers, marking what the company describes as CXL adoption “moving from evaluation into real-world deployment across hyperscale environments.”

The product design reflects a specific economic proposition. The DDR4-based Structera X 2404 lets operators repurpose existing DDR4 inventory behind current-generation servers, converting stranded memory into usable CXL-attached capacity at a fraction of the cost of provisioning new DDR5. The DDR5-based 2504 targets new builds where higher bandwidth justifies the DIMM cost. Both variants front 4 DDR channels per controller and attach via standard CXL interfaces, allowing rack architects to mix expansion capacity without changing the host CPU or platform firmware.

Marvell reports up to 4.8x higher inference throughput and an 82.7% reduction in time-to-first-token in GPU memory pooling configurations using Structera. These are vendor-published benchmarks; no hyperscaler is named and no deployment volume is quantified, so production-scale validation remains pending.

Attach Economics Could Make AI Memory Infrastructure a Volume Silicon Category

The controller math is what elevates this from a niche protocol story. Each Structera X controller fronts 4 DDR channels: the DDR5-based 2504 supports 8 DIMMs per controller, while the DDR4-based 2404 reaches 12 DIMMs at three DIMMs per channel. Card capacity is configuration dependent: a lean 4-DIMM build supplies roughly 512GB, requiring 2 cards per terabyte of expansion, while a fully populated 2504 with 128GB RDIMMs approaches 1TB on a single card.

Per-socket demand is climbing toward those numbers. Trillion-parameter-class models with long context windows push working sets toward multiple terabytes per CPU. At an 8TB per-socket memory footprint, a single server could support 8 to 16 CXL expansion controllers depending on DIMM density. That estimate is Futurum’s, derived from Marvell’s published channel and DIMM configurations, and should be treated as an upper bound pending real deployment data.

The implication is that CXL controllers may start to resemble NICs a decade ago: a per-server attach category whose unit volumes scale with server shipments. Futurum projects the data center CPU market will reach $76.6 billion by 2029, growing 34.9%, and every CXL expander attaches to a CPU socket. Agentic workloads that restore the CPU to the center of AI infrastructure pull memory expansion silicon along with them.

The rest of the Structera line deepens the attach story. The Structera A 2504 accelerator pairs 16 Arm Neoverse V2 cores at 3.2 GHz with up to 4TB of capacity and 200 GB/s of bandwidth for near-memory workloads such as DLRM and vector search, and the Structera S 30260 switch connects 16 or 32 hosts to 48TB of shared memory at 4TB/s aggregate bandwidth with under 460ns round-trip latency.

Portfolio Breadth Meets Entrenched GPU-Attached Memory and a Crowded Field

The competitive field is forming quickly. Astera Labs sells Leo CXL memory controllers into the same hyperscaler accounts, XConn ships CXL switches that compete with Structera S, and Montage Technology ships CXL memory expander controllers; Samsung, SK hynix, and Micron ship CXL memory modules while serving as Marvell qualification partners. Structera interoperability is validated on Intel Xeon, AMD EPYC, and Arm platforms, which keeps Marvell neutral across the CPU vendors.

The incumbent counter-move is architectural. NVIDIA keeps KV cache inside the NVLink domain, on HBM and Grace-attached LPDDR, and its rack-scale designs leave little PCIe real estate for third-party memory tiers; Photonic Fabric is Marvell’s answer at the pod level, though its 32TB shared tier must prove itself against NVIDIA’s tightly integrated memory hierarchy. The Bravera SC6 faces a friendlier path with its NAND-agnostic architecture aligning with hyperscaler sourcing strategy, and Q4 2026 sampling positions production SSDs for the PCIe Gen 6 server platforms arriving in 2027.

Marvell holds the broadest AI memory infrastructure portfolio in the market, and breadth is a real advantage when hyperscalers want one qualified architecture instead of stitched-together point products. The risk is that CXL software maturity, the latency gap versus local DRAM, and hyperscalers’ appetite for designing their own silicon could each compress the window before this category consolidates. Watch whether deployment disclosures catch up to the portfolio ambition.

What to Watch:

  • Public disclosure of a volume Structera X deployment within the next two quarters will validate the “evaluation to deployment” claim
  • Bravera SC6 sampling on schedule in Q4 2026 and NAND qualifications across at least three suppliers are the signals that the hyperscale sourcing flexibility pitch is landing
  • NVIDIA’s treatment of KV cache offload in its next rack-scale platforms, plus CXL 3.x product announcements from Astera Labs and XConn, will determine whether pooled memory becomes an open category or gets absorbed into proprietary GPU domains.

See the full press release on the Marvell website.


Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.

Other Insights From Futurum:

WEKA Engineers the AI Chassis to Conquer the Inference Power Paradox

Marvell Q1 FY 2027 Raises Full-Year Outlook on AI Data Center Demand

Marvell’s XConn Buy Yields a Two-Pronged Open Fabric Play Against NVLink

Featured Image: Marvell

Author Information

Brendan Burke, Research Director

Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers. 

Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.

Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.

Related Insights
Backblaze Q2 FY 2026 CoreWeave Deal Strengthens AI Storage Position
August 5, 2026

Backblaze Q2 FY 2026: CoreWeave Deal Strengthens AI Storage Position

Brad Shimmin and Brendan Burke, analysts at Futurum, review Backblaze’s Q2 FY 2026 earnings, focusing on B2 Cloud Storage growth, AI infrastructure demand, and the CoreWeave agreement....
Thales Launches Luna 8: A Quantum-Ready Security Solution for AI Era
August 5, 2026

Thales Launches Luna 8: A Quantum-Ready Security Solution for AI Era

Thales launches Luna 8, a quantum-resistant hardware security module designed for AI-era enterprise workloads. The new HSM addresses NIST's 2024 post-quantum cryptography standards, positioning Thales as a leader in enterprise...
Can Autodesk's New Robotics Lab Solve Florida's Housing Crisis?
August 5, 2026

Can Autodesk’s New Robotics Lab Solve Florida’s Housing Crisis?

Autodesk and University of Florida launch the nation's most advanced robotics industrialized construction lab, positioning the enterprise software leader to capture market share in a $762B industry projected to grow...
Amazon Q2 2026 AWS Momentum Accelerates as AI Investment Climbs
August 4, 2026

Amazon Q2 2026: AWS Momentum Accelerates as AI Investment Climbs

Futurum Research analyzes Amazon’s Q2 2026 earnings, focusing on AWS acceleration, AI infrastructure demand, custom silicon, agentic AI, and commerce execution....
August 4, 2026

Uncrewed Vessels Set to Revolutionize Anti-Submarine Warfare for Dutch Navy

Thales's selection to design uncrewed vessels for the Royal Netherlands Navy underscores OT/ICS security as defense's greatest challenge, with quantum-safe cryptography now standard....
SK hynix's HBF Technology: A Major shift for AI Infrastructure?
August 4, 2026

SK hynix’s HBF Technology: A Major shift for AI Infrastructure?

SK Hynix unveiled HBF specifications delivering 512GB capacity and 3TB/s bandwidth via UCIe interconnects, directly addressing the memory bottleneck constraining enterprise AI workloads as organizations prioritize generative AI investments....

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.