Two Clouds Pick Cerebras in One Week to Make Multi-Silicon Inference Fast

Two Clouds Pick Cerebras in One Week to Make Multi-Silicon Inference Fast

Gimlet Labs and General Compute each announced Cerebras deployments within a day of one another, with Gimlet Cloud targeting up to 3,000 tokens per second and General Compute financing a wafer-scale buildout for agentic coding with a $400 million debt facility. Futurum’s view is that these targets are achievable only because Cerebras’ CS-4 rack lets disaggregation software place decode on SRAM and prefill on GPUs inside one data center.

What Is Covered in This Article:

  • Gimlet Labs adding Cerebras wafer-scale compute to Gimlet Cloud with planned speeds of up to 3,000 tokens per second, 100 MW of capacity, and CS-4 access expected in 2027
  • Disaggregation methods spanning prefill-decode, attention-FFN, and speculative decoding across GPUs and the Cerebras Wafer Scale Engine
  • CS-4 enablers including 44 GB of on-wafer SRAM per wafer, 2.4 Tb/s of per-wafer Ethernet I/O, and 2 microsecond Direct Wafer Links
  • General Compute’s parallel multi-year Cerebras agreement for agentic coding, backed by a $400 million Upper90 debt facility
  • Competitive positioning against NVIDIA’s Rubin roadmap, AMD Helios pairings, SambaNova, and Cerebras’ own inference cloud

The News: Gimlet Labs and Cerebras Systems (NASDAQ: CBRS) announced on September 28 a collaboration to deliver ultrafast AI inference through Gimlet Cloud, a purpose-built disaggregated inference cloud spanning datacenter infrastructure to developer APIs. The companies plan speeds of up to 3,000 tokens per second for demanding agentic and real-time applications, with the first Cerebras-powered Gimlet Cloud datacenter expected to come online later this year. Gimlet is a launch partner for the Cerebras CS-4, plans to deploy 100 megawatts of Cerebras-powered inference capacity, and expects CS-4 access in 2027.

“Inference speed matters. It determines how productive AI can be. Fast inference creates magical user experiences and opens new markets,” said Zain Asgar, Co-Founder and CEO of Gimlet Labs.

The next day, General Compute announced a multi-year agreement with Cerebras to deploy ultra-fast inference for agentic coding workloads, with availability planned for Q1 2027 and the buildout financed by a $400 million debt facility from Upper90.

“Agentic coding is the clearest example. Agents make thousands of sequential calls, and latency compounds into wall-clock time,” said Finn Puklowski, Co-Founder and CEO of General Compute. Gimlet Labs, backed by Andreessen Horowitz and Menlo Ventures, plans to expand the collaboration across software, APIs, developer tooling, and production operations.

Two Clouds Pick Cerebras in One Week to Make Multi-Silicon Inference Fast

Analyst Take: Two clouds selecting Cerebras in one week is the first commercial test of whether multi-silicon inference can beat the homogeneous GPU cloud on the terms that matter to buyers: speed at production scale. Gimlet is the deeper technical story of the pair. Gimlet Cloud combines the Cerebras Wafer Scale Engine with GPUs in an integrated solution, using inference disaggregation to run each phase of model execution on the silicon best suited to it, and the companies plan speeds of up to 3,000 tokens per second for agentic and real-time workloads. Cerebras has already posted 4,400 tokens per second per user on GPT-OSS-120B, a figure attributed to Artificial Analysis, so the raw decode speed exists. The open question is the qualifier “at scale,” because per-user records measured at low concurrency on a 120 billion parameter model say little about serving frontier models to thousands of simultaneous agents.

Gimlet’s answer is to disaggregate the workload. Keep decode on SRAM, prefill on GPUs, and build the data center around that split from the start. Enterprise buyers are ready to pay for the result. In ETR’s AI Product Series Theme 3, September 2026 survey, 69.9% of respondents named performance among the most important factors when assessing a foundation model (n=511), ahead of cost at 68.7% and up from 63.8% in March 2025, and 78.8% named efficiency improvements the primary KPI for judging their own AI applications (n=600). Futurum sees the announcement as the clearest evidence yet that inference speed is graduating from a benchmark into a distinct market category with its own purpose-built clouds.

The timing matters because the CS-4 supplies the missing infrastructure. In Futurum’s coverage of the CS-4 launch at Supernova, the conclusion was that the substance of that generation sits in the rack rather than the silicon with doubled per-wafer power delivery, direct liquid cooling, and a networking overhaul. Those rack-level choices, presented in architectural depth at Hot Chips 2026, are what make Gimlet’s model buildable. At Hot Chips, Futurum observed every major chip team converging on memory efficiency and the elimination of synchronization penalties as the central design battleground and the Gimlet deal shows what that convergence looks like when it reaches the data center floor.

Disaggregation Is How a Speed Record Becomes a Production Number

The 3,000 tokens per second plan makes sense only as a composite of techniques, each mapped to different hardware. Gimlet describes three disaggregation methods on its inference cloud. Prefill-decode disaggregation runs the compute-bound prompt processing phase on GPUs and hands token generation to the Cerebras wafer, whose 44 GB of on-wafer SRAM per wafer moves data at 43.2 PB/s across three WSE-3 Turbo processors in a CS-4 rack, removing the HBM bottleneck that caps GPU decode speed. Attention-FFN disaggregation splits within each model layer, placing attention on GPUs and expert activation on SRAM-centric chips, trading some latency for much higher throughput. Speculative decoding disaggregation runs a small draft model at extreme speed on SRAM hardware while GPUs verify large batches of tokens efficiently. Stacked together, these methods explain the arithmetic: the CS-4 already claims more than 1,000 tokens per second on models exceeding 10 trillion parameters, and speculative decoding multiplies effective decode speed by two to three times when the draft model runs fast enough. Gimlet claims the combined approach yields 3 to 10x higher interactivity at a given throughput target, or 3 to 10x more throughput per kilowatt at a given interactivity target. The 3,000 tokens per second number is a plan stated in forward-looking language rather than a measured result. The architecture makes the target plausible and only production validation will make it a procurement fact.

The CS-4’s Ethernet Fabric Makes the Multi-Silicon Data Center Buildable

Fine-grained disaggregation has a physical precondition that the industry has been slow to price in. The heterogeneous hardware must sit in the same building, on a fabric everything can speak. Attention-FFN disaggregation ping-pongs between GPU and wafer at every model layer, so cross-datacenter links are disqualifying, and even prefill-decode splits degrade when the hop grows long. This is where the CS-4 announcements support the Gimlet deal. The CS-4’s Wafer I/O Module doubles off-wafer bandwidth to 2.4 Tb/s per wafer and 7.2 Tb/s per system, speaks standard RoCE v2 RDMA over Ethernet, and adds Direct Wafer Links joining wafers at 2 microseconds, with Arista Networks Etherlink switches scaling the fabric across racks. Standard Ethernet is what lets a neocloud that also fields NVIDIA GPUs, AMD hardware, and d-Matrix accelerators plug a wafer-scale decode tier into the same orchestration layer without proprietary interconnect islands. Cerebras’ decision to trade proprietary isolation for Ethernet membership has now produced its first purpose-built cloud customers.

100 Megawatts Puts Gimlet in Line Behind OpenAI for Cerebras Capacity

The scale commitments deserve as much scrutiny as the speed claims. Gimlet plans 100 MW of Cerebras-powered inference capacity, which at reported CS-4 rack power implies on the order of 700 to 800 racks, a material allocation from a supplier whose disclosed priorities run through OpenAI, G42, and AWS. Cerebras has contracted a manufacturing expansion of more than 10x in 2026 across three contract manufacturers and reported $25.4 billion in remaining performance obligations at Q2, so the capacity may exist on paper. Whether a venture-backed neocloud gets 2027 CS-4 allocation alongside a Master Relationship Agreement customer is a supply chain question, not a software one. The sequencing also leaves a gap the release does not resolve: the first Cerebras-powered Gimlet Cloud data center is expected online in late 2026, while CS-4 access arrives in 2027, which points to the initial deployment running current-generation systems at lower ceilings than the headline target. For Cerebras, the deal diversifies a customer base that concentration risk disclosures show is narrow, and it converts the CS-4’s Ethernet openness into a distribution channel.

Specialist Clouds, Not Hyperscalers, Are the Channel for Wafer Scale

Two specialist clouds committing to wafer scale in the same week points to where alternative silicon reaches production: the specialist cloud tier, where operators can design the data center around the workload, absorb novel racks, and finance accelerator purchases with debt rather than a hyperscaler’s capital committee. General Compute’s $400 million Upper90 facility shows the financing template, debt secured against inference capacity commitments rather than equity dilution, and its agentic coding focus shows the demand mechanism. Thousands of sequential calls per task compound decode latency into wall-clock time, so speed is the product. The multi-silicon cloud is becoming its own category with two distinct shapes. Gimlet orchestrates heterogeneous hardware behind one API and General Compute wraps a single specialty workload, coding agents, around the fastest available decode. Both are queueing for the same Cerebras capacity behind OpenAI and G42, and both move Cerebras from selling systems to selling through channels.

The strategic puzzle is that Cerebras operates its own fast-growing inference cloud, with cloud and other services revenue up 281% YoY to $126.0 million in Q2, so every neocloud deal adds a reseller beside the vendor’s own front door. The bull case is complementary reach. Gimlet brings multi-silicon orchestration, GPU capacity, and developer tooling that Cerebras does not field, General Compute brings a workload-specific buyer base in coding agents, and Cerebras gains token volume without owning every datacenter. The bear case is channel conflict as these clouds open public sign-ups and Cerebras expands its own API business against 600 MW of contracted capacity.

The competitive field is moving on both flanks. NVIDIA is productizing disaggregation as a single-vendor feature across its Rubin roadmap, including prefill-specialized silicon, which would compress the window in which multi-vendor orchestration is differentiating. AMD supplies Helios as a prefill partner in Cerebras’ own pairings while selling Instinct at the same time they evaluate the CS-4. SambaNova has published RDU plus GPU results and public clouds can bolt disaggregation software onto homogeneous fleets. Gimlet’s defensible position is physical: co-located heterogeneous hardware on one fabric with software that maps workload slices to silicon, a configuration that does not exist in public clouds today and cannot be replicated by software alone.

What to Watch:

  • Whether the first Cerebras-powered Gimlet Cloud datacenter comes online in 2026 and on which Cerebras system generation
  • Whether a disclosed benchmark validates 3,000 tokens per second at stated concurrency on a named production model
  • Whether Cerebras allocates CS-4 capacity to Gimlet in 2027 alongside OpenAI and G42 commitments
  • Whether General Compute’s Cerebras-powered agentic coding capacity reaches availability in Q1 2027
  • Whether NVIDIA’s single-vendor disaggregation stack narrows the multi-silicon advantage before Gimlet reaches public availability

See the complete details on the Cerebras website.


Sources

  1. Gimlet Labs Adds Cerebras to Deliver Ultrafast AI Inference through Gimlet Cloud Deployment Combines Cerebras Wafer-Scale Compute with Gimlet’s Inference Cloud to Deliver Up to 3,000 Tokens per Second, Cerebras

Declaration of generative AI and AI-assisted technologies in the writing process: This content has been generated with the support of artificial intelligence technologies. Due to the fast pace of content creation and the continuous evolution of data and information, The Futurum Group and its analysts strive to ensure the accuracy and factual integrity of the information presented. However, the opinions and interpretations expressed in this content reflect those of the individual author/analyst. The Futurum Group makes no guarantees regarding the completeness, accuracy, or reliability of any information contained herein. Readers are encouraged to verify facts independently and consult relevant sources for further clarification.
Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.
Read the full Futurum Group Disclosure.

Other Insights From Futurum:

Cerebras CS-4 Doubles Power and Triples Wafers in New Rack

Oracle Fusion Claw: Autonomous Enterprise Execution

SCSK and Vector Bet on Edge ECU Co-Creation to Win SDV Software

Author Information

Brendan Burke, Research Director

Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers. 

Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.

Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.

Related Insights
AMD Acquires World Labs for $8.2 Billion to Design Silicon Around World Models
September 30, 2026

AMD Acquires World Labs for $8.2 Billion to Design Silicon Around World Models

Brendan Burke, Research Director at Futurum, shares insights on AMD's $8.2 billion World Labs acquisition and why verified video generation leadership makes world models the workload AMD's next silicon generations...
NetApp Novus, PEAK:AIO and NetApp's Two-Market AI Strategy
September 30, 2026

NetApp Novus, PEAK:AIO and NetApp’s Two-Market AI Strategy

Nick Patience and Mitch Ashley, VPs and Practice Leads at Futurum, share their insights on NetApp Novus, the planned PEAK:AIO acquisition, and how NetApp is targeting AI factories and the...
SCSK and Vector Bet on Edge ECU Co-Creation to Win SDV Software
September 30, 2026

SCSK and Vector Bet on Edge ECU Co-Creation to Win SDV Software

SCSK and Vector Informatik partnered to create an integrated edge ECU software platform reducing development effort by 50–80%, positioning SCSK as a cross-border automotive software ecosystem integrator....
NXP Secures 300mm Supply Control at a Discount as VSMC's Singapore Fab Opens
September 29, 2026

NXP Secures 300mm Supply Control at a Discount as VSMC's Singapore Fab Opens

Brendan Burke, Research Director at Futurum, shares his insights on the VSMC Singapore fab opening and why NXP's 40% stake secures 300mm specialty supply at a discount while the full-ramp...
Snapdragon X Elite Expands Qualcomm's PC Reach With Googlebook
September 29, 2026

Snapdragon X Elite Expands Qualcomm’s PC Reach With Googlebook

Olivier Blanchard, Research Director at Futurum, shares insights on Qualcomm's Googlebook launch, Android integration, and competition with Intel....
Is Qualcomm’s New Ultra-Premium Handset 2nm Snapdragon 8 Elite Gen 6 Extreme SOC In A League Of Its Own?
September 29, 2026

Is Qualcomm’s New Ultra-Premium Handset 2nm Snapdragon 8 Elite Gen 6 Extreme SOC In A League Of Its Own?

Olivier Blanchard, Research Director at The Futurum Group shares insights on the new 2nm Snapdragon 8 Elite Gen 6, examining Qualcomm’s two-tier flagship strategy and the software support its advanced...

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.