At Fully Connected 2026, CoreWeave launched Forge as a unified environment for the AI loop, put NVIDIA Vera Rubin NVL72 into production with Cognition, mapped the enterprise path from distillation to reinforcement learning, and made its case on pace, performance, and partnership.
What Is Covered in This Article:
- CoreWeave Fully Connected 2026 in San Francisco on September 29-October 1 with 5,500 attendees
- CoreWeave Forge launch unifying Weights & Biases Models, OpenPipe post-training, marimo notebooks, serverless RL, and RL Rollouts
- Enterprise post-training adoption through distillation, fine-tuning, open weight models, and customer-written evals
- Cognition benchmarks of 4.8x inference throughput and 3.8x RL rollout throughput on NVIDIA Vera Rubin NVL72 versus GB200 NVL72
- Executive perspectives on CoreWeave differentiation distilled to pace, performance, and partnership with customers
- CoreWeave Partner Network with validated integrations from VAST Data, ClickHouse, CrowdStrike, Reflection, and an agent search layer from Exa, Parallel Web Systems, and You.com
The News: Fully Connected, CoreWeave’s AI cloud conference, convened in San Francisco on September 29-October 1, drawing 5,500 customers, partners, developers, and AI leaders to hear how AI is built and run in production. Futurum attended. The company organized its new product suite around a single concept it calls the “AI loop” of run, observe, curate, improve, evaluate, and repeat.
Four announcements filled that in. CoreWeave Forge, the headline launch, is a development layer that connects the full loop in one environment, built from CoreWeave’s acquisitions of Weights & Biases, marimo, and OpenPipe together with its own services, including the generally available ARIA research agent, the new Agent Lens observability tool, CoreWeave Sandboxes, CoreWeave Registry, and serverless post-training. MasterClass and Canva are already building on it, and Free, Pro, and Enterprise editions open the platform to teams without a procurement cycle.
On hardware, CoreWeave announced availability of NVIDIA Vera Rubin NVL72 with Cognition as the first customer anywhere running production workloads on the system, and said it will offer the NVIDIA Vera CPU, which NVIDIA positions as the first CPU built for AI agents, on bare metal under the same platform as the GPU fleet. The CoreWeave Partner Network rounded out the keynote, curating integrations validated under production load across infrastructure, data services, software, and models.
The through-line is that CoreWeave, already first to 1 active gigawatt and first to run NVIDIA Vera Rubin NVL72 in production, is now claiming the same lead in the environment where models and agents continuously improve.
CoreWeave Fully Connected 2026: Forge Brings Frontier Lab RL to Every AI Team
Analyst Take: CoreWeave built its reputation as the first AI-native cloud to reach 1 active gigawatt of capacity and Fully Connected 2026 showed where that power budget goes next. What stood out most is how CoreWeave Forge, the new unified environment for the AI loop, packages frontier lab reinforcement learning for every customer. RL has been the domain of labs with dedicated infrastructure teams. Any team with a defined task and a way to score it already has the raw material for RL. Infrastructure has been the barrier and CoreWeave Forge is built to remove it.
The company organized its new product suite around that gap, assembling its acquisitions of Weights & Biases, marimo, and OpenPipe with its own serverless services into the first managed version of the post-training workflow that produced Cognition’s Devin and its peers. Customer sessions grounded the thesis, showing enterprises entering the loop through distillation and evaluation while RL demand stays concentrated in AI-native labs. The risks are operational and competitive. Reward design stays with the customer, every vendor-supplied efficiency figure awaits independent validation, and the hyperscalers can bundle a response into platforms enterprises already pay for.

CoreWeave Forge Turns Reinforcement Learning Into a Managed Service
The RL pipeline inside CoreWeave Forge spans the stack. RL Rollouts, in preview within Dedicated Inference, loads each new checkpoint’s weights directly into the running deployment so the training loop keeps turning without a redeploy, while customers bring their own trainer, reward design, and environment. CoreWeave Sandboxes, now generally available, spin up isolated RL environments in parallel and feed scores back into Weights & Biases. Agent Lens and trace tagging turn production failures into the curated signal that defines the reward target and CoreWeave says Agent Lens detects 20% more critical failures at 10% the cost of a general purpose frontier LLM. Evaluations and model lineage in CoreWeave Registry show whether the new checkpoint beat the old one before it ships.

In the keynote’s live demo, a help desk agent went from misrouted tickets to a measurably better model through serverless RL, with only slightly higher latency. That is the frontier post-training workflow, reduced to a few clicks. CoreWeave’s claim that serverless RL trains 1.4x faster at 40% lower cost than a self-managed setup is internal and should be read as a pricing thesis until customers confirm it.
This product can define the integrated frontier lab experience for the long tail of enterprises. Hyperscalers market the same loop as fragmented product experiences, with AWS splitting it across SageMaker HyperPod, Bedrock, and separate observability tooling while Google and Microsoft spread it across Vertex AI and AI Foundry components that customers stitch together themselves. The other AI clouds, including Nebius, Together AI, and Lambda, supply capacity and support DIY experimentation, leaving teams to assemble the post-training stack on their own. CoreWeave Forge occupies the space between them with a connected loop that stays open across models, frameworks, and clouds. The open question is whether enterprise buyers graduate from fine-tuning to reward engineering, the one step Forge cannot do for them.
Enterprises Start With Distillation Before Moving to Reinforcement Learning
CoreWeave executives were candid that demand for reinforcement learning today comes overwhelmingly from AI-native companies such as Cognition, and that most enterprises, while keen, are not yet running it. Enterprise investment is flowing into fine-tuning and distillation instead. CoreWeave says customers typically come to distillation to cut inference costs and stay because a smaller, narrower model performs better on their specific task. Interest in open weight models has also climbed sharply over the past six months as models such as Z.ai’s GLM and Moonshot AI’s Kimi have improved. Customers test them head to head against frontier APIs and release them to small groups of internal developers before wider rollout. These post-training workloads tend to keep customers on the platform far longer than inference on whichever frontier model is newest, which matters for a company whose revenue still rests heavily on a small number of labs.
Our customer conversations showed how far ahead of their cloud tooling that AI-native customers already are on evaluation. A physical AI company said its customers write the test, arriving with decades of automation benchmarks and reliability thresholds near 99.99%, so its evals must mirror the customer’s yardstick rather than its own. An AI search partner maintains an internal benchmark whose questions change every week alongside a partner benchmark whose ground truth changes every hour, because a static test gets gamed. An agent-native team distills real-world user trajectories into targeted benchmarks, using small models to find where sessions break down, while an engineering software vendor admitted it was unprepared for the volume of evals that agentic workloads demand. That is the demand signal Forge is built against. Agent Lens converts production traces into exactly this kind of curated benchmark material, Sandboxes give each evaluation an isolated environment, and Registry lineage records whether the new checkpoint cleared the customer’s own yardstick.
CoreWeave’s enterprise target sits below the CIO, with the platform teams that allocate GPU spend inside large organizations as buyers and AI and ML engineers as day-to-day users. The go-to-market runs bottom-up, through the Weights & Biases free tier, pay-as-you-go pricing and usage signals that flag accounts ready for committed contracts. The Forge tooling layer runs on any cloud, including inside customer environments, and CoreWeave says most Weights & Biases customers do not use its infrastructure. However, the services layer of sandboxes, serverless RL and inference runs only on CoreWeave, the free tier bundles CoreWeave compute, and observability now carries the CoreWeave brand, so Forge’s openness is also the top of a funnel into CoreWeave’s own cloud. We think that is a sound strategy, although enterprise buyers will want evidence that the cross-cloud tooling receives the same investment as the CoreWeave-only services.
Cognition’s Vera Rubin Benchmarks Give the RL Pitch Its Hardware Proof
The hardware story reinforces the software one. Cognition, the applied AI lab behind Devin, stood up a Vera Rubin NVL72 cluster in early September and ran the first customer-executed inference benchmark on the platform, measuring a 4.8x increase in total token throughput for SWE-2 inference workloads and a 3.8x boost in output token throughput for reinforcement learning workloads, both against a GB200 NVL72 baseline. When asked about the advantage over GB300, the inference advantage compressed to 2.75x. A customer-run benchmark sits a tier above vendor marketing, and the RL rollout figure matters most for the Forge thesis because rollouts are the throughput-hungry step of the loop Forge sells.

The gap between SWE-2’s 4.8x on Vera Rubin and the 10x token throughput per megawatt CoreWeave measured on the DeepSeek R1 reasoning model is a property of the workload rather than the hardware. Theoretical benchmarks run a frozen open model at steady state. SWE-2 is served mid-RL, with multi-hour agent rollouts on weights that keep updating across mixed clusters on three continents. The test for AI clouds is how much of a theoretical gain they can deliver while the model is still training.
Multi-generational fungibility matters as much as the speedup. Cognition runs Vera Rubin under the same operating model as its GB200 and GB300 fleets, having scaled from bridge capacity to thousands of GPUs in under nine months. NVIDIA Volta silicon from 2017 remains in commercial service on CoreWeave Cloud. Software consistency, as much as early allocation, is becoming CoreWeave’s moat against Nebius, Together AI, and the other neoclouds racing to deploy the same racks.
CoreWeave’s Differentiation Message Distills to Pace, Performance, and Partnership
Asked what makes CoreWeave unique, executives gave multiple and sometimes mixed answers. One pointed to a combination of scale and speed, the efficiency of building very large clusters while growing quickly, two qualities that rarely arrive together. Another emphasized the interconnectivity of the Forge services, arguing that hyperscalers offer nominally equivalent services that customers would be hard pressed to make work together even on a single cloud, while the other AI clouds do not offer the full set at all. A third described working backward from what customers want to accomplish, acting as a trusted advisor rather than quoting GPU counts, and contrasted that with a legacy paradigm of bundled cloud services. A fourth argued that Weights & Biases, with more than 1 billion logged runs and tens of thousands of researchers using its tools, gives CoreWeave direct access to the developers and researchers that enterprise sales organizations never reach, and that when procurement objects, the ML researcher prevails.
These perspectives distill to three advantages: pace, performance, and partnership with customers. Pace shows up in nine months from concept to a dozen services and in bringing Vera Rubin NVL72 to production ahead of the field. Performance has third-party support, with Signal65 finding CoreWeave delivered higher throughput and lower cost per token on NVIDIA Blackwell, and Cognition’s customer-run Vera Rubin numbers extend that evidence a generation forward. Partnership is the hardest to copy, combining direct expert engagement from the CEO to the data center floor with a product feedback loop that runs from Weights & Biases’ developer base into service definition, the loop that produced Forge itself. CoreWeave now has more differentiation claims than one positioning statement can organize and the AI loop frame is its first attempt to unify them.
The Partner Network completes the picture with validated integrations spanning VAST Data, ClickHouse, LanceDB, CrowdStrike, IBM and Red Hat, and model partners including Reflection, while Exa, Parallel Web Systems, and You.com supply a shared search layer for agents through one integration.

The competitive risk remains with the incumbents. AWS, Google Cloud, and Microsoft Azure own the enterprise data gravity and can assemble a comparable loop over time. CoreWeave’s wager is that an integrated loop beats an assembled one and that performance leadership on NVIDIA’s newest racks keeps the labs anchored while Forge recruits the enterprises. Forge editions that start free make that wager measurable within quarters.
What to Watch:
- Whether RL Rollouts moves from preview to general availability with named customers beyond the launch demo
- Whether enterprise distillation and fine-tuning customers graduate to serverless RL within the next two quarters
- Whether Forge Free and Pro signups convert into enterprise contracts over the next two quarters
- Whether AWS, Google, and Microsoft answer with managed serverless RL in SageMaker, Vertex AI, and AI Foundry
- Whether CoreWeave invests in Forge’s cross-cloud tooling at the same pace as its CoreWeave-only services
- Whether Vera CPU deployments convert into agent sandbox demand at the concurrency CoreWeave projects
Read more details about CoreWeave Forge on the company website.
Sources
- Introducing CoreWeave Forge, CoreWeave
- CoreWeave Forge: Turn AI Iteration Into Compounding Improvement, CoreWeave
Declaration of generative AI and AI-assisted technologies in the writing process: This content has been generated with the support of artificial intelligence technologies. Due to the fast pace of content creation and the continuous evolution of data and information, The Futurum Group and its analysts strive to ensure the accuracy and factual integrity of the information presented. However, the opinions and interpretations expressed in this content reflect those of the individual author/analyst. The Futurum Group makes no guarantees regarding the completeness, accuracy, or reliability of any information contained herein. Readers are encouraged to verify facts independently and consult relevant sources for further clarification.
Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.
Read the full Futurum Group Disclosure.
Other Insights From Futurum:
CoreWeave Q2 FY 2026: AI Demand Drives Pricing and Capacity Growth
From Blaize to NVIDIA: Why Is AI Infrastructure Converging on Indonesia?
Adobe Embeds CX Intelligence Into ChatGPT at OpenAI DevDay
