AWS Serves Up NVIDIA GPUs for Short-Duration AI/ML Workloads

AWS Serves Up NVIDIA GPUs for Short-Duration AI/ML Workloads

The News: Amazon Web Services (AWS) launched Amazon Elastic Compute Cloud (EC2) Capacity Blocks for ML, a consumption model that lets customers reserve NVIDIA graphics processing units (GPUs) co-located in EC2 UltraClusters for short-duration machine learning (ML) workloads. You can read the press release on the AWS website.

AWS Serves Up NVIDIA GPUs for Short-Duration AI/ML Workloads

Analyst Take: NVIDIA has cemented its position as a leading GPU provider with its high-performance computing (HPC) and deep learning capabilities capturing significant market share, particularly among gamers, data scientists, and AI researchers. Hyperscale cloud providers are capitalizing on this demand by offering NVIDIA’s GPU-accelerated cloud instances, which cater to a wide array of workloads from complex AI modeling to graphics-intensive applications, thereby expanding access to these high-end computing resources without the upfront investment in physical hardware.

Against this backdrop, AWS has come up with a way to get around NVIDIA GPU demand issues while enabling customers to avoid making a long-term commitment to expensive GPUs to run short-term jobs. In his blog, Channy Yun, AWS principal developer advocate, compared this approach to making a hotel room reservation. The customer reserves a block of time starting and finishing on specific dates. Instead of picking a room type, the customer selects the number of instances required. When the start date arrives, the customer can access the reserved EC2 Capacity Block and launch P5 instances. At the end of the EC2 Capacity Block duration, any running instances are terminated.

The usage model provides GPU instances to train and deploy generative AI and ML models. EC2 Capacity Blocks are available for Amazon EC2 P5 instances powered by NVIDIA H100 Tensor Core GPUs in the AWS US East (Ohio) Region. The EC2 UltraClusters designed for high-performance ML workloads are interconnected with Elastic Fabric Adapter (EFA) networking for the best network performance available in EC2.

Capacity options include 1, 2, 4, 8, 16, 32, or 64 instances for up to 512 GPUs, and they can be reserved for between 1 and 14 days. EC2 Capacity Blocks can be purchased up to 8 weeks in advance. Keep in mind, EC2 Capacity Blocks cannot be modified or cancelled after purchase.

EC2 Capacity Block pricing depends on available supply and demand at the time of purchase (again, like a hotel). When a customer searches for Capacity Blocks, AWS will show the lowest-priced option to meet the specifications in the selected data range. The EC2 Capacity Block price is charged up front and will not change after purchase.

We see this usage model as a particularly good fit for organizations that need GPU for a single large language model (LLM) job and do not want to pay for long-term instances. This setup is especially valuable now with interest in generative AI peaking and GPU resources in great demand and priced at a premium.

Looking Ahead

Looking ahead, the GPU provisioning marketplace is poised for further innovation, with hyperscale cloud providers such as AWS leading the charge by offering flexible and cost-effective GPU access models akin to the EC2 Capacity Blocks. This approach not only circumvents the scarcity and high upfront costs of NVIDIA GPUs but also aligns with the growing enterprise demand for scalability and agility, especially as interest in generative AI peaks. AWS’s model, which facilitates short-term, high-intensity compute jobs without long-term commitment, is likely to become a blueprint for cloud services, offering a strategic advantage to organizations that engage in sporadic, resource-intensive tasks such as training LLMs.

Disclosure: The Futurum Group is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.

Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of The Futurum Group as a whole.

Other Insights from The Futurum Group:

AWS Storage Day 2023: AWS Tackles AI/ML, Cyber-Resiliency in the Cloud

AWS Announces New Offerings to Accelerate Gen AI Innovation

Google Cloud Set to Launch NVIDIA-Powered A3 GPU Virtual Machines

Author Information

Steven engages with the world’s largest technology brands to explore new operating models and how they drive innovation and competitive edge.

Dave focuses on the rapidly evolving integrated infrastructure and cloud storage markets.

Related Insights
RingCX Goes Agentic: Can AIR Pro Win Enterprise CCaaS?
August 26, 2026

RingCX Goes Agentic: Can AIR Pro Win Enterprise CCaaS?

RingCentral unveiled AIR Pro, an agentic AI suite designed for enterprise contact centers, featuring vertical readiness for healthcare, unified workforce engagement, and built-in compliance analytics at Customer Contact Week 2026....
Google's Vertical AI Bet Governance Matters More Than Models
August 26, 2026

Google’s Vertical AI Bet: Governance Matters More Than Models

Nick Patience, VP & Practice Lead for AI Platforms at Futurum, examines Google Cloud's new vertical AI platforms for legal and financial services, and asks whether governed connectors can finally...
nCino's Q2 FY2027: Agentic AI Banking Thesis Meets Margin Reality
August 26, 2026

nCino’s Q2 FY2027: Agentic AI Banking Thesis Meets Margin Reality

nCino's Q2 FY2027 results showcase agentic AI's banking impact: subscription revenues grew 10% to $143.5M, GAAP operating income swung to $13.6M profit, aligning with 86.6% of tech leaders prioritizing agentic...
FPT IS Bets on Vietnam's Insurance Gap With Atomi Digital MOU
August 26, 2026

FPT IS Bets on Vietnam’s Insurance Gap With Atomi Digital MOU

FPT IS and Atomi Digital partner to develop end-to-end insurance technology solutions targeting Vietnam's Insurance Gap. The MOU aims to capitalize on government initiatives to grow the sector from 1.8%...
Ransomware Hits 2026 Peak: Is Your Channel Ready for AI-Driven Attacks?
August 26, 2026

Ransomware Hits 2026 Peak: Is Your Channel Ready for AI-Driven Attacks?

NCC Group's July 2026 Threat Intelligence Report reveals 894 ransomware cases—a 22% surge and 2026 peak. The emergence of JADEPUFFER, the first fully autonomous AI attack agent, signals a critical...
AI Maps Cancer's Hidden States to Predict Winning Drug Combos
August 26, 2026

AI Maps Cancer’s Hidden States to Predict Winning Drug Combos

AI algorithms identified ultraconserved cancer cell states across patients and predicted synergistic drug combinations with ~90% accuracy, challenging assumptions about tumor heterogeneity....

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.