AMD’s MLPerf Inference 6.1 results turn ROCm software velocity into a measurable quantity. GPT-OSS-120B throughput rose up to 38% on unchanged MI355X hardware in one benchmark cycle, and the six-week release cadence AMD committed to in July 2026 gives that improvement a public schedule going forward. Futurum argues the rate ROCm sustains from this baseline now matters more than any single hardware comparison.
What Is Covered in This Article:
- ROCm round-over-round gains in MLPerf Inference 6.1 on fixed hardware, up to 38% on GPT-OSS-120B and 70% on Wan 2.2 single stream
- AMD’s six-week ROCm release cadence, announced at Advancing AI 2026 and enabled by the TheRock build system
- Why the cadence plus MLPerf makes ROCm velocity auditable from the 6.1 baseline forward
- AI-assisted kernel development inside AMD’s software organization
- Measurement caveats, including 6.0 baselines in NVIDIA comparisons, the MI455X gap, and a possible MLCommons shift to MLPerf Endpoints
The News: On September 16, 2026, AMD published its MLPerf Inference 6.1 results, its broadest submission to date, covering six models across Instinct MI355X, MI350X, and the MI350P PCIe card launched in May 2026. Every headline comparison in the round rests on software rather than new silicon. The same MI355X GPU that AMD submitted in round 6.0 delivered up to 38% more GPT-OSS-120B throughput at 8 GPUs, Wan 2.2 single-stream performance improved 70%, and 72 GPUs in 6.1 outperformed 94 GPUs in 6.0 on multi-node GPT-OSS-120B. In the closed division, an 8-GPU MI355X system led an 8-GPU NVIDIA B200 submission on GPT-OSS-120B by 33% in offline and 26% in server, and led an 8-GPU B300 submission by 14% and 13%. Crusoe submitted GPT-OSS-120B and DeepSeek-R1 on 512 MI355X GPUs, the largest GPU count in MLPerf Inference history, and seven partners landed on average within 4% of AMD’s own numbers.
The round arrives two months into a structural change in how ROCm ships. At Advancing AI 2026 in July, AMD moved ROCm from a roughly quarterly cycle to a fixed six-week release cadence, made possible by TheRock, the automated open-source build system that reached production with ROCm 7.14 on July 16. ROCm 10.0 followed on August 27 as the first release under the new cadence, introducing ROCm.AI with the ROCm CLI, AMD Skills for coding agents, and Hyperloom, an agentic system that profiles and optimizes inference workloads. AMD’s release materials claim an average 3.3x inference improvement over ROCm 7 on 8 MI355X GPUs.
AMD confirmed the GPT-OSS-120B runs used the MXFP4 model checkpoint as released by OpenAI with an FP8 KV cache, and committed to publishing step-by-step reproduction instructions. AMD did not submit MI455X results, positioning the new platform to participate in the upcoming MLPerf Endpoints, a rolling-submission format that measures inference under serving conditions.
AMD MLPerf Inference 6.1 Results Show ROCm Gaining 38% on the Same MI355X Hardware
Analyst Take: The most consequential quantity in AMD’s MLPerf Inference 6.1 submission is a rate of change. Between rounds 6.0 and 6.1, ROCm added up to 38% GPT-OSS-120B throughput and 70% Wan 2.2 single-stream throughput on hardware that did not change. For two years, the case against AMD Instinct centered on ROCm maturity, a quality argument that resisted measurement. A committed release calendar plus a public benchmark converts that argument into a number, and starting from the 6.1 baseline, the number can be checked every round by anyone who reads the submissions.
Futurum’s view is that sustained software velocity is the one variable that converts competitive CDNA4 silicon into procurement wins. The 6.1 round supplies the first clean data point. A customer who deployed MI355X for round 6.0 workloads holds 23% more effective multi-node capacity today without touching the hardware, and that capacity arrived through package updates on a published schedule.
The caveats are real and Futurum details them below. Several NVIDIA comparisons rest on round 6.0 baselines, the 8-GPU B200 figure for GPT-OSS-120B is a CoreWeave partner result, MI455X is absent, and one benchmark cycle of large gains says little about whether the trajectory continues. None of that changes what buyers should track. ROCm velocity is now observable on a schedule, and the burden of proof has shifted from AMD’s claims to AMD’s calendar.
MLPerf Rounds Now Function as a Public Audit of ROCm Velocity
Every result isolates software as the only moving variable. The 8-GPU GPT-OSS-120B configuration, the Wan 2.2 scenarios, and the multi-node GPT-OSS-120B runs all reused the MI355X silicon from round 6.0, with 288 GB of HBM3E and 8 TB/s of bandwidth per GPU held constant. The 38% and 70% gains are therefore pure software deltas, verified through MLCommons peer review rather than a vendor blog. Seven partners including Dell, Crusoe, MangoBoost, and MiTAC landed within 4% of AMD’s numbers on average, and some exceeded AMD by up to 2%. That spread is tight enough to treat AMD’s submissions as a planning baseline. It also means that a buyer who wants to verify the round-over-round gains can rerun the recipe on its own systems once AMD publishes the reproduction instructions it has committed to release.
The Wan 2.2 progression shows what a single cadence interval can contain. In 6.0, AMD optimized only the single-stream scenario, failed to complete an official offline run, and landed in the open division at 87% to 88% of B300. One round later, both scenarios are closed-division results at 111% and 118% of B300. First-time enablement to closed-division leadership within one cycle has historically been a trajectory NVIDIA’s software organization owned. AMD demonstrating it on a text-to-video workload, where kernel maturity is thin across the industry, suggests the velocity extends beyond well-worn LLM benchmarks.
AMD’s own release materials assert a larger figure over a longer window: an average 3.3x inference improvement from ROCm 7 to ROCm 10.0, measured internally on 8 MI355X GPUs. Futurum weights the MLPerf-measured gains far more heavily. The 3.3x claim spans cherry-pickable workloads and configurations that AMD selected, while the MLPerf deltas were measured on fixed configurations, submitted under closed-division rules, and reproduced by seven parties. The benchmark figure is the auditable floor.
The Six-Week Cadence Converts Software Maturity from a Narrative into a Schedule
The cadence commitment is the structural change beneath the benchmark result. ROCm shipped four feature releases in 2024 and moved at a roughly quarterly pace through mid-2026, with production and preview streams that diverged confusingly enough that version numbers jumped from 7.2 to 7.14. TheRock, the automated build and release system that reached production in July, collapsed that fragmentation into a single pipeline with public release candidates and nightly builds, and it is the mechanism that makes eight to nine releases per year credible. ROCm 10.0 on August 27 was the first release delivered on the new clock. The next arrives in early October, and each subsequent MLPerf round will span a countable number of releases.
CUDA’s advantage has never been a single benchmark lead but the confidence that performance on any new model arrives quickly and predictably. A fixed cadence, audited each round against a public benchmark, gives AMD its first structural answer to that confidence. If the next rounds show ROCm continuing to add meaningful throughput on the workloads buyers run, the total cost calculation for an MI355X deployment must include a software appreciation curve, and Futurum expects sophisticated buyers to start writing that curve into capacity models.
The risk cuts the other way with equal force. Software optimization on a fixed architecture eventually meets diminishing returns, and the easiest 38% is the first 38%. NVIDIA’s MLPerf history shows round-over-round software gains that decay toward the low single digits as workloads mature, which is visible in the Llama 2 70B results this round, where MI355X and B300 tie in offline and server because both vendors have exhausted the accessible optimizations. Futurum will treat round 6.2, or the first MLPerf Endpoints submissions, as the first true measurement of AMD’s per-release rate and the test of whether frontier-workload gains stay well clear of the mature-workload pattern.
AI-Written Kernels Are the Engine Behind the Rate, and the Hardest Part to Verify
The mechanism producing the velocity became clearer in a July analyst session with Vamsi Boppana, AMD’s SVP of AI, which Futurum attended. Boppana described a software organization in which AI agents now write substantial optimization code against ROCm’s open surface area, with a kernel-generation tool inside the ROCm.AI stack producing and tuning kernels under test harnesses that check accuracy at both the kernel and end-to-end model level. He was candid that agents will satisfy a poorly specified objective in clever ways, such as dropping to lower precision to hit a throughput target, so AMD constrains intermediate steps and enforces accuracy gates across standard evaluation suites. The engineering claim, in effect, is that AMD has industrialized the kernel optimization work that CUDA accumulated over 15 years of human effort, and that the six-week cadence is the shipping schedule for that industrial output.
Boppana also described the portability consequence. He recounted that an Anthropic engineer spun up Claude on the MI350 series and completed the port over a weekend, with the model having already absorbed the ISA and programmability details from AMD’s public documentation. The openness strategy and the agent strategy are the same strategy: everything AMD publishes becomes training surface for the agents that then write AMD’s own optimization code and its customers’ porting code. This creates a plausible flywheel and a verification burden. AI-generated kernels raise reasonable questions about correctness and benchmark-specific tuning, which is precisely why AMD’s disclosure of MXFP4 weights with an FP8 KV cache, the standard high-performance configuration for GPT-OSS-120B rather than an aggressive quantization, and its commitment to published reproduction instructions carry so much weight this round. AMD is preempting the tuning accusation before skeptics raise it.
The Measurement Has Gaps, and MLPerf Endpoints Would Change the Ruler
The 6.1 baseline carries known distortions. Several of AMD’s NVIDIA comparisons use round 6.0 baselines because NVIDIA did not resubmit the workload, including the Llama 2 70B interactive lead and the DLRM v3 results, where AMD was the only accelerator vendor this round. The 8-GPU B200 baseline for GPT-OSS-120B is a CoreWeave partner submission rather than NVIDIA’s own. The cleanest comparisons are GPT-OSS-120B against NVIDIA’s official 8-GPU B300 submission, where MI355X leads by 13% to 14%, and the 72-GPU GB200 result, where MI355X leads by 18% offline and 7% server with 95% scale-out efficiency.
MI455X is the larger gap. AMD attributes the absence to the July 31 deadline, which is credible for silicon introduced this year, but it means the velocity measurement currently covers only the CDNA4 generation, and gains on a new architecture typically start high and cannot be inferred from a mature one. MLCommons may replace fixed rounds with MLPerf Endpoints, a rolling-submission format measuring throughput, latency, concurrency, and interactivity under serving conditions. That transition would improve the realism of the measurement and complicate the arithmetic, since rolling submissions remove the clean round-over-round intervals that make a per-release rate easy to compute. Futurum’s expectation is that a rolling format ultimately favors AMD’s cadence, because a vendor shipping every six weeks benefits from a benchmark that accepts results every week rather than every six months.
What to Watch:
- Whether the next MLPerf round shows ROCm gains on GPT-OSS-120B continuing at a strong per-release pace or flattening toward the mature-workload pattern visible in Llama 2 70B
- Whether AMD ships its October and November ROCm releases on the six-week schedule without slippage
- Whether NVIDIA resubmits GPT-OSS-120B, Llama 2 70B interactive, and DLRM v3 in the next round, removing the 6.0 baseline caveats
- Whether MLCommons formally adopts MLPerf Endpoints, letting AMD publish MI455X results ahead of a spring 2027 round
- Whether third parties reproduce the 6.1 results from AMD’s published ROCm instructions and stay within the 4% partner spread
Read the complete details here.
Sources
- Call for Submission: Edge Agentic Inference Benchmark for MLPerf Inference v6.1, MLCommons, September 2026
- AMD Delivers Breakthrough MLPerf Inference 6.0 Results, AMD
- AMD Delivers Breakthrough MLPerf Training 6.0 Results, AMD
- AMD Delivers Breakthrough MLPerf Inference 6.0 Results, Reddit
- AMD Instinct GPU MLPerf Inference results, Principledtechnologies
- Breaking Down AMD’s MLPerf Training 6.0 Results, Tensorwave
- Technical Dive into AMD’s MLPerf Inference v5.1 Submission, AMD
- AMD’s MLPerf Inference 6.0 Results Show Strong …, Linkedin
- AMD Instinct MI355X Achieves MLPerf Inference v6.0 Gains with Over 1 Million Tokens per Second and Supports Scalable ROCm Stack, Storagereview
- MLPerf Inference v6.0: Dell Showcases Breakthrough Performance with AMD Instinct™ MI355X GPUs, Delltechnologies
- State of the Market Report: Semiconductors, Supply Chain, and Emerging Technology, Q3 2026
Declaration of generative AI and AI-assisted technologies in the writing process: This content has been generated with the support of artificial intelligence technologies. Due to the fast pace of content creation and the continuous evolution of data and information, The Futurum Group and its analysts strive to ensure the accuracy and factual integrity of the information presented. However, the opinions and interpretations expressed in this content reflect those of the individual author/analyst. The Futurum Group makes no guarantees regarding the completeness, accuracy, or reliability of any information contained herein. Readers are encouraged to verify facts independently and consult relevant sources for further clarification.
Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.
Read the full Futurum Group Disclosure.
Other Insights From Futurum:
Cerebras CS-4 Makes the Rack the New Chip by Doubling Power and Tripling Wafers
AMD Acquires Taalas to Advance AI Workload Optimization
AMD Q2 FY 2026: EPYC and Helios Fuel the Next AI Growth Phase
Author Information
Brendan is Research Director, Semiconductors, Supply Chain, and Emerging Tech. He advises clients on strategic initiatives and leads the Futurum Semiconductors Practice. He is an experienced tech industry analyst who has guided tech leaders in identifying market opportunities spanning edge processors, generative AI applications, and hyperscale data centers.
Before joining Futurum, Brendan consulted with global AI leaders and served as a Senior Analyst in Emerging Technology Research at PitchBook. At PitchBook, he developed market intelligence tools for AI, highlighted by one of the industry’s most comprehensive AI semiconductor market landscapes encompassing both public and private companies. He has advised Fortune 100 tech giants, growth-stage innovators, global investors, and leading market research firms. Before PitchBook, he led research teams in tech investment banking and market research.
Brendan is based in Seattle, Washington. He has a Bachelor of Arts Degree from Amherst College.

