Google Returns to the Frontier With Gemini 4 Argon

Google Returns to the Frontier With Gemini 4 Argon

Analyst(s): Nick Patience, Fernando Montenegro
Publication Date: October 1, 2026

Google has announced Gemini 4 Argon, its first frontier model in more than seven months, aimed at software engineering, legal and finance work, and cyber defense. Early independent tests put it at the level of OpenAI’s best model and behind Anthropic’s, and access starts with vetted cyber defenders, with paying customers to follow.

What Is Covered in This Article:

  • What Google announced with Gemini 4 Argon, including its 1-million-token output limit, pricing, and a staged rollout through the Fairwind Program.
  • How Gemini 4 Argon compares with models from OpenAI, Anthropic, and leading Chinese labs on early independent tests.
  • Why Argon’s low per-token prices do not necessarily translate into lower costs per task.
  • What a defender-first release of a flagship model means for security vendors, CISOs, and buyers outside the U.S.
  • Why the bottleneck in vulnerability management is moving from finding flaws to fixing them.

The News: Google has announced Gemini 4 Argon, describing it as its new frontier model for real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense. Argon is Google’s first frontier model since Gemini 3.1 Pro, following the company’s decision to skip the previously announced Gemini 3.5 generation, and it introduces a new naming scheme.

Gemini 4 Argon is rolling out first to an initial cohort of trusted cyber defenders through Google’s Fairwind Program, ahead of developers, enterprises, and consumers. Google launched Fairwind in September 2026 to stage access to government agencies and national cyber authorities, critical-infrastructure operators, and core technology platforms. The program has more than 650 participating partners in total, and participants must limit access to internal security, incident-response, or penetration-testing teams and use multi-factor authentication. For this cohort and for Google’s internal teams, Argon ships without its usual cyber guardrails.

Google is also participating in the U.S. government’s voluntary pre-release model access process. It says broader access will follow once it has gathered feedback from early testers and strengthened its safeguards, starting with paid API customers and Google AI Ultra subscribers, although it has not given a date.

Google has raised Gemini 4 Argon’s output limit to 1 million tokens, from 64,000, and the input context window remains 1 million tokens. Introductory pricing is $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 after the introductory period, with cached input priced 95% below the standard input rate.

Google reports a state-of-the-art 77.9% on DeepSWE v1.1, a long-horizon software engineering benchmark, first place at 51.3% on Zapier’s AutomationBench, and leading results on the Vals Index, Vals Finance Agent v2, and Harvey’s Legal Agent Benchmark. On security, Google says Argon can autonomously find, validate, and patch vulnerabilities, and that it ties for first place on CWE-bench v1, a remediation benchmark, at 68%. Inside Fairwind, Argon works with CodeMender, Google’s find-and-fix harness.

Wiz, which Google acquired in March 2026, is the named early user, through its Scan for Good initiative, which protects critical public infrastructure for free. Google says Argon found a critical flaw exposing sensitive personal data in healthcare software used by hospitals worldwide. Internally, Google says teams of Argon agents have freed more than 300 TiB of memory across its data centers and are migrating C/C++ codebases to Rust at scales of up to 800,000 lines.

Before broad release, Google says it is strengthening safeguards in four areas, covering cyber and chemical, biological, radiological, and nuclear (CBRN) misuse, indirect prompt injection, monitoring of the model’s reasoning and actions for misalignment, and hardening of the sandboxed environments used for training and evaluation.

Google Returns to the Frontier With Gemini 4 Argon

Analyst Take: Frontier model competition has shifted toward long-running professional work, and Gemini 4 Argon puts Google back in that contest, although not yet at the front of it. On the early independent evidence, Argon is level with OpenAI’s best model and behind Anthropic’s, which is a strong recovery for a company that went more than seven months without a new frontier model, yet short of the lead that Google’s own benchmark table implies. We think Argon’s most interesting qualities are the ones enterprise buyers will care about more than leaderboard positions, namely its honesty about what it does not know, its strength in business automation, and a security release model that gives defenders a head start. Its weaker points are the number of tokens it consumes and the distance between Google’s model and the products built around it.

How Gemini 4 Argon Compares With OpenAI, Anthropic, and Chinese Models

The clearest independent assessment so far comes from Artificial Analysis, which puts Gemini 4 Argon at 53 on its Intelligence Index at the highest reasoning setting, level with OpenAI’s GPT-6 Astra and Claude Fable 5.1 and 23 points above Gemini 3.1 Pro. Anthropic’s Claude Opus 5.5, at 58, and Claude Sonnet 5.5, at 56, remain ahead. Argon also trails Sonnet 5.5, Opus 5.5, and GPT-6 Astra on Terminal Bench 4, a test of agentic coding in a terminal, although it takes first place on Artificial Analysis’s version of AutomationBench and tops the Vals Index at 68.9%, the first Gemini model to do so.

Argon’s hallucination rate on the AA-Omniscience test is 15%, against 51% for GPT-6 Astra, which means it is far more likely to admit it does not know an answer than to guess. However, its accuracy on the same test is 13 points below Astra’s, so some of that low hallucination rate comes from caution. For legal and finance work, where a confident wrong answer costs more than no answer, we reckon that is the right trade-off.

On the current version of the index, the strongest Chinese models, Moonshot AI’s Kimi K3 and Z.ai’s GLM-5.3, score 44 and 45, roughly eight or nine points below Gemini 4 Argon. Kimi K3’s open weights, released under a revenue-tiered license, will keep it attractive to organizations that want control over where and how they run models, but on raw capability, Argon has opened a clear gap.

Earlier this month, Bloomberg reported that some Google employees find Gemini 4 less convincing in day-to-day work than its benchmark scores suggest, particularly on coding and front-end design, a characterization Google disputes. Developers and enterprises will settle that argument, but they cannot do so until Google sets a release date.

Why Gemini 4 Argon’s Low Token Prices May Not Mean Lower Costs

Google’s introductory pricing of $2 and $10 per million input and output tokens compares with $10 and $50 for GPT-6 Astra, and the 95% cache discount suits agentic workloads, where most tokens are repeated context. However, Artificial Analysis found that Argon uses an average of 62,000 output tokens per task, against 27,000 for Astra. A task that costs $1.99 at the introductory price, about 60% of Astra’s cost, therefore rises to $3.98 once the discount ends, roughly 20% more than Astra. Google will need to improve token efficiency, or keep the discount for longer, if price is to be part of its pitch.

Argon’s 1 million token output limit raises the same question, since a single run that uses the full limit would cost about $10 in output tokens at the introductory rate and $20 at list price. The limit should help with long code migrations, legal drafting, and financial analysis by reducing hand-offs between agent steps, although buyers will need to test whether long single runs beat shorter, well-orchestrated ones on both cost and accuracy.

Defender-First Release Reaches Google’s Flagship Model

Vetted defenders have been getting early access to cyber-capable models for some time, with Google’s Gemini 3.8 Flash Cyber, OpenAI’s cyber-tuned GPT-5.x line, and Anthropic’s Mythos all reaching them before, or instead of, the general public. Gemini 4 Argon differs because it is the general-purpose model Google intends to sell to developers, enterprises, and consumers, and trusted defenders get it first, without its usual cyber guardrails. We think that sequencing is the right call, and Google deserves credit for saying so openly and for naming the safeguards it wants in place before broad availability.

Access programs such as Google’s Fairwind, Anthropic’s Glasswing, and OpenAI’s Daybreak now look much alike, with vetted cohorts, use limited to internal security teams, multi-factor authentication, and monitoring, which makes the access gate part of the product. Microsoft has taken a different route by embedding its cyber model only inside MDASH, while NVIDIA favors an open approach that lets vendors such as CrowdStrike gate models of their own.

CrowdStrike and Palo Alto Networks sit in all three lab programs, and the one defender Google names as already using Argon, Wiz, is a Google company. ETR’s October 2026 data shows these vendors already running well ahead of the security sector on spending intent. That survey predates Argon, so this is an association rather than an effect, although it suggests early access is flowing to vendors that already have buyer momentum.

Argon’s early access also runs alongside a U.S. government pre-release process, and Europe was left out of Glasswing at launch, which adds a jurisdictional dimension to these programs. We would not overstate it, since the June 2026 restriction on foreign-national access was reversed within weeks, Fairwind claims partners globally, and vetting may well be the responsible choice at this capability level. For a CISO outside the U.S., however, being inside that early-access window is now a dependency worth tracking.

Finding Vulnerabilities Is Getting Cheaper Than Fixing Them

Google pitches Gemini 4 Argon’s security value on patching as much as on discovery, with CodeMender finding, validating, and fixing flaws, and CWE-bench, the security benchmark Google leads with, measuring remediation. We think that is the right target, because discovery is no longer the scarce resource, as Glasswing partners reported more than 10,000 high- or critical-severity findings within a month.

Futurum’s 1H2026 Cybersecurity Decision-Makers Survey shows where the harder problem sits. About half of the respondents asked about exposure management (N=99) named friction in routing patch work from security teams back to IT operations and DevOps among their top three obstacles to proactive risk reduction, and no other option was cited more often.

A model that writes a verified patch in minutes does not shorten the change-advisory process, open a maintenance window, or persuade a system owner to reboot a critical server. For open-source code, the strain falls on maintainers and the disclosure process, which is the coordination problem the White House’s Gold Eagle clearinghouse was set up to handle. In our view, the meaningful measure of Argon for defenders will be the time it takes to get a fix merged rather than the number of findings, and it is fair to ask Google, and every lab running these programs, to start reporting it.

Gemini 4 Argon and Google’s Enterprise Ambitions

Google is going after the enterprise work that pays, namely software engineering, legal and finance research, and security, and starting with cyber defenders gives it time to line up reference customers before general availability. The harder task is the products built around the model. OpenAI and Anthropic have moved beyond selling models to building their own products, including coding agents and work tools such as ChatGPT Work and Claude Cowork, and we think Google’s equivalents have yet to catch up. Google will need to bring Argon into Workspace and Google Cloud quickly if benchmark wins are to become an enterprise business.

What to Watch:

  • When Google sets a general availability date, and whether Vertex AI and Gemini Enterprise customers get Argon at the same time as API customers.
  • Whether independent developer testing confirms Google’s benchmark results, or the real-world coding weaknesses described by Bloomberg’s sources.
  • Whether Google improves Argon’s token efficiency before the introductory pricing ends, since the cost per task, rather than per-token price, will decide many enterprise purchases.
  • How long does the defender window last? Argon’s guardrail-free early access gives defenders a real head start only if broad release and comparable open-weight capability take long enough to arrive.
  • Whether Google publishes merged-patch or time-to-remediation results from Fairwind partners, which would be the clearest evidence that Argon reduces risk rather than adding to the backlog.

See the complete announcement blog post on Gemini 4 Argon on the Google website.


Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.

Other Insights From Futurum:

Anthropic Glasswing: AI Vulnerability Detection Has Crossed a Threshold

OpenAI’s GPT-6 Astra: Benchmarks, Cyber Risks, and Market Impact

So This Is How AIs Attack: Observations From the OpenAI/Hugging Face Incident

Featured Image: Google

Author Information

Nick Patience is VP and Practice Lead for AI Platforms at The Futurum Group. Nick is a thought leader on AI development, deployment, and adoption - an area he has researched for 25 years. Before Futurum, Nick was a Managing Analyst with S&P Global Market Intelligence, responsible for 451 Research’s coverage of Data, AI, Analytics, Information Security, and Risk. Nick became part of S&P Global through its 2019 acquisition of 451 Research, a pioneering analyst firm that Nick co-founded in 1999. He is a sought-after speaker and advisor, known for his expertise in the drivers of AI adoption, industry use cases, and the infrastructure behind its development and deployment. Nick also spent three years as a product marketing lead at Recommind (now part of OpenText), a machine learning-driven eDiscovery software company. Nick is based in London.

Fernando Montenegro serves as the Vice President & Practice Lead for Cybersecurity & Resilience at The Futurum Group. In this role, he leads the development and execution of the Cybersecurity research agenda, working closely with the team to drive the practice's growth. His research focuses on addressing critical topics in modern cybersecurity. These include the multifaceted role of AI in cybersecurity, strategies for managing an ever-expanding attack surface, and the evolution of cybersecurity architectures toward more platform-oriented solutions.

Before joining The Futurum Group, Fernando held senior industry analyst roles at Omdia, S&P Global, and 451 Research. His career also includes diverse roles in customer support, security, IT operations, professional services, and sales engineering. He has worked with pioneering Internet Service Providers, established security vendors, and startups across North and South America.

Fernando holds a Bachelor’s degree in Computer Science from Universidade Federal do Rio Grande do Sul in Brazil and various industry certifications. Although he is originally from Brazil, he has been based in Toronto, Canada, for many years.

Related Insights
Miro MCP Hits 16M Calls: AI Collaboration Goes Cross-Functional
October 1, 2026

Miro MCP Hits 16M Calls: AI Collaboration Goes Cross-Functional

Keith Kirkpatrick, Vice President, Research, at Futurum, Miro's MCP server exceeded 16 million calls since February 2026, with usage doubling since May and cross-functional adoption now outpacing engineering roles....
Salesforce Bets on AI Simulation to Defend CRM Dominance
October 1, 2026

Salesforce Bets on AI Simulation to Defend CRM Dominance

Salesforce's acquisition of Listen Labs embeds Agentic AI directly into its cloud platforms, enabling autonomous research agents and digital twins to compress customer research cycles from months to days across...
Oracle Nexus Adds Agents to Financial Crime Investigations
October 1, 2026

Oracle Nexus Adds Agents to Financial Crime Investigations

Keith Kirkpatrick, Research Director at The Futurum Group shares insights on Oracle Nexus, human oversight in financial crime investigations, and the need to demonstrate operational results....
OpenAI Moves Up the Stack and Competes With the Platforms It Powers
October 1, 2026

OpenAI Moves Up the Stack and Competes With the Platforms It Powers

Futurum Research’s Mitch Ashley, Nick Patience, and Vikram Rothnam examine how OpenAI used DevDay 2026 to move up the stack, claiming the work surface, the agent runtime, and the pricing...
SAS Bets on Partner Expertise to Close Its AI Go-to-Market Gap
October 1, 2026

SAS Bets on Partner Expertise to Close Its AI Go-to-Market Gap

SAS names Casey McGee as Executive Vice President and Chief Sales Officer to address documented weaknesses in go-to-market execution and ecosystem alignment, positioning the company to compete in the rapidly...
Why Enterprise AI Budgets Keep Growing Despite AI Safety Pledges
September 30, 2026

Why Enterprise AI Budgets Keep Growing Despite AI Safety Pledges

ETR data suggest Anthropic and OpenAI's pacing pledges function as competitive distance-setting more than proven risk mitigation, and it hasn't cost either lab a dollar in enterprise spend. Keyphrase: frontier...

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.