Palo Alto Networks Bets Frontier AI Can Make Pentesting Continuous

Palo Alto Networks Bets Frontier AI Can Make Pentesting Continuous

Analyst(s): Fernando Montenegro
Publication Date: September 28, 2026

Palo Alto Networks turned Unit 42’s frontier-model exposure analysis into an always-on subscription built on Claude Mythos 5, GPT-5.6-Cyber, and open-weight models. The service fits where exposure management is heading, but shared model access, tier economics, and the remediation last mile will decide its value.

What is Covered in this Article

  • Palo Alto Networks launched Unit 42 Continuous Frontier AI Defense, an annual service that pairs gated frontier models with open-weight models and pairs them with offensive security experts for always-on testing.
  • Unit 42’s own research makes the case for multiple models, since no single model catches more than 40% of vulnerabilities, and the two gated models overlap on fewer than 10% of findings.
  • Gated model access is shared with peers such as CrowdStrike and Google, and the model mix behind each subscription tier is a pricing lever buyers should question.
    Finding exposures is not the bottleneck; routing fixes through IT and DevOps is, and the service should be judged on exposures closed.

The News: Palo Alto Networks recently announced Unit 42 Continuous Frontier AI Defense, an always-on offensive security service sold as an annual subscription and available worldwide. Subscription options vary depending on which OpenAI, Anthropic, and open-weight models a customer uses.

The service starts with a full-estate baseline scan, then tests continuously as the environment changes. A proprietary multi-model harness routes each task to the model best suited for it, including the gated Claude Mythos 5 and GPT-5.6-Cyber. Unit 42 validates end-to-end attack paths across web apps, APIs, cloud infrastructure, code repositories, and network assets, then delivers prioritized fixes, code-level guidance, and virtual patch recommendations.

It extends the point-in-time Frontier AI Exposure Analysis Unit 42, launched in April. Palo Alto Networks cites six months of internal testing, 100+ customer engagements, and a $17M R&D investment, and says internal use produced a year’s worth of penetration testing results in three weeks. OpenAI and Anthropic provided launch quotes.

Palo Alto Networks Bets Frontier AI Can Make Pentesting Continuous

Analyst Take: A point-in-time penetration test has always been a snapshot of a moving target, and frontier models that find and chain exploits at machine speed make that snapshot age faster than most programs can schedule the next one. Palo Alto Networks is betting the answer is a service rather than a product: Unit 42 people, a proprietary harness, and gated models, sold by the year.

The direction fits where exposure management is heading, so what deserves scrutiny is more practical: who gets privileged access to these models, what the cheaper tiers give up, and whether findings turn into fixes. Buyers also appear more comfortable than the evidence suggests. In the Futurum 1H 2026 Cybersecurity Decision Maker Survey, 68% of respondents (N=929) agreed they are highly confident in detecting and defending against AI-generated attacks, while Unit 42 reports its exposure analysis found exposures in every customer it assessed.

An Honest Argument for More Than One Model

Unit 42 backs the harness with its own model research, and the findings argue against single-model hype. By its measurement, no single model catches more than 40% of vulnerabilities in a complex environment, and Claude Mythos 5 and GPT-5.6-Cyber overlap on fewer than 10% of the exposures they find. If that holds up outside Palo Alto Networks’ own testing, it makes a strong case for orchestrating across models, which is what the harness is built to do.

Packaging this as a service also fits the problem. Validation (confirming that an exposure is real, reachable, and consequential in this specific environment) is the stage of exposure management that most resists productization, and an exploit that lands is about as clean a proof as the discipline gets.

It is also a large pool to sell into. Services is the biggest segment in Futurum’s 1H 2026 Cybersecurity Market Sizing & Forecast, at $77.7B of a $336B market in 2025.

Preferential Access, or Table Stakes?

The press release leans on “exclusive access to gated capability models.” Frontier labs are extending preferential access to their most capable cyber models, today Mythos and GPT-5.6-Cyber, and whatever follows them, to large security vendors as a group. That access is exclusive relative to customers, who cannot simply buy these models, but much less so relative to peers: Palo Alto Networks was a Project Glasswing launch partner alongside CrowdStrike, Microsoft, Cisco, Google, and others, and OpenAI’s August Daybreak expansion put GPT-5.6-Cyber in the hands of 16 security vendors and consultancies.

Then there is cost. Gated models sit at the top of the price list; Anthropic, for example, lists Mythos 5 at $10 per million input tokens and $50 per million output tokens, 2.5 times its Opus 5.5 pricing, and continuous offensive testing is likely to burn through a lot of tokens. Palo Alto Networks is candid that the harness exists partly to manage “the cost of frontier AI at scale,” that open-weight models sit in the mix, and that subscription options vary with the models used.

To us, that makes model mix a pricing lever the customer can only partly see, since tiers name the models but not how the harness routes work among them. If no single model covers more than 40%, a buyer on a tier weighted toward open-weight models should ask what coverage they are giving up, and how Unit 42 would demonstrate the difference.

There is a small irony here, too. In Palo Alto Networks’ own threat narrative, open-weight models stripped of guardrails are the attacker’s tool of choice, and the same class of model now sets the cost floor of the defense.

A Crowded Field Arriving at the Same Place

CrowdStrike’s Frontier AI Readiness and Resilience Service is close to a mirror image, pairing frontier-model scanning with red team prioritization and remediation routed through its platform. Google is taking an own-model path with AI Threat Defense, combining Gemini, Wiz, CodeMender, and Mandiant.

Autonomous pentest specialists such as XBOW, Horizon3.ai, and Pentera were running continuous testing well before “frontier” became a product adjective, and consultancies including Accenture and IBM now reach the same OpenAI models through Daybreak.

The model labs are the wildcard. Anthropic and OpenAI both supplied quotes for this launch, but both distribute their cyber models through many partners at once, and it is fair to ask how long a harness stays a differentiator when the model owner can see across all of them.

Exposures Closed, Not Exposures Found

The 51% reduction in mean time to remediate comes from Palo Alto Networks’ internal deployment, where security and engineering report into the same company. Customer environments are messier. In the survey’s exposure management module, friction routing patching work from security back to IT operations and DevOps was the most-cited operational hurdle, named in the top three by 50% of respondents.

That is the gap a continuous service has to close, and it is an organizational problem more than a technical one. Always-on testing that produces more validated findings than system owners can absorb risks becoming a faster way to build a backlog.

We would also argue that continuous need not mean uniform. Testing cadence should track how fast each part of the estate changes and what it is worth, which points toward tiering by risk rather than running everything at the same tempo.

The measure that matters, in our opinion, is exposures closed rather than exposures found. The remediation guidance and the pairing with Frontier Virtual Patching point in that direction, and we would like to see Unit 42 report results to customers on those terms.

For more information, read the full announcement from Palo Alto Networks here. Read the Unit 42 blog post here.

What to Watch

  • Will customers see what they bought? Watch whether Unit 42 discloses which models ran and what coverage the cheaper, open-weight-heavy tiers give up.
  • Does gated access stop being scarce? If frontier labs keep widening access, competition shifts to harness quality, threat intelligence, and people.
  • Do the remediation numbers hold up with customers? The 51% gain is internal. Customer data, where security does not control engineering, is the real test.
  • Is this a funnel into the platform? Findings flow toward Frontier Virtual Patching and Cortex; watch how well the service serves estates built on other vendors’ controls.
  • How do pentest specialists and the channel respond? XBOW, Horizon3.ai, Pentera, and MSSPs reselling pentests face a well-funded services rival; expect price pressure or new partnerships.

Disclosure: Futurum is a research and advisory firm that engages or has engaged in research, analysis, and advisory services with many technology companies, including those mentioned in this article. The author does not hold any equity positions with any company mentioned in this article.
Analysis and opinions expressed herein are specific to the analyst individually and data and other information that might have been provided for validation, not those of Futurum as a whole.
Read the full Futurum Group Disclosure.

Other Insights from Futurum:

Why AI Learned to Attack Before It Learned to Defend

Can Frontier Virtual Patching Close the AI Exposure Gap?

Anthropic Glasswing: AI Vulnerability Detection Has Crossed a Threshold

So This Is How AIs Attack: Observations From the OpenAI/Hugging Face Incident

Author Information

Fernando Montenegro

Fernando Montenegro serves as the Vice President & Practice Lead for Cybersecurity & Resilience at The Futurum Group. In this role, he leads the development and execution of the Cybersecurity research agenda, working closely with the team to drive the practice's growth. His research focuses on addressing critical topics in modern cybersecurity. These include the multifaceted role of AI in cybersecurity, strategies for managing an ever-expanding attack surface, and the evolution of cybersecurity architectures toward more platform-oriented solutions.

Before joining The Futurum Group, Fernando held senior industry analyst roles at Omdia, S&P Global, and 451 Research. His career also includes diverse roles in customer support, security, IT operations, professional services, and sales engineering. He has worked with pioneering Internet Service Providers, established security vendors, and startups across North and South America.

Fernando holds a Bachelor’s degree in Computer Science from Universidade Federal do Rio Grande do Sul in Brazil and various industry certifications. Although he is originally from Brazil, he has been based in Toronto, Canada, for many years.

Related Insights
Cohesity Extends Cyber Resilience to the AI Agents Themselves
September 28, 2026

Cohesity Extends Cyber Resilience to the AI Agents Themselves

Fernando Montenegro, VP at Futurum, analyzes Cohesity's Agent Resilience launch at Catalyst 2026, which brings backup and rollback to AI agents and ties recovery to business continuity....
WidePoint's DHS Contract Protest: Setback or Speed Bump?
September 26, 2026

WidePoint's DHS Contract Protest: Setback or Speed Bump?

The GAO sustained TurningPoint Global Solutions' protest of WidePoint's DHS CWMS 3.0 contract award. Futurum Group analyzes whether this represents a major setback or routine federal contracting process, examining WidePoint's...
Nomios Acquires Orbcom to Plant Its Flag in Iberia
September 24, 2026

Nomios Acquires Orbcom to Plant Its Flag in Iberia

Nomios, backed by Keensight Capital, acquired Orbcom, a €15M Portuguese cybersecurity firm with 80 professionals, establishing its first Iberian presence and gaining Palo Alto Networks Diamond Innovator Partner status....
VAST DataEnclave Unifies Proprietary Models and Sensitive Enterprise Data
September 23, 2026

VAST DataEnclave Unifies Proprietary Models and Sensitive Enterprise Data

Brad Shimmin, VP and Practice Lead at Futurum, analyzes how VAST DataEnclave leverages NVIDIA Confidential Computing to manage proprietary AI models and sensitive data as unified operating system resources....
Splunk .conf26 Trust is Key for the Agentic Era
September 22, 2026

Splunk .conf26: Trust is Key for the Agentic Era

Fernando Montenegro, VP at Futurum, analyzes Splunk .conf26, where Cisco made trust and cost the gate on the agentic SOC and positioned the data and control layer beneath the agents...
AI Fuels Record Cyberattacks in Italy, and a Channel Opportunity
September 22, 2026

AI Fuels Record Cyberattacks in Italy, and a Channel Opportunity

Exprivia's Q2 2026 threat report reveals record AI-driven cyberattacks on Italian organizations, as cybersecurity becomes a top revenue driver for global channel partners....

Book a Demo

Welcome

The vision behind everything in Futurum’s Custom Research practice is this: research should show you what is happening, what comes next, and what to do about it. It should be personal to each audience, easy for people to grasp, and structured so LLMs can reason over it accurately. And it should be fast and turnkey; you want answers now, not another project to carry for quarters.

Whether you are defining business, channel, or go-to-market strategy; evaluating vendors or justifying ROI; or commissioning research to fill an emerging market need, we have your back, with a program that answers your questions with the objectivity and credibility to drive real decisions.

To do it, we bring unmatched data to bear: Futurum research, surveys, and market projections; validated market feeds; ETR’s 15 years of insight from 10,000 technology decision-makers; G2’s buyer and user data; and what our analysts hear every day. Add leading primary collection, from AI-moderated voice interviews to surveys and analyst-led interviews, all turnkey, and every project comes out credible, nuanced, and actionable.

And we don’t just drop the results in your lap. For internal work, we provide analyst-led sessions, interactive dashboards, and a range of formats. For market-facing work, Futurum delivers turnkey activation and amplification that actually gets seen, by people and by LLMs, through our media and share of voice. This is research that moves decisions and markets.

We will meet you wherever you are, from a fast-turn brief to a multi-year program, and shape the work to your goals, timeline, and budget. The right program for your moment.

If any of this is useful, I would love to talk.

Benjamin Brown, VP Custom Research, Futurum Research

Benjamin Brown

VP, Custom Research · The Futurum Group

Newsletter Sign-up Form

Get important insights straight to your inbox, receive first looks at eBooks, exclusive event invitations, custom content, and more. We promise not to spam you or sell your name to anyone. You can always unsubscribe at any time.

All fields are required






Thank you, we received your request, a member of our team will be in contact with you.