Last May, when the Cloud Security Alliance quietly documented threat actors deploying autonomous AI frameworks for what it called "persistent exploit operations," the news landed with a thud among security professionals who'd been bracing for exactly this moment. Not surprise, exactly. More like confirmation of something they'd been watching unfold in real time.
The shift had been building for months: traditional vulnerability scanners—the kind that flag thousands of potential issues and leave security teams drowning in alerts—were giving way to something far more unsettling. AI-powered red teams that don't just identify vulnerabilities but chain attack steps together, write working exploit code, and validate that patches actually work. They think, in other words, more like a human penetration tester than a search algorithm.
The distinction matters. Traditional scanners generate noise. Autonomous systems generate proof.
---
Consider the numbers, directional as they are. The penetration testing market stood somewhere between $1.82 billion and $1.94 billion in 2023, depending on which analyst you ask—figures drawn from 2024 reports that may not reflect current market conditions. By 2030, forecasts from those same reports point toward $4.57 billion to $5.24 billion—roughly 16.6% compound annual growth. Adjacent markets are moving even faster: Breach and Attack Simulation hit $1.65 billion in 2026 and is projected to nearly double by 2032.
These figures, subject to the usual analyst caveats and the passage of time since their publication, still tell a story about where enterprise budgets are flowing. Validation over detection. Proof over possibility.
Several open-source frameworks have emerged as testing grounds for this approach, and their adoption trajectories hint at broader industry appetite. Strix, which went public in November 2025, accumulated 37,500 GitHub stars within months—a metric that means something in developer circles, even if it's an imperfect proxy for actual usage. The project claims to orchestrate multi-agent penetration testing: enumerating attack surfaces, chaining exploitation steps, producing proof-of-concept code. By last April, Strix reported more than 80,000 users running north of 1,300 pentests daily.
Those numbers, it should be said, are company-reported and should be taken with caution until independently verified.
PentAGI, another open-source contender, picked up somewhere between 17,000 and 18,000 stars in the spring and early summer of last year. Other frameworks followed: Pentest-Swarm-AI, RedAmon (focused specifically on large language model security), T3MP3ST. Each positions specialized agents to coordinate reconnaissance, exploitation, and reporting—mimicking, more or less, how a human red team operates.
Commercial platforms have moved in parallel, though with less transparency about their methods. Horizon3.ai's NodeZero reported performing over 100,000 pentests as of February 2025, the company said, posting 102% year-over-year ARR growth. Pentera achieved $100 million in annual recurring revenue as of January 2026 and launched "Pentera Peer," an AI interface for its adversarial exposure validation platform, two months later. NetSPI rolled out AI-powered continuous pentesting services in May. Assail launched its "Ares" autonomous red team product—focused on application-layer attacks—in late March.
What distinguishes these tools from traditional scanners isn't just performance. It's architecture. A Help Net Security article from November 2025 described the new systems as exhibiting "operator-like reasoning." They don't stop at flagging a possible SQL injection. They attempt the injection, capture the database response, document the exploit chain, then retest after remediation to confirm the fix held.
One security engineer, speaking on condition of anonymity because they weren't authorized to discuss internal tools, put it more bluntly: "The old scanners told you there might be a problem. These things prove there is one."
---
Three forces are converging to accelerate this shift, and none of them are particularly comforting.
First: the attacker advantage is automating faster than most enterprises realize. That Cloud Security Alliance research from May documented AI-assisted zero-day patterns and frameworks like Strix appearing in adversary toolkits alongside temporal knowledge graphs for maintaining persistent access. Separately, reports surfaced last September of HexStrike-AI, a tool weaponizing multiple Citrix vulnerabilities at speed. If attackers are automating end-to-end exploitation—and the evidence suggests they are—defenders running signature-based scans are fighting with yesterday's playbook.
Second: enterprise AI adoption has outpaced governance by a comfortable margin. According to the World Economic Forum's 2026 report released last May, 77% of organizations are using AI in cybersecurity operations, with 94% of cyber leaders viewing AI as a defining competitive force. Yet a Cloud Security Alliance survey from April revealed that 82% of enterprises harbor unknown "shadow" AI agents in their environments. Sixty-five percent experienced AI agent-related incidents in the previous twelve months. An EY survey from March showed 99% of leaders expect AI to transform proactive and defensive strategies, but only 26% have fully rolled out governance frameworks.

The gap is uncomfortable. ETR research from May found 68% of respondents rate AI agents as 4 or 5 out of 5 in importance to cybersecurity's future. Governance and control-plane approaches? Still unsettled. Gartner named agentic AI oversight a top cyber trend for last year in February. Forrester predicted a public breach caused by agentic AI and warned that most enterprises weren't operationally ready.
They may still not be.
Third: regulatory pressure is beginning to formalize what "AI red teaming" actually means, though the definitions remain fluid. The EU AI Act's main applicability date arrived last August, with draft guidance from May explicitly referencing penetration testing and red teaming for high-risk AI systems. In the United States, NIST issued a concept note in April for a critical infrastructure AI profile and launched an AI Agent Standards Initiative in February. The UK's AI Security Institute continued publishing pre-deployment evaluation methodologies through mid-year. DHS guidance situates AI red teaming under Executive Order 14110, and CISA's Secure-by-Design updates from May increasingly reference continuous validation.
Regulatory frameworks don't yet mandate specific testing depth. But the direction is clear enough: compliance teams will need to show not just that they scanned for vulnerabilities, but that they validated exploitability and retested fixes. That's a higher bar than most organizations currently meet.
---
Strix offers a window into how the open-source model is reshaping this space, for better or worse. The framework runs locally or in CI/CD pipelines, integrates with the Caido proxy for precise HTTP request manipulation (announced last June), and exposes a REST API for pentest lifecycle automation. An April update introduced "context-aware pentesting" with a persistent memory layer, allowing agents to retain knowledge across test runs—an architectural choice that mirrors how human testers build intuition about a target.
The project's blog posts frame agents as "new users" and emphasize validated exploitability over raw vulnerability counts. An April post titled "Open Source Isn't Dead" positioned the dual open-source and platform model: framework remains free, enterprise layer adds scheduling, state management, auto-fix capabilities, compliance integrations.
Whether the numbers hold—80,000 users, 15 billion tokens processed daily, 78,000 vulnerabilities reported since launch—remains unverified externally. What's less ambiguous is adoption pattern. Anecdotal reports on forums like Reddit from July describe developers wiring Strix into staging environments and bug-bounty workflows, treating it as a complement to tools like Nuclei or PentestGPT.
Horizon3.ai's NodeZero represents the commercial autonomous pentesting lineage. Over 100,000 pentests conducted, the company said in February 2025, and it continues publishing research on safe, controllable autonomy. Pentera's $100 million ARR milestone and March launch of "Pentera Peer"—an AI copilot for adversarial exposure validation—signal enterprise willingness to pay for platforms that go beyond scanning.
Bugcrowd, HackerOne, and Cobalt, historically crowd-sourced pentesting platforms, are introducing AI-augmented PTaaS offerings. Yet a Cobalt report covered last June found fewer than 10% of cybersecurity professionals trust AI-only testing tools. Over three-quarters reported their AI vulnerability scanners missed critical flaws. The same research noted that AI-related findings carried a higher high-risk rate—32%—but lower resolution rates, suggesting the technology excels at finding edge cases but struggles with actionable remediation guidance.
Microsoft's PyRIT and NVIDIA-backed garak (version 0.14 released last February) represent open-source toolchains focused on LLM and agent security, mapping findings to frameworks like the OWASP LLM Top 10 and MITRE ATLAS. RedAmon aggregates these tools for reproducible LLM red teaming. Academic frameworks like OpenRT and BlackIce continue exploring agent-based adversarial testing at the research frontier.
The Cloud Security Alliance's May documentation of Strix among frameworks observed in attacker toolkits introduces an uncomfortable wrinkle: dual-use is inherent. The same multi-agent coordination that helps defenders validate security posture can help adversaries scale exploitation.
One researcher, speaking at a closed-door security conference last fall, described it as "handing both sides the same weapon and hoping good intentions carry the day." The comment drew nervous laughter.
---
The industry is converging on a model where "validated exploitability" becomes the baseline expectation—perhaps faster than many anticipated. Exposure validation platforms, pentesting-as-a-service, and AI agent security tools are circling the same territory: evidence-first reporting that demonstrates not just that a vulnerability exists, but that it can be exploited and subsequently fixed.

Governance will determine which projects survive. Gartner's Hype Cycle for Agentic AI, published last May, flagged early-stage need for security, governance, and cost controls. Forrester's predictions warned that a portion of AI spending would be deferred as enterprises grappled with operationalizing agents. The ETR survey from May found control-plane approaches still fragmented, with no dominant orchestration model emerging.
Regulatory timelines will accelerate formalization. The EU AI Act's August applicability—with extended deadlines for certain high-risk systems stretching into 2027 and 2028—creates a compliance timetable that favors platforms demonstrating continuous testing and retest loops. NIST's critical infrastructure AI profile and the UK AISI's pre-deployment methodologies point toward similar expectations in other jurisdictions.
Skepticism persists, and probably should. The Cobalt report's finding that fewer than one in ten professionals trust AI-only testing reflects pragmatic wariness: automation can scale testing, but judgment—deciding which vulnerabilities matter, understanding business context, weighing exploitability against operational impact—remains a human domain.
For now.
What's less debatable is the direction. Attackers are deploying autonomous frameworks for reconnaissance and exploitation. The Cloud Security Alliance documented it. Defenders running weekly Nessus scans are increasingly outmatched. The shift from signature-based flagging to proof-of-exploit validation isn't aspirational anymore—it's underway, driven by open-source frameworks that iterate faster than procurement cycles, commercial platforms crossing nine-figure ARR, and regulatory regimes beginning to codify what "AI security testing" actually means.
The question isn't whether autonomous red teams will reshape penetration testing. It's whether enterprises can govern them before the first high-profile breach proves they should have. And whether, when that breach arrives, anyone will be particularly surprised.
