The chatbot's politeness should have been the first warning sign. Ask nicely, and it would share confidential employee records—names, salaries, performance reviews—with anyone curious enough to try. Elsewhere, an AI coding assistant that had cut development time by 40% was committing API keys to public repositories, poisoned by a cleverly corrupted Stack Overflow answer it had ingested without question.
These aren't thought experiments. Security teams found both vulnerabilities in production systems during early 2026, according to interviews with companies building automated red-teaming tools. And they're hardly alone.
As enterprises unleash thousands of AI agents—some approved, many decidedly not—a new industry is scrambling to hack these systems first. The task is more daunting than it sounds. Traditional security testing falters against AI's probabilistic nature, while the agents themselves can query databases, dispatch emails, and authorize purchases on behalf of users. An April 2026 survey found that 82% of enterprises might have AI agents they are unaware of. Sixty-five percent reported having AI-agent-related security incidents in the past year.
The agents being deployed today will soon handle progressively higher-stakes decisions—approving loans, scheduling surgeries, configuring cloud infrastructure. The race between those hunting vulnerabilities and those exploiting them shows no signs of slowing.
Governance Lagging, Badly
The AI agent security problem arrived faster than most organizations anticipated, perhaps faster than they were willing to admit. Gartner projects a 47% increase in AI spending, reaching $2.59 trillion in 2026. Information security budgets are projected to reach $244 billion in 2026, driven in part by AI and generative AI concerns.
But governance? That's lagged badly behind deployment. Deloitte reported that around 80% of organizations lack mature agent governance frameworks—a polite way of saying most companies have no idea what their AI is doing. The Cloud Security Alliance survey was blunter: 54% of enterprises reported between 1 and 100 unsanctioned AI agents operating within their environments. These rogue agents, often deployed by individual teams without security review, can access sensitive data, interact with external systems, and make decisions affecting customers and employees.
Recent incidents underscored the stakes. Researchers disclosed EchoLeak (CVE-2025-32711) in late 2025, a zero-click indirect prompt injection vulnerability in Microsoft 365 Copilot that enabled unauthenticated data exfiltration. In January 2026, the "Reprompt" attack demonstrated how attackers could pilfer data from Microsoft Copilot through carefully crafted prompt injections. By August, security researchers were warning of a "Copilot worm" spreading through Microsoft Word documents via cross-domain prompt injection attacks hidden in document formatting—essentially turning office documents into vectors for AI exploitation.
Why the Red-Teaming Market Is Accelerating Now
Three forces are converging: regulatory pressure, proven attack vectors, and a dawning recognition that traditional testing approaches simply don't work for AI.
The regulatory drumbeat has grown louder. The EU AI Act's Article 55 requirements for general-purpose AI systems with systemic risk went into enforcement on August 2, 2026, mandating "conducting and documenting adversarial testing" using standardized protocols. Providers can adhere to the GPAI Code of Practice to demonstrate compliance until harmonized standards emerge, but the obligation remains clear. Across the Atlantic, the Office of Management and Budget's M-24-10 memo from March 2026 established governance and risk management requirements for federal AI use, explicitly defining AI red-teaming and directing agencies on risk practices.
The attack surface, meanwhile, keeps expanding in ways that would have seemed exotic just a few years ago. Microsoft's Cyber Pulse report from February 2026 outlined foundational governance areas for agents: visibility, least privilege, real-time monitoring, interoperability, built-in protections. But academic research published throughout 2026 demonstrates that vulnerabilities extend well beyond simple prompt injection. Researchers have documented successful attacks via RAG poisoning, tool metadata manipulation, memory drift, over-privileged tool access, and goal hijacking—terms that sound like science fiction but represent very real exploit paths. A June 2026 paper, "Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation," catalogued the expanding threat landscape facing multi-step agentic systems.
Perhaps most importantly, security leaders are realizing that static security evaluations are fundamentally broken for systems that adapt and reason.
"Static security evals are fundamentally broken for AI agents," Ibrahim Abdu, co-founder of Fabraix, wrote on LinkedIn in July 2026. Traditional penetration testing assumes you can enumerate vulnerabilities in a fixed codebase—you find the bug, you patch it. AI agents generate novel responses to novel inputs, interact with changing external data sources, and operate across extended conversational contexts. The old playbook doesn't apply.
Consolidation and Emergence

The vendor landscape reflects both opportunity and the inevitable consolidation pressure of a hot market. Gartner's June 2026 report on "Emerging Tech: Agent and LLM Security Spend" noted rapid growth but predicted significant vendor landscape churn as larger platforms absorb red-teaming capabilities.
That consolidation is already visible, and moving fast.
OpenAI announced its acquisition of Promptfoo on March 9, 2026, bringing the open-source and enterprise red-teaming tool's CLI-based automated testing and evaluation frameworks directly into OpenAI's ecosystem. Zscaler acquired SPLX around November 2025, rebranding its capabilities as "AI Red Teaming" and integrating them into the Zero Trust Exchange. Check Point bought Lakera in September 2025, combining its capabilities into "AI Guardrails" and "AI Red Teaming" offerings.
Smaller, specialized players are carving out niches as well. Fabraix, a two-person team from Y Combinator's Summer 2026 batch, describes their Nyx system as "the world's frontier hacker for AI agents"—a claim that's bold even by startup standards. The automated red-teaming agent reportedly found vulnerabilities in agents at dozens of Fortune 500 companies and achieves a 78% attack success rate on the AgentHarm benchmark versus 67% for GPT-5.6 Sol, according to the company's July 2026 launch post. The team—Ahmed Aly, former Monzo payments fraud lead, and Ibrahim Abdu, an ex-Meta engineer who built AI debugging agents—incorporated Fabraix Ltd in London on January 30, 2026.
Fabraix's approach exemplifies the new generation of automated testing. Nyx maintains more than 10,000 jailbreak and attack strategies, uses adversarial self-play to evolve tactics, and tests across chat, voice, browser, and coding agents without requiring source code access. The company introduced an "Adversarial Cost to Exploit" (ACE) benchmark, measuring the dollar and token cost required to breach an AI agent—framing security as an economic problem rather than a binary pass-fail assessment. It's a clever reframe that resonates with CFOs and security teams alike.
Mindgard, a Lancaster University spinout, positions its DAST-AI platform for continuous red-teaming across prompts, guardrails, and tool integrations. Customer case studies from January 2026 describe mission-critical AI applications in highly regulated enterprises undergoing attacker-aligned testing. CalypsoAI markets "Agentic Warfare" capabilities—the marketing arms race is in full swing. HiddenLayer released a red-teaming module in April 2026. Each vendor emphasizes continuous, automated testing over point-in-time manual reviews, recognizing that AI agents require monitoring, not just auditing.
Larger platforms, for their part, are bundling red-teaming with runtime guardrails and governance tools. Zscaler's partnership with OpenAI, announced April 15, 2026, leverages OpenAI models within Zscaler's red-teaming workflows—the company has used OpenAI models for AI red-teaming since 2024. Check Point's June 2026 Agent Security Early Access program includes integrations with Google Cloud's Gemini Enterprise Agent Platform.
OpenAI itself disclosed "GPT-Red" in mid-July 2026, an internal automated red-teaming system used to adversarially train GPT-5.6. The company devoted significant compute to safety self-improvement, positioning automated red-teaming as a core component of responsible AI development rather than just an external security practice. When the company building the most capable AI models in the world starts treating adversarial testing as infrastructure, others tend to follow.
What Comes Next
Several shifts seem likely, though predicting the arc of an industry this young carries its own risks.
First, cost-based and distributional security metrics will likely gain traction over simple pass-fail evaluations. Academic papers published in July 2026 criticized single-seed benchmark scores and advocated for reporting attack success rate distributions and economic cost-to-exploit. Fabraix's ACE framework and similar approaches may reshape how organizations evaluate and procure agent security tools—or at least how vendors pitch them.
Second, the definition of "red-teaming" for agents will expand beyond jailbreak prompts to encompass tool poisoning, memory manipulation, and multi-turn attack chains. Research published throughout 2026 demonstrates that real-world agent vulnerabilities often involve indirect attacks on retrieval systems, tool metadata, or conversation history rather than direct prompt injection. NIST's March 2026 analysis of a large-scale agent red-teaming competition highlighted these multi-layer challenges, which are harder to defend against precisely because they're harder to conceptualize.
Third, defense-in-depth will become the norm. No single control stops agent attacks—a reality that's simultaneously sobering and liberating for security teams. Microsoft's February 2026 security guidance, NIST frameworks, and vendor materials all emphasize layered defenses: input monitoring, guardrails, automated red-teaming, least-privilege tool access, runtime containment. Organizations that treat red-teaming as a checkbox rather than one layer in a comprehensive security architecture will remain vulnerable, full stop.
Fourth, regulatory compliance will accelerate formalized testing. With EU enforcement of GPAI adversarial testing requirements beginning August 2, 2026, model providers and enterprises deploying high-impact agents face pressure to demonstrate documented, standardized red-teaming workflows. U.S. federal agencies already operate under OMB guidance requiring AI risk management. This won't eliminate vendor hyperbole about "frontier" capabilities, but it should push the industry toward more rigorous, reproducible testing practices. Should being the operative word.
The Stakes Are Rising

The question facing CTOs and security leaders isn't whether to red-team AI agents. That debate is over. The question is how to do it continuously, comprehensively, and credibly—before attackers write the headlines instead of appearing in case studies.
The agents handling customer service inquiries today will be approving credit lines tomorrow, configuring production infrastructure the day after. Each capability expansion multiplies the potential damage from a successful exploit. Security teams accustomed to patching known vulnerabilities in deterministic code now face systems that generate novel behaviors in response to adversarial inputs they've never seen before.
It's a fundamentally different problem. And the companies racing to solve it—whether two-person startups or AI giants—are betting that enterprises will pay handsomely for answers before their agents become liabilities.
