When RunSybil closed a $40 million funding round from Khosla Ventures in March 2026, the pitch sounded almost reckless: autonomous agents that continuously probe live production systems, hunting for vulnerabilities, with no human operator watching over their shoulder. Two years ago, established security vendors would have dismissed the concept as fantasy. Today, the company counts Cursor, Notion, Baseten, Turbopuffer, and Thinking Machines among its customers.
That roster tells you something. These aren't enterprises moving cautiously through procurement committees. They're AI-forward startups moving at a pace that makes traditional security tooling look quaint. And they represent the leading edge of a broader, quieter transformation playing out across developer infrastructure—one where the tools built for yesterday's software are being replaced wholesale by systems designed from first principles for AI workloads.
This isn't the familiar story of incumbents slapping "AI-powered" labels onto existing products. It's something more fundamental. When your application fires off dozens of LLM calls per user request, when autonomous agents are orchestrating multi-step workflows without human intervention, when observability means tracking not just latency but token costs and reasoning chains—well, the old tools start breaking in ways their creators never anticipated.
What the Data Actually Shows
The observability market offers perhaps the clearest window into this shift. Grafana Labs' 2026 survey—fielded between October 2025 and January 2026 with responses from 1,363 practitioners—turned up some striking numbers. Ninety-two percent see value in AI surfacing anomalies. Seventy-seven percent support autonomous actions, though nearly a quarter remain skeptical. More revealing, perhaps: OpenTelemetry adoption has hit 57% for metrics, 50% for traces, 48% for logs.
That OpenTelemetry uptake matters because it signals something bigger than tool preferences. The project now includes semantic conventions specifically for generative AI—attributes like gen_ai.provider.name, token counts, cost fields that simply didn't exist in traditional observability schemas. MLflow's GenAI tracing, with updates running through April, now supports over 40 one-line autolog integrations spanning agent frameworks from LangChain to Google's Agent Development Kit. Arize Phoenix, an open-source effort, added ATIF (Agent Trajectory Interchange Format) ingestion this April, converting agent trajectories into OpenTelemetry span trees.
Gartner weighed in with a forecast on March 30, predicting that by 2028, half of GenAI deployments will include dedicated LLM observability investments—up from 15% in 2026. The GenAI models market itself is projected to surpass $25 billion in 2026, reaching $75 billion by 2029. IDC's March analysis described "rapid shifts driven by agentic AI architectures" reshaping vendor strategies.
The most concrete signal, though? Palo Alto Networks agreed in November 2025 to acquire Chronosphere for $3.35 billion, explicitly framing the deal as marrying observability "built to scale for the data demands of the AI era" with Cortex for autonomous remediation. The transaction is expected to close in the second half of fiscal 2026.
Three Converging Pressures
Three forces are accelerating this infrastructure transition, and they're colliding in 2026 with perhaps more velocity than anyone expected.
Regulatory deadlines, for one. The EU AI Act's general obligations take effect August 2, imposing logging and oversight requirements on high-risk AI systems. OMB Memorandum M-24-10, issued back in March 2024, already mandates minimum risk practices and inventories for safety- and rights-impacting AI across federal agencies. These aren't abstract compliance exercises. They demand the ability to retain agent traces, document risks, demonstrate human oversight—capabilities that retrofitted observability tools handle awkwardly at best.
Model capabilities are outpacing defensive tooling, too. OpenAI released GPT-5.4 in March, followed by GPT-5.4-Cyber in mid-April—a "cyber-permissive" variant reserved for vetted security professionals. Anthropic launched Project Glasswing on April 7, using its Mythos model for vulnerability discovery through a limited preview program. Early testing, reported by Axios on April 21, shows these models accelerating exploit development when given large token budgets. The UK AI Security Institute ran tests with 100 million token allocations. One executive framed it to Axios as the shift from "human speed vs machine speed." Which is a polite way of saying the offense is pulling ahead.
Third, standardization is happening faster than the typical enterprise technology cycle would suggest. OpenTelemetry's GenAI semantic conventions provide a common language that multiple vendors now support: MLflow, Phoenix, Langfuse (which added OpenTelemetry ingestion in February 2025), LangSmith. Harrison Chase, LangChain's CEO, wrote in a February 12 blog post that "agent observability should work no matter how you build." He noted that many LangSmith customers don't use LangChain's open-source framework but rely on the observability platform for traces and evaluations. Companies like Clay, Harvey, and Vanta are using production agent traces as a "system of record" for documenting agent logic—treating observability as infrastructure, not instrumentation.
The YC Lens

Y Combinator's Winter 2026 batch—Demo Day was March 24—offers a useful snapshot of how the next wave of founders thinks about AI infrastructure.
Sentrial, covered on March 12, positions itself as "AI agent failure detection" with drop-in observability for LangChain. Terminal Use, announced the same day, describes itself as "Vercel for filesystem-based agents" with full observability baked into the deployment layer. Chamber, launching March 16, tackles GPU infrastructure orchestration with an "AI teammate" that includes observability for cluster health. Sonarly, announced in mid-February, deploys AI agents specifically for production alert triage and fixes.
These aren't companies adding AI features to existing products. They're building observability into the fabric of AI-native deployment and orchestration. The distinction matters.
Established vendors are adapting, naturally, but their approach reveals underlying architectural assumptions. Datadog announced integration with Google's Agent Development Kit on February 6 and lists "Distributed AI Observability" and GPU monitoring in investor materials from February 12. Honeycomb announced AI features and MCP (Model Context Protocol) integrations on March 11. Grafana's survey found that 95% of respondents want AI to "show its reasoning," and 26% cite "too much manual context input" as a barrier. Which suggests the challenge isn't just capturing data but contextualizing it for AI workloads in a way that makes sense without constant human shepherding.
Security's Starker Divergence
The security side shows an even more pronounced split. MarketsandMarkets projects the Penetration Testing as a Service (PTaaS) market at $0.72 billion this year, growing to $1.98 billion by 2031—a 22.6% compound annual growth rate. But the competitive dynamics are fracturing in interesting ways.
Incumbents like HackerOne introduced "Agentic PTaaS" on January 26, emphasizing continuous testing with expert verification. Synack's Sara platform, detailed in April materials, stresses human-validated AI pentesting, claiming 64% of enterprises prefer that approach. These companies are automating workflows while keeping humans in critical loops. It's a sensible position, arguably the conservative one.
RunSybil represents a different wager: fully autonomous continuous pentesting with no humans involved in execution. The company's $40 million raise in March and its customer roster suggest early traction for that model. Assail emerged from stealth on January 13 with Ares—autonomous agents for API, web, and mobile pentesting—going generally available February 1. Both position themselves as infrastructure purpose-built for continuous validation in an AI-era development velocity that traditional pentesting cycles can't match.
The question hanging over this market is whether enterprises will trust fully autonomous agents probing production systems. Or whether the human-in-the-loop model proves more durable than its champions expect.
What Comes Next

The next eighteen months should clarify which architectural bets pay off. Two threads stand out.
First, the observability-security convergence will likely accelerate. Palo Alto Networks' Chronosphere acquisition signals that direction clearly enough. Microsoft open-sourced an Agent Governance Toolkit on April 2, referencing OWASP's Agentic Top 10 risks and emphasizing runtime policy enforcement—effectively merging observability with guardrails. Expect more deals pairing security validation with AI-native monitoring as companies realize these aren't separate problems but two sides of the same coin.
Second, the gap between "AI-assisted" and "AI-native" will widen. Grafana's survey showed 77% support for autonomous actions, but vendor messaging in 2026 has shifted from "assistive insights" to autonomy and remediation. The question isn't whether AI will triage alerts or detect anomalies—that's table stakes now. It's whether the underlying infrastructure was designed for agentic workflows or retrofitted from traditional architectures. The latter approach can work, for a while. Until it can't.
Several research papers published in early 2026 highlight observability gaps that matter to this transition. An arXiv paper from March 27 argues that "output-only feedback fails in complex chains" and calls for intermediate state observability. Another from April 6 contends that current tools like OpenTelemetry and Langfuse treat governance as downstream rather than runtime enforcement. These aren't academic exercises, exactly. They're pointing to architectural decisions that determine whether your observability stack can handle agents making real-time decisions or whether it's just logging after the fact. There's a difference.
CTOs and engineering leaders building AI-powered products face a choice that rhymes with the cloud migration decisions of the 2010s: adapt existing tools or adopt AI-native infrastructure. The companies placing early bets—Cursor using RunSybil for continuous pentesting, Harvey and Vanta using LangSmith for agent observability—suggest the transition is already underway. By Gartner's timeline, half of GenAI deployments will include dedicated LLM observability by 2028.
Which means the infrastructure shift isn't coming. It's here. The only question is how long it takes everyone else to notice.
