The promise sounds straightforward enough: autonomous software that handles customer service, manages infrastructure, even ships code without human handholding. Eighty-four percent of enterprise leaders told Zapier in a January survey they plan to increase AI agent investments this year. The catch? Most of those initiatives never escape the pilot phase.
It's not a compute problem. It's a trust problem. Specifically, nobody wants to hand autonomous systems the keys to production databases, billing platforms, or supply chains without knowing—really knowing—what those systems will do when the unexpected happens.
Which is why a race you probably haven't noticed is gathering speed. While the industry argues over which large language model benchmarks highest, a parallel infrastructure build-out is underway: constructing the digital gyms where AI agents train before they touch anything that matters.
Polymath, a two-person team out of Y Combinator's Winter 2026 batch, just raised $8 million to scale exactly that. The seed round, led by Base10 and Cervin Ventures on April 16, pulled in SurgePoint Capital, Founders Future, Berkeley Frontier Fund, Y Combinator itself, and a handful of former YC partners. It's the kind of raise that signals investor conviction that simulation infrastructure isn't a nice-to-have—it's foundational.
The timing tracks. Gartner forecast last August that 40 percent of enterprise applications would feature task-specific AI agents by the end of 2026, up from barely 5 percent in 2025. That's not incremental growth. That's an order-of-magnitude leap. But surveys from ITPro in January show many of those projects stuck at proof-of-concept, gated by safety and reliability concerns that nobody's quite figured out how to address.
The gap between what agents can theoretically accomplish and what enterprises will actually allow them to do remains uncomfortably wide. Simulation infrastructure is emerging as the bridge—perhaps the only bridge—across that divide.
When Benchmarks Meet Reality
The shift from generative AI to agentic systems has moved from speculation to engineering challenge. METR's March 2025 analysis of long-horizon task completion found that the length of tasks agents can reliably complete with 50 percent accuracy has doubled roughly every seven months over the past six years. If that trend holds—and METR's ongoing TH1.1 updates suggest it is holding—models are approaching the ability to autonomously handle workflows that span days, not minutes.
BCG weighed in with its own forecast in February, projecting up to $200 billion in net-new demand for tech services driven by agentic AI scaling beyond pilots. KPMG's Q4 AI Pulse survey, released January 15, found that half of executives plan to spend between $10 million and $50 million over the next year just to productionize agents. The capital is there. The appetite is there.
What's missing is confidence.
The problem is deceptively simple: you can't beta-test an AI agent on live customer data. You can't let it learn by breaking billing systems or experimenting with supply chain dashboards. And synthetic benchmarks built for code generation—can this model write a function that sorts a list?—don't translate to real-world operational complexity. If agents are going to complete tasks that span hours or days, deploying software or managing infrastructure or handling customer inquiries across half a dozen interconnected systems, they need environments that mirror the messiness of production. Without the consequences.
Regulation, Standards, and the MCP Effect
Several forces are converging to make simulation infrastructure not just useful but necessary. Start with regulation. The EU AI Act's obligations for General-Purpose AI models and high-risk systems take effect August 2, 2026. The Act's FAQ makes clear that AI agents qualify as AI systems and may be classified as high-risk depending on their use case, which creates a compliance requirement for pre-deployment testing. NIST's AI Risk Management Framework is moving in a similar direction; an April 7 concept note outlined a Critical Infrastructure Profile that would formalize risk controls for agents operating in sensitive contexts.
Then there's technology. NVIDIA used its March GTC conference to deepen its simulation stack for what it calls "physical AI." The company released Omniverse DSX Blueprint to general availability, announced partnerships with FANUC, ABB, YASKAWA, and KUKA integrating Omniverse and Isaac, and previewed Isaac Sim 6.0 for long-running agent workflows. Siemens Energy integrated Isaac Sim and Metropolis into its Noedra grid-health digital twin. KION and GXO are using Omniverse-powered warehouse twins to train and test fleets of autonomous forklifts before those machines ever touch a real pallet.
The industrial playbook—train in simulation, validate in digital twins, deploy to hardware—is now being adapted for software agents. It's an obvious parallel, though software carries its own complications.
The Model Context Protocol is accelerating this adaptation. MCP emerged as the de facto tool-integration layer in 2026, with extensive vendor adoption following. Industry trackers estimate more than 10,000 live MCP servers and roughly 97 million SDK downloads per month as of early this year. ExpressVPN shipped an MCP server in March. Anthropic, OpenAI, and Gemini all support MCP clients. That standardization means simulation platforms can plug into the same tool ecosystems agents will use in production, which makes test environments significantly more realistic.
It also introduces new attack surface. OX Security identified potential vulnerabilities requiring patches across MCP SDKs, prompting updates across the ecosystem. Which underscores, perhaps more clearly than anything else, the need for pre-deployment hardening. If the tooling layer itself is vulnerable, letting agents loose without rigorous testing becomes reckless.
A Crowded, Fast-Moving Category

The simulation-for-agents category is crowded and moving fast. Salesforce announced "eVerse" on March 27, a simulated environment designed to expose edge cases and handoff failures for customer experience agents before they interact with real customers. LivePerson launched Syntrix on March 3, positioning it as an agent evaluation and live-agent training platform that uses simulated customer personas and policy adherence testing. Virtue AI introduced Agent ForgingGround on March 17, offering more than 50 production-grade simulated enterprise environments—Gmail, ServiceNow, Atlassian—for security testing and red-teaming of agents.
Centific went further on March 12 with "Environments-as-a-Service," simulated enterprises for reinforcement-learning-trained agents. The pitch: spin up a realistic corporate IT environment, complete with tickets, workflows, approval chains, and let your agent learn by doing without risking anything real. Parallel Web Systems, which raised a $100 million Series B at a $2 billion valuation on April 29, is building web infrastructure—indexing, APIs—specifically for agents.
The infrastructure layer is thickening.
Polymath's angle is production realism at the benchmark level. The company's Horizon-SWE benchmark, released February 2 and updated February 6, evaluates end-to-end software engineering agents in what it describes as a "production-grade, tool-rich environment." The benchmark includes 73 tools, a live application, synthetic traffic, and verifiers that check not just code correctness but DevOps reliability and engineering quality. Leading models score around 25.5 percent on strict pass/fail. A partial-credit composite metric shows the leader at 60.4 percent. Horizon-SWE uses MCP for tool access, mirroring how agents operate in the wild.
Founders Dylan Ma, who previously worked at Hume AI and AWS and holds an NSF Fellowship, and Naren Yenuganti, who spent time at Plaid and Amazon, position Polymath as building the "environment layer" for agents. The YC launch post from around March describes the company's work as creating "simulated worlds where AI agents learn to operate autonomously over long horizons," combining running applications, real tools, and multi-step tasks.
The vision echoes Applied Intuition, the autonomous-vehicle simulation company that raised a $600 million Series F at a $6 billion valuation in June 2024. If Applied Intuition is the gym for self-driving cars, Polymath wants to be the gym for software agents. It's an ambitious parallel, though software presents fundamentally different failure modes.
The Benchmark Wars
Polymath is hardly alone in rethinking how to evaluate agents. February and March saw a wave of new benchmarks designed for long-horizon, interactive tasks. LongCLI-Bench arrived February 15. AMA-Bench on February 26. ATBench, focused on long-horizon agent safety, appeared April 2. The HORIZON diagnostic suite launched April 13. WebArena, a browser-based task benchmark, introduced a "Verified" variant in February to address scoring variance and make results more deterministic. BrowserGym, another web navigation benchmark, received an active update in March.
The pattern is clear: benchmarks are converging toward longer task horizons, verifiable outcomes, and safety diagnostics. The "Verified" movement—deterministic scoring with partial-credit taxonomies—aims to close the gap between evaluation and production performance. Polymath's Horizon-SWE fits this trend. By measuring feature correctness, DevOps reliability, and engineering quality in addition to code generation, it reflects the multi-dimensional nature of real software work.
This matters because enterprises don't care if an agent can write a function. They care if it can ship a feature, monitor its deployment, handle rollback if something breaks, and document the process. That's what Horizon-SWE tries to capture. The 25.5 percent pass rate for the leading model suggests we're still early—very early. But the benchmark's structure—stateful systems, 73 tools, deployment and monitoring—provides a template for what production-ready evaluation might look like.
What Comes Next

The simulation layer for AI agents will likely consolidate around a few themes. Regulatory pressure will favor companies that can demonstrate pre-deployment testing rigor. The EU AI Act's August 2, 2026 deadline and NIST's evolving framework both push in that direction. Enterprises will want audit trails showing that agents were tested in realistic environments before going live. Simulation platforms that integrate with regulatory sandboxes—the EU AI Act makes sandboxes operational from August 2, 2026—will have a structural advantage.
Standardization around MCP and similar protocols will make it easier to build high-fidelity simulations. If agents use the same tool interfaces in training as in production, the fidelity gap narrows. But standardization also introduces monoculture risk. The vulnerabilities in MCP SDKs identified earlier this year are a reminder that widely adopted protocols become attractive targets. Security will be a recurring tension.
The market will likely split between horizontal platforms—companies like Polymath building general-purpose simulation and evaluation infrastructure—and vertical players focused on specific domains. Salesforce's eVerse is for CX agents. NVIDIA's stack is for physical AI and industrial twins. Centific is betting on enterprise IT workflows. Whether the horizontal or vertical approach wins may depend on how quickly enterprises move from single-use-case agents to multi-domain orchestration. That shift feels inevitable, but timing is uncertain.
What's certain is that simulation and evaluation infrastructure will be table stakes for agent deployment at scale. The difference between a chatbot and an autonomous agent is operational risk. Chatbots can hallucinate; agents can delete databases, bill customers incorrectly, or make decisions with regulatory consequences. The industry is realizing, somewhat belatedly, that you can't train agents on production and hope for the best.
Polymath's $8 million seed, announced the same week that Parallel Web Systems raised $100 million and Virtue AI launched its ForgingGround platform, signals that investors see this infrastructure layer as foundational. The race isn't just to build smarter models anymore. It's to build the environments where those models can safely learn to do real work.
And if METR's trend line holds—task horizons doubling every seven months—the need for those environments is arriving faster than most enterprises are ready for.
