Arga Labs, a San Francisco startup barely out of its Y Combinator cohort, has closed a $10 million seed round to build what amounts to safe playgrounds for artificial intelligence agents—digital twins of the enterprise software stacks they're meant to navigate.
General Catalyst led the financing announced Tuesday, with participation from Box Group, Emergence, Gradient, and SV Angel. The capital arrives as a wave of companies discover that AI agents dazzling in controlled demos often stumble when set loose on real customer data and production integrations—a gap Arga aims to fill.
The company spins up what it calls "stateful copies" of third-party SaaS platforms: Stripe, Slack, Google Workspace, Salesforce, Jira, GitHub. Developers can train autonomous agents in these replicas, which mirror real APIs and command-line interfaces without touching live systems. Think of it as a flight simulator for software bots.
"Can the agent correctly identify that these two are the same company?" CEO Phillip Li asked TechCrunch, describing the kind of edge case that trips up agents in production but gets missed in traditional testing.
Li's question underscores a broader challenge. Enterprises rushing to deploy AI agents face a testing conundrum that existing tools haven't solved. Agents that need to operate across multiple systems—customer support bots querying HubSpot and Slack simultaneously, or sales assistants reading from Salesforce while writing to Linear—operate in messy, interconnected environments where mocks and static test data fall short.
Arga's platform distinguishes itself from stateless API mocks by maintaining state across interactions. Each sandbox captures every call an agent makes, along with responses, latency spikes, and side effects. The system pulls context from project management and monitoring tools like Jira, Linear, GitHub, Sentry, and PostHog, feeding that information into agent-generated test suites that, in the company's phrasing, "know what to validate and why." Results post back as GitHub pull-request checks, slotting into familiar developer workflows.

"Having a repeatable sandbox environment is very important," Yuri Sagalov, managing director at General Catalyst, told TechCrunch—a statement that reflects venture capital's current bet on infrastructure for AI agents rather than the agents themselves.
Arga emerged from Y Combinator's most recent batch and drew attention from investors at the accelerator's Demo Day. Li recently posted on LinkedIn that customers have run north of 100,000 tests across Arga's sandboxes over a 16-week stretch, though the company hasn't published independent case studies. Its website lists logos from Weave, Rho, Slash, and Y Combinator, though these are company-provided claims not independently verified as customer endorsements.
The founders bring résumés from Amazon and Stripe. Li built an internal developer tool at Amazon that reportedly saves "10+ recurring weeks per year" of engineering time, according to the company's Y Combinator profile. He studied cognitive systems at the University of British Columbia and competed on Canada's Junior National Fencing Team—an unexpected biographical detail that mirrors the startup world's penchant for recruiting polymaths.

Co-founder and CTO Akira Tong worked as a software engineer at Stripe and a quantitative analyst at Goldman Sachs. He skipped high school, graduated university at 19, and played Identity V professionally. The company's team has grown from four people at its Y Combinator launch to six by late summer—a lean operation for a startup promising to replicate entire software ecosystems.
Arga competes in a crowded field. Langfuse and AgentOps have raised capital recently, while Vijil secured $17 million in late 2025, all targeting various facets of agent testing and observability. What sets Arga apart, in theory, is its focus on replicating stateful third-party services rather than concentrating solely on tracing or security. The company's documentation shows support for MCP server configuration export and local agent config merging, technical features aimed at developers who want to test agents with the same rigor they apply to traditional software.
The seed funding will support hiring and infrastructure expansion as Arga scales its library of service twins and the orchestration layer underneath. The startup plans to support enterprise customers running agents that span dozens of third-party integrations, a use case that seems inevitable as companies push AI assistants deeper into their operational stacks.

Whether Arga's approach proves essential or merely clever engineering remains to be seen. But the underlying premise—that AI agents need specialized testing environments that traditional mocks can't provide—reflects a maturing understanding of what it actually takes to ship autonomous software into production. The honeymoon phase of agent hype, it seems, is giving way to the harder work of making them reliable.
