Hue, a two-person startup emerging from Y Combinator's latest cohort, has launched a testing platform designed to address one of the thorniest problems in deploying AI agents: the gap between controlled demos and chaotic production environments. The San Francisco company's software captures complete production runs of AI agents and converts them into repeatable tests with simulated environments that preserve the tools, data, and state from actual usage.
The timing reflects mounting frustration among developers. Engineers have been discussing the need for what they call "trace evidence" from real agent failures, not just synthetic test cases that miss the edge cases production inevitably surfaces.
Hue's approach hinges on capturing the full context of how agents interact with APIs, external tools, and data streams during live operations. The platform then reconstructs those interactions in what the company describes as "isolated simulated worlds" for each test run. According to documentation on the company's site, the agent executes in the developer's own process while a disposable simulated environment runs within Hue's infrastructure. That simulated world becomes authoritative, recording an ordered journal of every action the agent takes.
The company claims its simulated APIs and Model Context Protocol servers replicate real-world behavior down to the edge cases. Whether that holds up under the pressure of genuinely unpredictable production traffic remains to be seen, but the architecture at least attempts to solve a problem that has bedeviled teams building autonomous agents.
Rather than locking users into proprietary tracing, Hue built its system on OpenTelemetry standards. That decision allows teams to layer Hue's testing on top of existing observability stacks. Documentation shows developers running both Langfuse and Hue simultaneously through dual OpenTelemetry exporters, a flexibility that may prove critical for adoption in enterprises already committed to specific monitoring tools.
The startup shipped Python and TypeScript SDKs in mid-September, according to package registries on PyPI and npm. The Python release moved from version 0.1.0 on September 15 to version 0.6.1 ten days later, a pace suggesting active iteration. The platform includes a command-line interface with utilities for running evaluations and integrating with popular coding agents like Claude Code, Cursor, VS Code, and Windsurf.

Founder Pedigrees and Product Roots
CEO Aneesh Kethini was the second AI product manager at Datadog, where he built Bits AI, the SRE agent the observability giant announced in December 2025. Before that he worked as a quantitative trader at SIG, a background that perhaps explains the emphasis on systematic testing and reproducibility. CTO Josh Le brings engineering experience from Apple, Gusto, and the home automation company August, along with stints at Wayfair and two Y Combinator alumni, SendBlue and Sixtyfour. He was also a Kleiner Perkins fellow, a detail that signals he's run in circles where rigor matters.
The agent testing category has grown crowded. Runloop launched a benchmark orchestration platform with Weights & Biases integration last spring. Around the same time Hue was taking shape, AIVAX introduced what it called "Agentic Tests" that transform actual production jobs into regression checks. Established players like Langfuse, Traceloop, and Baserun already offer evaluation pipelines for LLM applications. OpenAI added testing utilities to its Agents SDK in late September, and an arXiv paper published that same month detailed record-and-replay techniques for converting agent incidents into CI regression tests.
Hue claims backing from investors affiliated with Y Combinator, Cursor, Clay, and Vercel, though the company has not disclosed funding amounts or valuation. The startup lists SOC 2 Type II, ISO 27001, and HIPAA compliance on its website, claims that would be unusual for such an early-stage company if verified. A link to a trust center suggests the founders are at least taking security theater seriously, whether or not they've completed full audits.

The platform is available now, with documentation hosted at docs.hue.run. Whether Hue can carve out defensible territory in a space where both incumbents and well-funded newcomers are racing to solve similar problems will depend on execution speed and whether the OpenTelemetry bet pays off in enterprise adoption. For now, the product exists, and teams wrestling with unreliable agents in production have another tool to consider.
