It took Judgment Labs all of four months to close two funding rounds—$32 million total, both led by Lightspeed Venture Partners, both announced on the same day, May 12, 2026.
That kind of velocity isn't unheard of in San Francisco, particularly not for a startup tackling the sort of problem that keeps CTOs up at night: What happens when your AI agent goes rogue in production?
The answer, according to Judgment's three twentysomething founders, is that you need infrastructure to catch it. Think error monitoring, but for autonomous systems that can book flights, write code, or approve expense reports without asking permission first.
The Observability Gap
Judgment positions itself as "Sentry for agents," borrowing the mental model of a tool most developers already know. Where Sentry alerts engineers when traditional software crashes, Judgment watches what AI agents are doing—tracking reasoning traces, tool calls, retries, the digital breadcrumbs that reveal whether an agent is behaving or spiraling into expensive nonsense.
The platform runs real-time evaluations using LLM-as-judge scorers or bespoke rules written by engineers. When something breaks—or looks like it might—alerts fire and the problematic interaction gets routed into a dataset for retraining. There's pattern clustering to group similar misbehaviors, plus A/B testing for models and prompts. Documentation leans heavily on OpenTelemetry-based tracing, with an emphasis on server-side evaluations that won't add latency. SOC 2 Type II compliance is claimed, along with self-hosting options for the compliance-obsessed.
It's not a revolutionary pitch. But maybe that's the point.
Young Founders, Old Playbook

CEO Alex Shan is 22, a Stanford NLP Group alum who studied under Chris Manning—who also happens to be an angel investor in the company. Chief Scientist Andrew Li, 23, came from Together AI, where he was an early research hire. CTO Joseph Camyre, also 23, spent time at Datadog as a systems engineer, which likely explains why the company thinks like a monitoring shop rather than a research lab.
The trio released an open-source SDK called "judgeval" late last year, with the last tagged release (v0.23.3) on December 1, 2025. LinkedIn pegs the team at somewhere between 11 and 50 employees as of April.
Founded in 2025—yes, barely a year ago—Judgment has now raised $32 million across consecutive seed and Series A rounds. The Information reported the combined financing valued the startup at roughly $175 million post-money, a figure the company has declined to confirm publicly. That's a steep climb for a company that could still measure its age in months, but it tracks with the frothy investor sentiment around anything that makes deploying AI agents less terrifying.
Regulatory Winds (and Timing)

The fundraise came just days after U.S. and Five Eyes intelligence agencies issued joint guidance calling for continuous monitoring and traceability of AI agents. Coincidence? Perhaps. But the regulatory pressure is real, and it's created sudden urgency around tools that can demonstrate compliance and flag dangerous behavior before it reaches end users.
Judgment is hardly alone in chasing this. Braintrust raised $80 million in February at an $800 million valuation. LangSmith, Arize's Phoenix, and Patronus all offer variations on agent evaluation and monitoring. Even security vendors like Cyberhaven are bolting "agent observability" features onto their existing products, trying to answer the question enterprises are asking louder by the quarter: What are these autonomous systems actually doing inside our infrastructure?
Lightspeed partner James Alcorn said in the announcement that the firm is betting on Judgment's ability to "productize agentic evaluations" for companies deploying agents in production. The quote has the slight flatness of investor-speak, but the underlying thesis is straightforward enough—whoever wins the observability layer for AI agents will likely do so by signing marquee customers first, fast.
The Customer Silence

Judgment says its platform is live at "a growing list of agent-native companies," though it hasn't published a customer roster. The press release included a testimonial from Aqil Naeem, CEO at E3 Group, praising the platform's ability to pinpoint failure points and measure performance lifts. Beyond that, public proof points remain thin.
That's not unusual for an early-stage infrastructure play, but it does raise the obvious question: How many companies are actually deploying agents at the scale that would justify purpose-built monitoring infrastructure? The market might be nascent, or it might be arriving faster than anyone expected.
Either way, Judgment now has $32 million and a crowded field to navigate. The money will presumably go toward expanding the team, building out enterprise features that check procurement boxes, and racing to sign the kind of reference customers that turn a credible pitch into an inevitable category leader.
Whether the founders can move as fast as their fundraising calendar remains to be seen.
