The problem, as Shubham Palriwala and Parth Ajmera kept discovering, wasn't that their evaluations were broken. It was that evals only catch the failures you're already looking for.
Users would abandon conversations mid-stream. No error logs, no stack traces—just silence. AI agents would technically complete their tasks while quietly violating company policies in ways no one had thought to test. Feature requests would bubble up organically in chat transcripts, only to vanish into the noise before product teams even knew they existed.
So the two engineers built Agnost AI, an analytics platform designed to surface the kinds of agent failures that slip past traditional testing frameworks. The San Francisco startup emerged from Y Combinator's accelerator program and claimed to process upward of a million events daily as of mid-2024, working with engineering teams at companies including Google, Exa, and Corgi.
Whether the market needs another observability tool—this one focused specifically on conversational AI—remains to be seen. But Agnost's pitch addresses a real tension in how companies evaluate AI agents: the difference between measuring what you expected might go wrong and discovering what actually did.
Built on OpenTelemetry, for Better or Worse
At its core, Agnost ingests conversation data through OpenTelemetry, the increasingly ubiquitous open standard for observability in distributed systems. Teams already instrumenting their agent frameworks can redirect their existing OTLP exporters to Agnost's hosted collector without installing yet another SDK—a pragmatic choice that lowers the barrier to adoption but also ties the platform's fate to OpenTelemetry's trajectory.
The system supports the usual suspects: Agno, Anthropic, CrewAI, LangChain, Vercel AI SDK, OpenAI SDK, OpenAI Agents SDK, plus voice providers like Telnyx and Trellus. Google's MCP Toolbox for Databases, an open-source project, even explicitly documents OTLP export to Agnost in its telemetry setup. Yuan Teoh, a software engineer at Google involved with that project, noted the team "integrated comprehensive observability features" using the platform—though the comment stops short of a full-throated endorsement.
What Gets Measured (and What Gets Missed)

The platform organizes conversations into two primary data structures: Intents and Violations. Intents try to capture what users were attempting to accomplish, even when—perhaps especially when—the agent never understood or completed the task. Violations flag moments where an agent broke defined policies: confirming identity before exposing account details, reciting the cancellation policy, escalating safety concerns.
Ishan Goswami at Exa found it "useful in tracking analytics and error rates," according to testimonials on Agnost's site. The same page highlights quantitative wins captured in snapshots from mid-2024: 1,247 feature requests surfaced for a company called Odysser, 16 out of 18 autonomous pull requests merged for something called Lopus AI. Whether those numbers still reflect current usage or have since been surpassed isn't entirely clear.
That autonomous PR capability deserves scrutiny. When Agnost identifies a pattern of failures or violations, it can theoretically investigate the issue, gather context from the repository, prepare a code change, and open a pull request against prompts, tool definitions, or test harnesses. The documentation—wisely—cautions teams to treat these generated diffs like any engineering contribution, complete with proper testing and security review. Still, the idea of an observability tool writing its own fixes feels either prescient or overambitious, depending on your tolerance for automation.
The Pricing Math and Competitive Landscape
Agnost offers a free tier for up to 1,000 events monthly with seven-day retention—enough to kick the tires, probably not enough to run a production service for long. The Starter plan costs $49 a month for 10,000 events and 30-day retention. Pro tier jumps to $499 monthly for up to a million events, 90-day retention, and direct Slack access to the founders (a detail that signals either intimacy or scale constraints). Enterprise customers get the usual: self-hosted VPC deployments, custom SLAs, tailored workflows.
Setup purportedly takes minutes. Teams can install an "Agnost AI skill" via npm or configure environment variables to export OpenTelemetry spans directly. Events appear in the dashboard within seconds, assuming the integrations cooperate.
But Agnost enters a crowded field. Tools like Langfuse, LangSmith, Helicone, and Arize Phoenix already serve teams building with large language models, and industry comparison matrices throughout last year focused heavily on tracing, evaluation management, cost tracking, and latency monitoring. Reddit threads from around the time of Agnost's launch noted Y Combinator backing multiple agent observability startups while attempting to distinguish where Agnost fits.
The founders position their tool as product analytics rather than pure observability. Where most platforms target the regression problem—did the model produce the expected output given known inputs?—Agnost focuses on murkier questions: Where do users get stuck? What do they actually want? Why do they drop off without saying why?
The People Behind It

Palriwala and Ajmera bring complementary technical backgrounds, if not deep experience in building venture-backed companies. Palriwala was reportedly the youngest engineer on Cisco's analytics team, became the first hire at Formbricks, and contributed to projects at Bitcoin, OWASP, and the Linux Foundation. Ajmera studied computer science at IIT Madras, built Spark data pipelines at Microsoft, and led GPU technology work at Infurnia.
The company raised $250,000 from Entrepreneurs First and Transpose Platform in May 2024, following earlier pre-seed backing referenced in December 2023. The timeline suggests an extended period between initial funding and Y Combinator, though the exact sequence isn't entirely transparent.
A series of blog posts from the founders between late spring and summer last year detail their thinking: detecting churn signals before users articulate dissatisfaction, establishing post-launch improvement cadences, analyzing coding agents differently than customer service bots, measuring quality at human handoff points. The writing reflects genuine thought about product analytics for AI-native applications, even if some of the ideas remain aspirational.
The platform has achieved SOC 2 Type 1 compliance and is working toward Type 2, which remained in progress as of mid-2024, according to its trust center. The company lists 25 policies and 35 controls but no subprocessors, suggesting data remains entirely within Agnost's infrastructure—a selling point for security-conscious teams, assuming the claim holds up under scrutiny.
An Open Question
Whether the market will ultimately support both traditional LLM observability platforms and conversation-specific analytics tools isn't yet clear. For teams shipping agents that hold genuinely open-ended conversations with users—customer service bots, coding assistants, research tools—the distinction between testing known failure modes and discovering unknown ones may prove meaningful.
Or perhaps not. Evals validate specific behaviors you thought to check. Analytics surfaces the behaviors you didn't know to test. That sounds compelling in a pitch deck, but whether it's a venture-scale insight or a feature that eventually gets absorbed into existing platforms is still an open bet.
For now, Agnost occupies a niche: small enough to move quickly, focused enough to avoid competing directly with well-funded incumbents, and early enough that the rules of AI observability haven't fully hardened. What happens next depends less on the technology—which seems sound, if not revolutionary—and more on whether companies building conversational agents decide they need this particular lens on their production systems.
The founders, at least, seem convinced. The rest of the market is still making up its mind.
