Agnost AI has spent the past months studying a peculiar problem: the artificial intelligence agents companies deploy in production often break in ways no one notices. A tool call returns nothing, a hallucinated link gets passed along, behavior drifts just enough that users give up without filing a bug report. Traditional infrastructure monitoring catches none of it.
Now the San Francisco startup, which emerged from Y Combinator's Summer 2026 accelerator program, says it's built a solution that goes beyond simply logging those failures. The company announced on August 20, 2026 that it can convert production agent logs into custom-trained language models purpose-built for specific tasks. In testing against Anthropic's Claude Opus 4.8, Agnost's first workload-specific model achieved 88% task success compared to 72% for the frontier model, the company reported. Median latency dropped 90%; inference costs per completed task fell 95%.
Those figures come from an evaluation across 780 held-out customer traces, which Agnost used to validate what CEO Shubham Palriwala describes as a fundamentally different approach to improving agent reliability.
"Your agent logs aren't just for debugging. They're your training data," Palriwala wrote in the company's launch post. It's a pitch aimed squarely at enterprises running conversational agents on narrow, repetitive workflows where consistency matters more than creative reasoning.
Building Models From What Went Wrong
The technical approach diverges from standard fine-tuning. Rather than adapting an existing large language model, Agnost generates evaluation datasets by parsing historical production traces and then trains small, workload-specific models from scratch. The company ingests agent traces through an SDK, OpenTelemetry integration, or its Model Context Protocol connector, according to documentation on the company's site.
What Agnost calls "silent failures" are the instances infrastructure teams typically miss: an agent that appears to finish a task but skips required steps, returns incorrect information, or subtly frustrates users in ways that don't trigger alerts. The platform clusters these failures by intent, then constructs training datasets that isolate where and why things broke down.
The company's first deployment—an identifier extraction agent for an unnamed customer—jumped from 71.5% task success to 87.9%. Median latency collapsed from 4.8 seconds to less than half a second, while tail latency fell from 11.2 seconds to 1.35 seconds. Cost per thousand completed tasks went from $42 to $2.30, Agnost said in its launch materials.
"This is one model for one workload, not a claim that specialist models should replace frontier models everywhere," the company noted, a rare moment of restraint in a sector fond of sweeping pronouncements.
A Two-Person Team With Observability Roots

Palriwala, who joined Cisco's analytics division as its youngest engineer before becoming the first hire at Formbricks, has a background in open-source contributions to Bitcoin, OWASP, and Linux Foundation projects. Parth Ajmera, the CTO, previously built terabyte-scale Spark pipelines at Microsoft and led GPU technology at Infurnia after studying computer science at IIT Madras.
The duo raised a pre-seed round from Entrepreneurs First and Transpose Platform, Palriwala disclosed in a January LinkedIn post. While Agnost hasn't publicly disclosed the amount, EF's standard global pre-seed investment offers up to $250,000, with an optional $125,000 MFN note available through Transpose for participants in Delaware and San Francisco.
Agnost now claims to process more than a million messages daily across its customer base, a figure the company has shared publicly but not verified through independent filings. Customer logos on the company's homepage include Google, Exa, Corgi Insure, Lindy, and Orchid, though the nature of those relationships remains unclear. The startup ranked third in Product Hunt's Product of the Day rankings when it launched there in late August.
A Crowded Market Exploring Similar Territory

Agnost isn't alone in recognizing that production agent logs hold untapped value. Langfuse, acquired by ClickHouse earlier this year, allows users to export production generations and corrections to build fine-tuning datasets. HoneyHive positions dataset creation from real traces as a core feature, according to an updated site published in July. Braintrust described workflows from production traces to evaluation datasets in articles published over the late spring. OpenPipe offers pruning rules to convert chat logs into efficient fine-tuning pipelines, and Weights & Biases' Weave product links agent traces to multi-turn evaluations.
None of those competitors, however, market end-to-end training of task-specific models as a primary service. Microsoft Research released AgentRx in March, a framework for automated failure localization in agent trajectories that underscored how difficult it is to diagnose where multi-step agent processes break down.
Agnost sells its analytics platform separately, with public pricing listed on the company's site: a free tier, a $49 monthly Starter plan, $499 for Pro, and custom Enterprise packages. The company has not disclosed pricing or delivery timelines for the custom model-training service announced last month, nor has it revealed which base models it trains, whether it hosts the resulting models or deploys them within customer infrastructure, or how long the process takes from trace ingestion to delivery.
"Today, we're launching custom models at Agnost AI," Ajmera wrote on LinkedIn. "They performed better than Opus 4.8 on 780 held-out traces while being much faster and cheaper." A live demo sits on the company's homepage alongside a walkthrough video published with the Y Combinator launch.
What remains to be seen is whether enterprises will trust a two-person startup to train models on their production data—and whether workload-specific models trained on failure logs can scale beyond narrow use cases into something resembling a defensible business.
