Understudy Labs, a two-person company in Y Combinator's latest batch, has launched an inference platform that it claims can reduce large language model expenses by 80% — a bold, unverified pitch aimed at engineering teams whose monthly Claude and GPT bills sometimes stretch into six figures.
The San Francisco startup's approach sidesteps the usual trade-off between cost and performance. Instead of simply routing requests to cheaper off-the-shelf models, Understudy automatically trains specialized versions from a customer's own production data, then deploys them only after they outperform the pricier frontier APIs on that company's evaluations.
It's a technique that addresses a mounting problem in enterprise AI. As companies scale their use of models from OpenAI and Anthropic, inference costs balloon. Manual fine-tuning offers one escape route, but it requires machine learning expertise many teams lack. Understudy's founders, both Instacart alumni who led ads machine learning positions, built their platform to make that shift automatic.
The system functions as a drop-in proxy. Teams redirect their API calls to Understudy's endpoint, which captures every request and response in production. Those traces become training data for smaller, open-weight models. When a newly trained model beats the incumbent on a customer's own benchmarks, it gets promoted into the routing table. Deployment can happen on Fireworks, AWS Bedrock, Vertex AI, or a customer's own GPUs, though hosted infrastructure is optional.
CEO Luis Manrique previously worked as the first product manager on Instacart's Carrot AI consumer team and later led ads machine learning there before joining Gumloop as a founding technical hire, where he reportedly drove around $2 million in first-year revenue. His co-founder, Aamir Poonawalla, spent nine years at Instacart building the ads auction platform and experimentation framework. Poonawalla, who holds a master's in computer science from Georgia Tech, is going through Y Combinator for the second time.
Benchmarks suggest substantial savings, though with caveats
The company has published its own performance numbers, which have not been independently verified. In one customer relationship management example, a trained model scored 13% higher on evaluations than Claude Sonnet 4.6 while consuming a quarter of Sonnet's token budget, according to Understudy's internal testing. Latency on an 8-billion-parameter route fell to 369 milliseconds compared to 1.935 seconds for Sonnet, a roughly fivefold improvement with six times lower token costs on a 90-trajectory validation set.
Another test involved sentiment labeling across nearly 40,000 YouTube comments. Understudy's post-trained 30-billion-parameter model processed the task for $2.82, versus $12 for Claude Sonnet and $140 for Opus, with 99.3% three-way agreement on dense labels, according to the company's own benchmarks. The company notes a 90-times price gap between Sonnet's $18 per million tokens and Qwen3-8B's $0.20 per million tokens.
What Understudy hasn't disclosed is what it charges for the gateway and training service itself. The platform remains in private preview, with access granted by request.
A crowded field, but a different angle
Understudy enters a market where inference providers are racing to undercut frontier pricing. Together AI claimed in July it could deliver provisioned throughput at up to 90% below Claude Opus pricing for reserved capacity workloads. Fireworks AI and Baseten publish per-token rates for open models. OpenRouter and Portkey AI offer multi-provider routing with cost optimization built in.

Distil Labs, another startup in this space, makes a similar "up to 80%" cost reduction claim and has cited a 68% savings case with a customer called Knowunity.
Where Understudy distinguishes itself, at least in theory, is automation. Rather than simply routing to pre-existing cheaper alternatives or requiring customers to manage fine-tuning themselves, it trains models from production traces and handles promotion decisions based on performance thresholds the customer defines.
The company has open-sourced some agent tooling under an MIT license on GitHub, including local evaluation and prompt optimization utilities designed to work with Claude Code and Cursor. Updates have appeared as recently as the past week.
Still early, with questions ahead
"We started Understudy because we believe every company using AI should be accumulating intelligence, not accumulating API bills," the founders wrote on their Y Combinator profile. It's an appealing framing, though the startup has yet to publish customer logos or independent case studies.
The platform's architecture separates data-plane traffic from control-plane management, meaning model serving continues even if the control plane fails. Response headers expose which model handled each request, allowing teams to audit routing decisions in real time.
Whether enterprises will trust a startup to sit in the critical path of their AI infrastructure, capturing every prompt and response, remains an open question. So does the matter of how well models trained on one company's narrow use case actually generalize, and whether the promised savings hold as workloads evolve. For now, teams curious enough to find out can request access through the company's website.

