Mendral's bill went down when they started paying more per token. Which sounds wrong—until you understand what they stopped doing.
The YC-backed developer tools company upgraded to Claude Opus in early 2026, Anthropic's premium model that runs $5 per million input tokens compared to $1 for its budget sibling. Their costs for handling CI failures dropped 80%. The secret wasn't the model itself. It was what happened before most requests ever reached it.
By early 2026, Mendral had built something closer to triage than a pipeline. The company routes only the thorniest 20% of CI failure investigations to Opus. The other four out of five? They get intercepted by Haiku—Anthropic's economy model—at a fraction of the cost. "A triager match costs around 25x less than a full investigation," the company explained in a March 6 blog post that laid out the architecture in granular detail.
What Mendral stumbled onto, then systematically documented across three technical posts between February and March, is a recognition that's quietly spreading through the AI infrastructure world: most production systems burn money on overkill. They send log parsing and metadata extraction to frontier reasoning models when cheaper alternatives handle those tasks perfectly well.
Perhaps more importantly, they demonstrated that intelligent routing between model tiers can deliver better results at dramatically lower cost than the brute-force approach of throwing everything at a single expensive model.
How the Economics Actually Work
The system Mendral designed operates as a three-tier hierarchy, each layer doing different work at different price points. Haiku 4.5—the budget model—handles log parsing and initial triage. Sonnet 4.6, the mid-tier option, gathers evidence and executes SQL queries against the company's ClickHouse database. Opus 4.6, the expensive one, performs root-cause analysis and generates non-trivial fixes.
The pricing structure matters here: Haiku costs $1 input and $5 output per million tokens. Sonnet runs $3 and $15. Opus hits $5 and $25.
Now consider the distribution. Haiku processes roughly 65% of all input tokens but represents just 36% of Mendral's total LLM spend. The expensive model thinks. The cheap model reads. When a continuous integration pipeline breaks, the Haiku agent first decides whether it recognizes the failure pattern. If it does, the investigation stops there—problem solved, cheap. If not, it escalates to Sonnet for evidence gathering or directly to Opus for complex reasoning.
It mirrors, in a way, how engineering teams naturally organize themselves. Junior developers handle routine issues. Senior engineers tackle the novel problems. The difference is that LLM tokens don't complain about being underutilized or argue about promotion cycles.
Pulling Context Instead of Pushing It
Mendral's second major innovation involves data handling, and it runs counter to a pattern that's become surprisingly common in production AI systems. Instead of stuffing raw logs into prompts—a strategy that bloats token counts and can degrade model performance—their agents query for exactly what they need.
"Let the agent pull context, don't push it," the company wrote in their March post. It's advice that sounds simple but requires rethinking how you structure the system.
Here's how it works in practice: Opus delegates targeted SQL queries to Haiku sub-agents. Those sub-agents extract specific information from ClickHouse and return structured summaries. The orchestrator's context window stays clean. Token consumption stays bounded.
The approach eliminates what Mendral calls "anchoring bias"—when models fixate on irrelevant details buried in massive log dumps instead of reasoning about the actual problem. Give a model 10,000 lines of logs and watch it get distracted by warnings that have nothing to do with the failure you're investigating.
This strategy aligns with a broader trend. The Register reported in April that text-to-SQL systems are experiencing a renaissance as language models improve at translating natural language into database queries. Mendral's implementation demonstrates the cost benefits: you pay for small, targeted queries instead of embedding thousands of lines of logs in every prompt.
The sub-agent fan-out is deliberately constrained to one level—a simple guardrail that prevents runaway costs. Opus can spawn Haiku agents to fetch data, but those agents don't spawn additional agents. It's unglamorous but effective cost control.
The Multipliers You Don't See Coming
Anthropic's pricing structure includes several less-obvious factors that can inflate bills if you're not paying attention. Opus 4.7, which shipped in April, uses a new tokenizer that can consume up to 35% more tokens for identical text compared to earlier versions. Fast mode for Opus 4.6 carries a 6x multiplier—$30 input and $150 output per million tokens. Tool-use functionality adds 313 to 346 tokens of system prompt overhead to every single request.
Mendral's architecture minimizes exposure to these multipliers, though perhaps not by design initially. By routing most requests to Haiku and Sonnet, they avoid the Opus tokenizer tax on routine work. Their agents use tools strategically rather than universally. And they've optimized their prompts to maximize cache hits.
Prompt caching is where things get interesting mathematically. Anthropic offers a 90% discount on cache reads—you pay just 0.1x the normal cost—after paying a 1.25x premium to write the cache initially. AWS Bedrock added a one-hour cache window option for select Claude models in January, useful for long-running agent sessions. Mendral caches system prompts, tool schemas, and documentation—stable content that gets reused across many requests.
The math works if you get at least one or two cache reads per session. After that, you're saving tokens. Stack that with the 50% discount from Anthropic's Batch API for non-interactive workloads, and the optimizations compound.
An Industry Grappling With Cost
Mendral's approach arrives at a moment when cost concerns threaten to derail production AI deployments across the industry. Gartner warned in April that over 40% of agentic AI projects may be canceled by the end of 2027 due to rising costs, unclear value, and technical risks. Axios reported the same month that some AI use cases already cost more than equivalent human labor—a sobering data point for anyone betting on immediate economic returns.
The pricing landscape remains fluid, which complicates planning. OpenAI's GPT-5.5 runs $5 and $30 per million tokens. GPT-5.4 costs $2.50 and $15. Google's Gemini 3.1 Pro commonly costs around $2 and $12, though rates vary by region and volume. DeepSeek advertises materially lower headline rates for certain models. All providers offer cache discounts and batch pricing that can halve costs for suitable workloads.
But raw pricing obscures the real story, as Mendral's results illustrate. Model capability matters as much as cost per token.
Research published in January on "shepherding"—where cheaper models provide hints to guide expensive models—showed 2.8x cost reductions compared to simple routing or cascading approaches. The older FrugalGPT paper from 2023 demonstrated up to 98% cost cuts through intelligent routing and prompt adaptation.
Mendral's 80% reduction validates these patterns in production. Using Haiku for root cause analysis produces worse results. Using Opus for log parsing wastes tokens. The key is matching model capability to task complexity, then building the infrastructure to enforce that matching at scale. Easier said than done.
Infrastructure That Actually Prevents Waste

The operational harness matters as much as the routing logic, maybe more. Mendral runs on Inngest for durable execution, which means a failed step doesn't force the system to re-run expensive LLM calls from scratch. State persists between steps. Retries happen at the right granularity. Their Firecracker microVMs support suspend and resume, eliminating idle compute charges while agents wait for external services.
These aren't flashy optimizations. They're the unglamorous work of production engineering—the difference between a clever demo and a system that runs reliably at scale without bleeding money. It's the kind of work that doesn't make for good conference talks but shows up directly in the monthly bill.
Mendral imposes budgets per step and caps sub-agent depth. They maintain evaluations on routing thresholds because model capabilities improve over time and work that required Opus six months ago might run fine on Sonnet today. The routing rules aren't static; they evolve with the models. Which means someone has to actively manage them.
A Pattern That's Starting to Spread
AWS introduced Bedrock Intelligent Prompt Routing in 2024, which automatically routes between Haiku and Sonnet-class models to balance quality and cost. The fact that cloud providers built native routing suggests this pattern has moved beyond research labs into production infrastructure. It's becoming table stakes.
But provider-level routing can't capture application-specific knowledge about which tasks actually need frontier reasoning versus simple pattern matching. Mendral's architecture works because they've specialized their agents for CI/CD workflows—different prompts, different tools, different data structures than a general-purpose coding assistant would use.
"Same model weights, completely different harness," they wrote in a March 30 post.
That specialization is where the cost savings really compound. A generic agent routing layer saves money. An agent designed from the ground up around tiered reasoning? It saves more.
Mendral was founded by Sam Alba and Andrea Luzzardi, both Docker and Dagger alumni, and went through Y Combinator's Winter 2026 batch. Their technical posts read more like engineering documentation than marketing—complete with specific token counts, cost breakdowns, and SQL query examples. The kind of detail that suggests they've actually run this in production rather than just theorizing about it.
What It Means for Everyone Else

The broader implication is that cost optimization for AI agents isn't just about negotiating better rates or waiting for hardware to get cheaper. It's architectural. Multi-tier routing, structured data access, prompt caching, durable execution, and specialized agent design compound in ways that slash costs while maintaining—or in some cases improving—quality.
As Opus 4.6 and Sonnet 4.6 ship with million-token context windows at standard pricing, the opportunity to build sophisticated multi-agent systems grows. But so does the risk of runaway costs if those systems aren't designed with economic constraints baked in from the start.
Mendral's 80% reduction demonstrates what's possible when you treat model selection as an engineering problem rather than a checkbox on a feature list. Four out of five CI failures never need frontier reasoning. Which means the expensive model should be the exception, not the default.
That sounds obvious in retrospect. Most good engineering decisions do.
