A four-person team from Y Combinator has built software that promises to slash developers' AI bills by routing most requests to their own computers before touching the cloud — a gambit that tests whether the industry's race toward ever-larger models has left cheaper alternatives sitting on the table.
Conifer, which emerged from Y Combinator's most recent batch and opened its gateway to users in July 2026, operates on a simple premise: not every task needs GPT-4. The San Francisco startup claims its routing system can cut the volume of paid tokens by as much as 80%, steering requests first to open-weights models running on local hardware, then to budget cloud alternatives, and only escalating to frontier APIs when necessary. The company states on its website that it bills cloud tokens at the model's own rate with no markup.
The approach reflects a broader tension in the AI infrastructure market. As enterprises watch inference costs mount, a cluster of startups has emerged to promise smarter routing, caching, and fallback strategies. Conifer's particular angle is aggressive localism: let the developer's own machine handle what it can, bill nothing for those calls, and treat cloud providers as overflow capacity rather than the default.
"Stop paying cloud prices for every token," Michael Jeffords, a co-founder, wrote in the company's launch announcement. "We route 80% of requests to your hardware at no cost." Whether those savings hold across diverse workloads remains to be seen; Conifer has not published independent benchmarks or named customers.
Three tiers, zero markup
The gateway organizes inference into what Conifer calls Tier 0, 1, and 2. Tier 0 runs entirely on the user's machine — open-weights models executed through a Rust-based local runtime that the company says decodes up to 60% faster than llama.cpp on Apple Silicon chips. Those inferences cost nothing. Tier 1 taps efficient cloud models. Tier 2 routes to frontier offerings from OpenAI, Anthropic, and others. The system selects the cheapest tier capable of handling the request, according to documentation the startup published.
Developers set a routing policy: auto lets the gateway choose, balanced trades off cost and capability, or best always picks the strongest model regardless of price. If a developer names a specific model in the request, Conifer serves it exactly as specified, with no silent substitutions. Every response includes headers disclosing which model ran, whether the call was routed, and the cost in nanodollars. The system returns a 402 HTTP status code before executing any call that would breach a developer-set spending limit.

The gateway works with existing OpenAI and Anthropic client libraries pointed at https://api.conifer.build. Conifer bills cloud tokens at list rates, according to its billing documentation, with nothing added on top. Bring-your-own-key requests carry zero fees, the company says, with the underlying provider charging the customer directly.
That pricing puts Conifer in direct competition with OpenRouter, which charges a 5% fee on BYOK usage beyond plan allowances, and with Ramp's Router, which advertises 40% average savings. Other gateways, including Portkey, Helicone, and LiteLLM, position themselves as zero-markup conduits priced on log volume or request throughput rather than token counts.
Rust runtime, open-source SDK
Conifer released a signed macOS desktop application called Juniper and a command-line tool that runs across Mac, Windows, and Linux. The CLI fetches GGUF model files and runs them locally. Windows and Linux desktop apps are in development. The startup open-sourced its SDK at ConiferKit/use-conifer on GitHub, with TypeScript, Python, and MCP support listed in recent commits.

Jeffords and co-founder Charles Muehlberger, who carries the title of CTO, lead a team of four. In community posts, the founders have suggested that between 70% and 80% of workload volume can run locally, depending on the task mix. Coding agents and customer-support bots, they argue, generate high-volume, multi-turn conversations where cheaper models handle the majority of requests. The company has not published enterprise case studies or disclosed which organizations are using the system in production.
Conifer's documentation notes that the gateway supports prompt caching for Anthropic models, with cache reads and writes metered separately at catalog rates. A free routing-decision endpoint at /v1/route returns the gateway's pick and fallback chain without executing the completion, which the company says is useful for logging or external orchestration.
The broader question Conifer is testing is whether the infrastructure layer can meaningfully alter the economics of inference before the model providers do it themselves. OpenAI and Anthropic have both introduced cheaper tiers and caching mechanisms. Startups like Conifer are betting there's still arbitrage to be captured by developers willing to run open models on their own silicon.

