The CFO at a mid-sized insurance company had a straightforward question last fall: Why was the AI budget outpacing revenue growth by a factor of three? The engineering team's answer was less straightforward. They were generating billions of tokens monthly through OpenAI's API, and while the system worked beautifully, each incremental customer was costing more to serve than anyone had anticipated.
That conversation, playing out in variations across hundreds of companies, has triggered what amounts to an industry-wide reckoning. The experimental phase of enterprise AI is over. What's replacing it isn't abandonment—it's optimization, and it's happening fast.
Call it the great unbundling of artificial intelligence. After two years of racing to adopt whatever worked, companies are now asking a more uncomfortable question: what should intelligence actually cost? For a growing number, the answer involves walking away from the big proprietary APIs that got them started.
The shift is visible in the data, though you have to look carefully. By late 2025, roughly 18% of U.S. firms had adopted AI in some form, according to Federal Reserve figures published this past April. Among individual workers, penetration was far deeper—around 41% were using generative AI tools by November. The velocity has been extraordinary, perhaps more than the industry expected. The invoices have kept pace.
When the Bill Comes Due
The economics are brutal in their simplicity. Proprietary models—the OpenAI and Anthropic APIs that powered the first wave of adoption—run between $12 and $30 per million output tokens, according to a July analysis from CUNY AI Lab. Open-weight alternatives clock in at $0.87 to $4.40. When you're in production, generating the kind of volume that real applications demand, those deltas don't just add up. They multiply.
Together AI, one of the more aggressive inference providers, claims cost reductions approaching 90% when companies switch to open-weight models with provisioned throughput. Even discounting for vendor enthusiasm, the magnitude is difficult to wave away.
The pressure hits hardest at companies that moved quickly. Early adopters built atop OpenAI and Anthropic because the APIs were fast, capable, and available when nothing else was. Speed came with a price tag that made sense when AI was a line item. Now that it's a cost center, the math looks different.
Enter what's being called the migration wave—a term that undersells what's actually happening. This isn't just switching vendors. It's a fundamental rearchitecting of how companies think about model deployment, cost management, and vendor dependency.
Y Combinator's recent batch included Understudy Labs, a startup tackling exactly this migration problem with open-source tools designed to help companies make the leap without breaking production systems. The timing isn't coincidental. A confluence of forces—better open models, real cost pressure, infrastructure improvements that actually matter—has turned what was once a fringe engineering concern into something that shows up in board meetings.
What "Open Weight" Actually Means
The terminology itself signals something. "Open weight" has largely displaced "open source" in AI discussions, a shift you can trace through industry publications and model release announcements over the past year. The distinction is more than semantic.
Open-weight models provide access to parameters and architecture without necessarily open-sourcing training data or the full codebase. It's a middle path between proprietary black boxes and fully transparent systems—one that gives companies enough control to self-host and optimize while avoiding some of the complexity that comes with truly open development.
Three things had to converge to make this viable.
First, the performance gap narrowed. Meta's Llama 3.1, which dropped in July 2024, demonstrated parity with frontier models on enough benchmarks to matter. DeepSeek's V3 and V4 families kept pushing. So did Alibaba's Qwen variants and models like Databricks' DBRX. You can quibble with individual benchmark claims—and people do—but the trajectory is unmistakable. Research from the UK's AI Security Institute noted that capability gaps between open-weight and closed models are shrinking, particularly in specialized domains.
Second, the infrastructure caught up, though "caught up" might be generous. NVIDIA's Blackwell GB200 NVL72 architecture started rolling out across cloud providers in late 2025 and early this year, claiming performance improvements of up to 30× over H100s for LLM inference. Software optimizations—TensorRT-LLM, evolving features in vLLM, quantization techniques like FP8—compounded those gains. Microsoft reported a 35% throughput increase for Llama 3.1 405B on its H200-based Azure VMs in documentation updated this spring.
Third, and perhaps most importantly, the tooling ecosystem matured past the science project phase. Model serving platforms like Together AI, Fireworks, and Baseten now offer managed inference for open-weight models with API compatibility layers. Gateway solutions enable hybrid routing strategies. Companies don't have to choose between all-proprietary and all-open anymore; they can route selectively based on task complexity and cost tolerance.
Understudy Labs, founded by Luis Manrique and Aamir Poonawalla—both veterans of large-scale ML infrastructure at Google and Instacart—fits into this third category. Their angle is specific: convert production traces into evaluation datasets, then use those evals to validate that cheaper models can actually handle specific workloads before routing traffic away from frontier APIs.
The product, currently in private preview, functions as a drop-in gateway compatible with OpenAI and Anthropic formats. It proxies existing traffic, logs interactions, and gradually shifts slices of live requests to managed or self-hosted open-weight models. Some components are MIT-licensed and available as open source, while the company maintains a hosted control plane.
Their thesis, outlined in research published in May: production traces plus human judgment become "eval contracts" that gate model substitution. Only when a cheaper model clears held-out evaluations does traffic route to it. Frontier models stay in the mix for tasks where premium capability demonstrably changes outcomes. It's a hedge against over-optimization—saving money without breaking what works.
How It Works in Practice

The implementations vary, but a pattern has emerged. Wholesale replacement is rare. What companies are doing instead is hybrid routing: identify the 5–15% of queries that genuinely benefit from frontier model capabilities, keep those on premium APIs, and route classification, extraction, and tool-call scaffolding through smaller open-weight models in the 7B–35B parameter range.
Cursor, the AI-powered code editor, detailed their approach in a blog post this summer. Working with Together AI, they built a pipeline to move internal open-weight models into production endpoints, leveraging Blackwell hardware with TensorRT-LLM quantization. The result was low-latency inference at scale. Specific cost figures weren't disclosed, but the implication was clear enough.
InterWiz, an AWS customer, claimed cost reductions approaching 90% by migrating to Amazon Bedrock with prompt caching and intelligent routing across models including Llama variants. The case study, published by AWS in July, illustrates the appeal of cloud-native tooling for enterprises reluctant to manage their own inference infrastructure.
At the other end of the spectrum, CUNY AI Lab published a detailed account of their operational deployment entirely on open weights, complete with cost and benchmark comparisons. A legal tech company case study claimed a 60% inference cost cut after fine-tuning Llama-3-8B and serving it on their own A100s, though independent verification was lacking.
The pattern extends beyond individual companies. LangChain's State of Agent Engineering survey, conducted late last year and published in June, found that 57% of respondents already have agents in production. The implication: this isn't experimental anymore. These are live systems generating real costs, which makes them prime targets for optimization.
The Ecosystem Reacts
Major cloud providers have recognized where this is heading. AWS launched Transform model-to-model migration assessments in June, supporting migrations from OpenAI, Gemini, and Anthropic SDKs to open alternatives via LiteLLM and Ollama. It's infrastructure-as-code for model substitution, automating what would otherwise demand significant engineering effort.
Upbound released Modelplane, an open-source control plane for AI inference fleets, in late June, explicitly aimed at organizations "moving to open-weight models." Even OpenAI entered the territory, releasing gpt-oss-120b and gpt-oss-20b as open-weight models under Apache-2.0 licensing in mid-July, though these are self-host only and not available through their API.
The inference layer is becoming its own battleground. Andreessen Horowitz published analysis in late July arguing that inference represents "the largest market in AI" and that companies will consolidate around end-to-end inference stacks rather than fragmented point solutions. The firms positioning themselves as migration facilitators—Together, Baseten, Fireworks, and now Understudy—are betting that thesis holds.
There's also an emerging data layer around this shift. OpenRouter, a multi-model API aggregator, has been publishing usage telemetry showing rising open-weight token share through 2025 and into this year. The "State of Open Source AI" report, released in July, claims open-weight models are capturing meaningful share, though methodology matters and platform-specific data shouldn't be confused with industry-wide trends.
The Hidden Complexity

The per-token price comparisons are real, but they obscure as much as they reveal. Several academic papers this year have cautioned against naive cost modeling.
One arXiv paper from June showed that concurrency and utilization patterns dominate effective cost, with the same H100 hardware yielding anywhere from $0.21 to $15.25 per million output tokens depending on batch size and load. Another July paper examined inference economics for enterprise coding agents and found that high prompt-cache hit rates can invert the API-versus-self-hosted comparison entirely. In their specific workload, effective API cost dropped to $0.57 per million tokens while amortized on-premises cost remained at $2.83—the opposite of what list prices would suggest.
The takeaway? Realized TCO depends on workload characteristics: input/output token ratios, caching efficacy, latency requirements, model selection, quantization strategies. Migrating blindly can backfire if the new architecture doesn't match the use case.
This is part of what Understudy is attempting to solve—not just providing cheaper inference, but validating that the cheaper option actually works for your specific production traffic before you commit. Their benchmark examples claim a 13% higher evaluation score versus Sonnet on certain tasks with a tuned Qwen model, though those figures are vendor-reported and lack third-party audit. The company's zero data retention policy, updated in May, addresses another common migration concern: data privacy during trace capture and evaluation.
Policy Enters the Picture
Cost isn't the only force shaping adoption decisions. Policy is beginning to assert itself, and it complicates things.
The EU AI Act includes provisions exempting certain documentation requirements for models released under free and open-source licenses with public parameters. The European Commission has been actively promoting open and sovereign AI ecosystems; communications this past June explicitly boosted support for open digital ecosystems in public administrations.
In practice, this means European public-sector buyers and regulated industries are increasingly evaluating open-weight alternatives as part of compliance and sovereignty requirements. Deutsche Telekom's Industrial AI Cloud, which went live earlier this year in partnership with NVIDIA, exemplifies the trend—a sovereign hosting platform that secured a German federal government contract in May. France has piloted Mistral models in inter-ministerial AI assistants since last fall.
The U.S. presents a different picture. NIST's AI Risk Management Framework and Generative AI Profile have become de facto standards for federal procurement and large enterprise governance. The framework doesn't mandate open weights, but it emphasizes transparency, safety testing, and continuous evaluation—requirements often easier to satisfy with models you can inspect and host.
At the same time, policy debates about restricting certain Chinese open-weight models like DeepSeek have intensified, creating uncertainty for companies that built infrastructure around them. A July report noted that enforcement of any outright ban would be nearly impossible given how downloadable weights circulate, but the regulatory risk introduces another variable into model selection decisions.
What Happens Next

The migration to open-weight models isn't replacing proprietary APIs wholesale. What's emerging instead is something more nuanced: frontier models for genuinely hard problems, open-weight specialists for routine operations, and intelligent routing between them.
The cost downtrend will continue, driven by hardware improvements, software optimizations, and competitive pressure among inference providers. But realized savings will accrue unevenly. They'll favor teams that invest in workload analysis, evaluation pipelines, and operational maturity around model management—precisely the kind of unsexy infrastructure work that doesn't make for good conference talks but determines whether your AI strategy is sustainable.
For startups like Understudy, the opportunity lies in making that investment easier—turning migration from a six-month infrastructure project into a gradual, validated transition. The open-source components and eval-driven approach suggest positioning for a world where model substitution becomes routine rather than exceptional.
Broader forces will shape how quickly this unfolds. If policy continues tilting toward transparency and sovereignty requirements, open weights gain structural advantages beyond cost. If certain international models face restrictions or scrutiny, American and European open-weight initiatives—Meta's Llama, Mistral, Snowflake's Arctic—may see accelerated enterprise adoption. If frontier closed models leap forward again in capability, the migration calculus shifts back.
For now, the trend line is clear. AI workloads are moving from "whatever works" to "whatever works at acceptable cost," and the acceptable cost threshold keeps falling. The companies building bridges to that future—whether through tooling, infrastructure, or better models—are betting that optimization eventually becomes mandatory.
Given the budget pressures engineering leaders face today, that's not a difficult case to make. The question is no longer whether companies will migrate, but how quickly they can do it without breaking what they've built.
