When the cloud credits dried up last spring, the founder of a generative AI customer service startup watched her burn rate nearly double overnight. She declined to be named, citing investor sensitivities, but her story has become common enough that venture capitalists now have a term for it: the credit cliff.
Application-layer AI startups are incinerating $2 million seed rounds in 12 to 15 months, burning cash roughly 30 percent faster than traditional SaaS companies did at similar stages—a pattern reported across founder posts and finance blogs published between April and June of this year, though it remains informed commentary rather than official benchmark data. The difference comes down to one line item that barely existed three years ago: inference cost. That's the variable expense of running AI models at scale, and it has quietly become the dominant budget-eater for young companies building atop foundation models, even as the price per token has plummeted.
The dynamics are creating an awkward mismatch. Carta data from late last year showed the median time from seed to Series A had stretched past 20 months. Do the math: a startup that exhausts its seed in 12 to 15 months faces a funding gap, forcing either a bridge round or a painful slowdown in growth just when momentum matters most.
Enterprises, meanwhile, are grappling with their own version of the problem. According to Vista Equity Partners, enterprises spent an average of $7 million on model usage in 2025, up from $2.5 million the year before, based on analysis published in June. Some organizations now report monthly inference bills in the tens of millions, Vista said, drawing on a synthesis of survey data from Deloitte. The spending surge prompted OpenAI CEO Sam Altman to acknowledge publicly in June that "AI token costs are becoming a huge issue." He added, with a half-laugh that did not quite mask the seriousness, "It's kind of a meme now that 'My company spent my entire 2026 budget in Q1.'"
The root cause is what the industry calls agentic AI—systems that autonomously execute multi-step tasks. These architectures routinely consume 5 to 30 times more tokens per task than simple chat applications, according to Vista Equity's synthesis of third-party data published earlier this year. Gartner said in a March 25 press release that token consumption is growing faster than unit prices are falling, even as the research firm forecasts a greater-than-90-percent decline in the cost of performing inference on a trillion-parameter model by 2030. TechTarget wrote in July that "inference is the largest long-term expense" in enterprise AI deployments. McKinsey, in a note the same month, said sustainable margins in most enterprise software use cases require cost declines measured in multiples, not percentages.
A Different Kind of Burn
The economics hit startups differently than they hit established companies. Traditional SaaS businesses paid hosting costs that scaled gradually as they added customers. AI application startups pay inference costs that scale directly with usage: every query, every token, every agentic loop adds to the bill. The variable expense structure compresses runway because early adopters tend to be power users, generating high usage relative to the revenue they bring in.
Cloud credits mask the problem initially. AWS Activate provides up to $200,000 in credits for startups, according to program pages updated this year. But finance advisors say burn can jump 40 to 60 percent when those credits expire. One fractional CFO who advises a handful of AI-native startups described it as "the moment founders realize they've been flying blind on their true unit economics."
Per-token list prices have, in fact, fallen dramatically. OpenAI's GPT-4o costs $2.50 per million input tokens and $10 per million output tokens, according to developer documentation verified in September. OpenAI's GPT-5.6 family, announced July 29, priced its Luna tier 80 percent lower than its Sol tier. Anthropic made Claude Sonnet 5 pricing permanent at $2 per million input tokens and $10 per million output tokens, the company said in September. Stanford's AI Index 2025 documented a 280-fold drop in the cost of GPT-3.5-equivalent performance from November 2022 to October 2024.
Yet total spend keeps rising. Agentic systems execute tool calls, maintain long contexts, retry failed outputs, and generate verbose responses—behaviors that rack up token counts faster than any API price cut can offset. OpenAI said in its July 29 GPT-5.6 announcement that its agentic harness reduces "context bloat." Gartner said March 25 that frontier reasoning models should be gated to high-margin tasks, not sprayed across every workflow.

An ArXiv paper published June 10 found that effective cost per million tokens can vary from $0.21 to $15.25 on identical H100 hardware purely due to utilization patterns. The underutilization penalty near idle can reach 36 times the optimized rate, the researchers said—a sobering reminder that raw silicon performance tells only part of the cost story.
Learning to Optimize or Die Trying
Character.AI reduced its serving costs by at least 33 times since launching in 2022, company founder Noam Shazeer wrote in a June 20, 2024 blog post. "It now costs us 13.5 times less to serve our traffic than it would cost a competitor building on top of the most efficient leading commercial APIs," Shazeer wrote at the time. The company moved inference to AMD Instinct MI300 and MI325 accelerators with DigitalOcean this year, doubling throughput and cutting cost per token by roughly 50 percent, according to a case study published by AMD.
Cognition, the startup behind the Devin coding agent, moved to metered self-serve pricing April 14. "It lets us keep investing in products that are genuinely expensive to run," the company wrote at the time. On June 29, Cognition announced a multi-model routing system called Fusion. "You pay a very expensive price" when calling external models without cacheable context, the company noted. The routing system delivered up to 60 percent cost improvement on coding tasks, Cognition said, though the gains varied by workload.
Runway raised $315 million at a roughly $5.3 billion post-money valuation and contracted with CoreWeave for access to GB300 NVL72 hardware, according to a February 10 Sacra profile. The company ported its Gen-4.5 video model from Hopper to Vera Rubin NVL72 architecture in a single day, the profile said, signaling an aggressive push to optimize inference economics before they become an existential problem.
Microsoft shifted its 365 Copilot monetization structure twice in 2026. On April 15, the company paywalled most Copilot features in Office apps, according to Windows Central. On June 16, Microsoft moved Copilot Cowork to pay-per-use billing. The shifts suggest even Microsoft, with its enormous scale advantages, is still figuring out how to make the unit economics work.
The Infrastructure Gold Rush
Venture capital has flowed to companies promising to reduce inference costs, and the pitch decks practically write themselves. Together AI raised an $800 million Series C on July 1. CEO Vipul Ved Prakash framed closed frontier LLM cost structures as "unsustainable in production" for many use cases, the company said. Fireworks raised a $1.5 billion Series D on July 15, emphasizing serverless inference. Baseten raised a $300 million Series E on February 5 with positioning around "production inference that won't break your bank."
Carta data from the first quarter showed more than 60 percent of venture dollars on its platform went to AI companies. Early-stage AI valuations ran elevated relative to non-AI peers, a premium that assumes inference costs will eventually come under control.
Hardware vendors have pushed infrastructure upgrades with near-messianic fervor. NVIDIA and SemiAnalysis claimed multi-fold throughput gains and dramatic cost-per-token reductions on reasoning workloads with Blackwell GB200 and GB300 systems, according to company pages and benchmarking published this year. Vendor claims of cents-level cost per million tokens are benchmark-specific, though, and seasoned operators treat them as directional rather than universal guarantees.
The AEA Journal of Economic Perspectives published an analysis in summer showing that the quality-adjusted "price of intelligence" fell roughly 1,000-fold over recent years. Open-weight models cost approximately 90 percent less than comparable closed models on a quality-adjusted basis, the journal said, hinting at where the long-term price pressure might come from.

Juniper Research said in a summary reported September 2 that Chinese AI models on marketplaces can be "up to 90 percent cheaper" than U.S. alternatives, shifting developer workloads and bifurcating the market into price-driven and quality-driven segments. Geopolitical tensions add complexity, but developers under burn-rate pressure are increasingly willing to route non-sensitive workloads to the cheapest provider.
The Road Gets Harder
The runway compression creates uncomfortable choices. Founders are responding with aggressive optimization, but there's only so much room to maneuver. OpenAI's July 29 GPT-5.6 announcement claimed a 20 percent end-to-end serving cost reduction from kernel and stack optimizations. The company recommended strict context control, prompt caching, and structured output constraints to reduce retries. Multiple operators emphasized reducing output verbosity, summarizing pre-tool outputs, and eliminating unnecessary context as primary cost levers—unglamorous work that nevertheless determines whether the company makes it to the next round.
Software stack choices matter more than many founders initially realize. The vLLM project published a post April 22 detailing FP8 key-value cache quantization. Practitioners using TRT-LLM, SGLang, speculative decoding, and continuous batching reported 15-to-35-times relative cost reductions compared to prior-generation stacks at specific workloads, though gains are workload-dependent and require deep technical expertise to implement.
Model routing is becoming standard practice. Gartner said March 25 that small or flash models should handle routine tasks while frontier reasoning is reserved for high-margin work. The approach trades model sophistication for cost control, which works until competitors figure out how to deliver better results at similar cost.
Regulatory overhead is beginning to add compliance burden, particularly in Europe. The EU AI Act's enforcement for general-purpose AI and transparency requirements begins August 2, according to the official AI Act Service Desk. High-risk system rules phase in December 2, 2027 and August 2, 2028, adding another layer of operational expense to an already tight budget picture.
Energy constraints may ultimately limit supply, creating a different kind of bottleneck. The International Energy Agency said in an update this year that AI-focused data center electricity consumption surged 50 percent in 2025 and forecast that data center power demand will double by 2030. The U.S. Energy Information Administration said between January and April that it expects the strongest four-year load growth since 2000, fueled by data centers. The Federal Energy Regulatory Commission launched an action June 18 to speed large load interconnections, acknowledging grid constraints that could slow the buildout of inference capacity.

Palo Alto Networks CEO Nikesh Arora said in a July interview that "we need to see the pricing for AI come down by 90 percent to become affordable." McKinsey echoed the sentiment in its July note, saying that despite more than $700 billion in hyperscaler capital expenditure committed this year, most enterprise use cases still need multiple-fold cost declines to reach sustainable margins.
Gartner's March 25 forecast assumes per-inference costs will fall more than 90 percent by 2030. If token consumption growth outpaces those declines—and current trends suggest it might—application-layer startups face a stark choice. They can build custom infrastructure to capture vendor margins, a capital-intensive path that few seed-stage companies can afford. They can constrain usage to extend runway, which risks losing to faster-moving competitors. Or they can raise larger rounds at compressed intervals, accepting dilution and higher investor expectations.
The compressed-runway era is forcing founders to treat inference cost not as an operational detail to optimize later, but as a first-order product constraint from day one. The startups that survive will be the ones that figured that out earliest.
