Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Climate / Social Tech iconClimate / Social TechMay 15, 2026

GridCARE Raises $64M to Unlock Hidden Power for AI Data Centers

GridCARE Raises $64M to Unlock Hidden Power for AI Data Centers
Ai InfrastructurePower Infrastructure+3
SaaS iconSaaSMay 15, 2026

YC-Backed Vector Raises $10M Series A for Contact-Level Ad Platform

YC-Backed Vector Raises $10M Series A for Contact-Level Ad Platform
YcAd Tech+3

Founders Mentioned

Unknown

Expanse

saas icon
SaaS

Unknown

Expanse

saas icon
SaaS
SaaS iconSaaS
May 15, 2026
Ai InfrastructureCost OptimizationEnergy EfficiencyPower Infrastructure

The Race to Make AI Inference 100x Cheaper—And Why It Matters

As AI inference costs plummet 280x in three years, startups like Doubleword are pushing new limits—but power grid constraints may determine who wins the efficiency wars.

The Race to Make AI Inference 100x Cheaper—And Why It Matters

There's a number that keeps showing up in pitch decks and engineering blog posts, and it's the kind of number that makes investors pause mid-sip of their coffee: 280-fold. That's how much the cost of running GPT-3.5-level inference has fallen between November 2022 and October 2024, according to Stanford's latest AI Index. OpenAI now cuts its prices in half if you're willing to wait a day for results—a policy that took effect March 31, 2026. A London startup called Doubleword talks openly about 99% discounts on certain jobs.

The economics have inverted so quickly that even people inside the industry seem a little stunned by it.

But here's the thing nobody's quite saying out loud yet: the plummeting costs might be solving the wrong problem. Because while the price per query has cratered, the actual constraint—the thing that determines who wins and who gets left behind—isn't compute efficiency anymore.

It's power. As in, can you plug in enough data centers to meet demand? Do you have electrons to burn?

The Overnight Discount

The simplest arbitrage in artificial intelligence right now doesn't involve some breakthrough in transformer architecture or exotic new semiconductor design. It's much more prosaic than that. Just be willing to wait.

OpenAI's Batch API, with pricing guaranteed through at least March 2026, drops costs by half for any job that can tolerate a 24-hour turnaround. Together AI and Fireworks.ai have rolled out similar tiered structures—cached inputs at half price, overnight windows with steeper cuts. Doubleword, which raised roughly $12 million in May 2025 and rebranded from its earlier incarnation as TitanML, has pushed this logic to its extreme. The company's async pricing for DeepSeek-V4-Pro can hit $1.31 per million tokens overnight. Compare that to what you'd pay for instant responses elsewhere, and you start to see why people are interested.

Their case studies read like cost-cutting fantasies. OpenMed, working on a medical imaging dataset, annotated over 119,000 images for about $450 using frontier vision-language models routed through Doubleword's infrastructure. The same work through Anthropic's standard API? Closer to $7,500, according to their estimates—a 94% haircut. Dataiku built out a synthetic dataset for privacy testing at $50, a job that might have run $1,000 through conventional providers.

The insight underneath all this is almost embarrassingly straightforward: when nobody's sitting there waiting for an answer, you can cram far more work into the same silicon. Continuous batching, speculative decoding, aggressive compression of the key-value caches that models use to track context—all those techniques that might add a few milliseconds of delay in real-time use become essentially free optimizations when you've got hours to play with. Doubleword's engineers have a pithy way of putting it: "When no one is waiting, the optimization target changes: not latency → IQ per dollar."

Which is to say, you're no longer racing the clock. You're racing the efficiency curve.

The Engineering Behind the Collapse

None of this is magic, though it can feel that way when you watch the numbers. It's really a story about open-source software, hardware iteration, and algorithmic tricks compounding in ways the market is still digesting.

Start with the serving engines. vLLM arrived in 2023 with a technique called PagedAttention that treats GPU memory the way an operating system manages RAM—result, throughput jumps of 2x to 4x. SGLang, with updates rolling out through 2026, added deterministic inference and faster cold starts. Doubleword's implementation can now spin up models on NVIDIA's B200 chips in about 10 seconds. Earlier versions? Try 12 minutes. TensorRT-LLM, NVIDIA's optimized runtime, layers on additional performance tuning, especially for quantized models where you trade a bit of precision for speed.

Then you get into key-value cache compression, which sounds technical because it is. Papers from the past year or so show you can quantize these caches down to 2 bits with barely noticeable quality loss. One study, KIVI, reported throughput improvements approaching 3.5x. Doubleword's stack folds in these techniques alongside what they call "QueueSpec" speculative decoding—basically, a smaller model drafts tokens that a larger one verifies, accelerating generation when the guesses are good.

NVIDIA's Blackwell architecture, the GB200 and NVL72 variants, promises cost and energy reductions up to 25x per inference compared to the previous Hopper generation. Those are NVIDIA's claims, for trillion-parameter models in real-time scenarios, and the reality will likely be messier. Still, even if you discount the marketing by half, the improvements are material. And they stack with software optimizations in ways that make isolating any single factor nearly impossible.

Consider: a Salesforce production study published in April 2026 documented tail-latency reductions above 50% and throughput gains near 4x using a modular inference setup. Those translated to 30-40% cost savings at scale. Not in a lab. In production, with real customer traffic.

Inference Eats the World

Digital illustration for article section "Inference Eats the World" in "The Race to Make AI Inference 100x Cheaper—And Why It Matters" - A clean, minimalist conceptual image representing continuous AI inference overtaking initial trainin...

Gartner made a call last October that felt significant at the time and looks more so now: AI infrastructure spending for inference, they said, would overtake training spending sometime in 2026. The logic tracks. Training a frontier model is expensive—wildly so—but it's a one-time hit. Inference is the recurring cost that scales with every user query, every API call, every background job you run. As AI features metastasize across applications, inference becomes the ongoing burn rate.

You can see the gold rush in the funding rounds. Together AI raised $305 million in February 2025. Fireworks.ai closed a $250 million Series C at a $4 billion valuation last October, explicitly to scale inference capacity. GPU infrastructure providers like CoreWeave and Crusoe have expanded at breakneck speed, hunting for power-advantaged regions to plant their data centers.

The hyperscalers, unsurprisingly, are adapting. AWS SageMaker integrated NVIDIA's inference microservices in early 2024, smoothing the path to optimized deployments. The distinction between "cloud provider" and "specialized inference platform" is getting fuzzy as everyone optimizes for the same bottleneck: squeezing more intelligence out of every watt.

The Electron Problem

Digital illustration for article section "The Electron Problem" in "The Race to Make AI Inference 100x Cheaper—And Why It Matters" - A conceptual, minimalist representation of the rising electricity demand for AI facilities, featurin...

Which brings us to the uncomfortable part of the story.

Unit costs are in freefall, yes. But total bills? Those face an entirely different constraint.

The International Energy Agency's analysis from earlier this year found that data center electricity demand grew 17% in 2025. AI-focused facilities? Up roughly 50% year-over-year. The capital expenditure of five major tech companies topped $400 billion last year, with projections for a 75% bump in 2026. Much of that money is going toward securing power capacity before someone else does.

Morgan Stanley's outlook suggests U.S. data center power demand could hit 74 gigawatts by 2028—with a shortfall approaching 49 gigawatts. Carbon Direct's May analysis of grid interconnection queues found over 10 gigawatts of new data center capacity announced just since last October. Only about 30% of that is expected online by 2027. The queues themselves exceed 300 gigawatts of planned capacity. Projects waiting, sometimes for years, for grid access.

The practical upshot: data center projects are getting delayed or outright canceled despite screaming demand. Some jurisdictions have imposed moratoria on new facilities. The constraint isn't chip availability or access to capital. It's kilowatt-hours, plain and simple.

This creates an odd paradox. Software innovations can deliver 100x cost reductions on specific tasks—Doubleword's customers demonstrate this—but those gains assume you can actually get power to begin with. The winners of the efficiency wars may well be the companies with the best grid relationships, not the cleverest algorithms. Which is perhaps more than the founders expected when they set out to optimize inference engines.

Regulation Enters the Chat

Technical and power constraints aren't the only forces reshaping the landscape. Regulatory frameworks are starting to influence architectural choices in ways that advantage certain approaches over others.

The EU AI Act went into force in August 2024, with most obligations kicking in by August 2026. General-purpose AI model requirements arrived last August. The framework creates risk-based compliance demands that nudge enterprises toward approaches offering greater transparency and control—often self-hosted or hybrid deployments rather than pure API consumption through some distant provider.

The NIST AI Risk Management Framework and its Generative AI Profile, published last July, provide similar guidance stateside, shaping how enterprises evaluate and procure AI systems. Together, these frameworks create a kind of downstream pressure: if you need to demonstrate model provenance, explain decisions to regulators, or guarantee data residency, the cheapest API call might not represent the lowest total cost of ownership.

Export controls on advanced accelerators add yet another variable, affecting chip availability and pricing across geographies. The competitive landscape for inference infrastructure is increasingly about regulatory arbitrage as much as technical innovation. Not exactly the story the engineers want to tell, but there it is.

What It Means for Builders

Digital illustration for article section "What It Means for Builders" in "The Race to Make AI Inference 100x Cheaper—And Why It Matters" - A minimalist and conceptual visualization representing the race for optimized inference and latency-...

The race to cheaper inference is real. But it's not one race—it's at least three, with different winners.

For latency-sensitive applications—chatbots, real-time assistants, interactive tools—the optimization target remains milliseconds per token. Specialized hardware and model quantization matter more here than batch scheduling. Cost reductions are meaningful but incremental. Think 3-5x over two years rather than 100x.

For throughput-oriented workloads—data labeling, synthetic data generation, evaluation pipelines, document processing—the async dividend is genuinely enormous. This is where Doubleword's case studies live, where OpenAI's Batch API economics make sense. Companies with substantial workloads that can tolerate hours of latency can see order-of-magnitude cost drops by rearchitecting around batch processing. The savings are real if you're willing to rethink your workflows.

For agentic systems and background pipelines—increasingly common as AI moves from user-facing features to infrastructure—the economics shift again. Multi-step reasoning chains, tool use, long-running processes create opportunities for hybrid approaches that route requests to different inference tiers based on urgency and complexity. The platforms that win here will be the ones that make that routing invisible to developers.

But the power constraint cuts across all three categories. Even with perfect software efficiency, you still need electrons. The companies locking in long-term power purchase agreements now, building in regions with surplus generation capacity, or partnering early with utilities are buying optionality that no amount of technical cleverness can replicate.

Perhaps the clearest signal is that the conversation has shifted from "can we afford to deploy AI?" to "can we get enough power to deploy it at the scale we want?" When your constraint moves from dollars to watts, the efficiency battle has already been won.

The infrastructure battle, though? That one's just getting started.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • GridCARE Raises $64M to Unlock Hidden Power for AI Data Centers
  • YC-Backed Vector Raises $10M Series A for Contact-Level Ad Platform
  • Big Tech's AI Brain Drain: $18B Follows Researchers to Startups
  • DeepMind's David Silver Raises $1.1B to Build AI That Learns Like Humans
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.