Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSMay 14, 2026

Ardent Launches 6-Second Database Cloning for AI Agent Testing

Ardent Launches 6-Second Database Cloning for AI Agent Testing
YcAi Agents+3
SaaS iconSaaSMay 14, 2026

From UX to AX: Why 40% of Enterprise Apps Are Redesigning for AI Agents

From UX to AX: Why 40% of Enterprise Apps Are Redesigning for AI Agents
Ai AgentsEnterprise Software+3

Founders Mentioned

Unknown

Expanse

saas icon
SaaS

Unknown

Expanse

saas icon
SaaS
SaaS iconSaaS
May 14, 2026
Gpu CloudAi InfrastructureCost OptimizationCloud InfrastructureResource Optimization

The $725B GPU Waste Crisis: How Predictive Tools Unlock Idle Capacity

With GPU utilization averaging just 5%, startups like Expanse are pioneering predictive resource optimization to unlock billions in wasted compute before it hits the scheduler.

The $725B GPU Waste Crisis: How Predictive Tools Unlock Idle Capacity

The numbers refuse to reconcile. Big Tech's capex, including the top eight cloud providers, will reach an estimated $725 billion in 2026, analyst projections suggested back in May. Microsoft has warned it will remain capacity-constrained through year's end, at minimum. CoreWeave—a data center upstart that's become synonymous with the GPU gold rush—recently crossed the 1-gigawatt threshold and is virtually sold out of 2026 capacity, with plans to expand to 5 gigawatts by 2030.

And yet.

Across tens of thousands of production Kubernetes clusters analyzed in a report published April 21, 2026, the average GPU utilization rate sits at 5%.

Five percent. The figure, published by Cast AI in their 2026 State of Kubernetes Optimization Report released on April 21, represents perhaps the starkest disconnect between capital deployment and operational efficiency in the history of modern computing. If that baseline holds—and there's little reason to believe it won't—the industry's capex binge is leaving hundreds of billions of dollars in compute capacity gathering dust while companies continue scrambling to lock down more chips.

The crisis, it turns out, isn't really about scarcity anymore. It's about waste. Waste on a scale the industry has never quite seen.

A Problem That's Been Hiding in Plain Sight

Low utilization didn't start with the generative AI boom, but it has metastasized alongside it. A Microsoft Research study published in April 2024—examining 400 internal deep learning jobs—found median GPU utilization hovering around 50% or lower. The truly troubling detail? Nearly half the inefficiency stemmed from basic data operations, and 85% of the issues could be fixed with minor tweaks to scripts or code. Academic research published late last year noted that "real deployments continue to report average utilization near 50%," a figure that now looks almost generous when set against Cast AI's data.

The Cast AI report, which covers production environments through April of this year, suggests the situation has deteriorated. Alongside that 5% GPU figure, CPU utilization averaged 8% and memory just 20%. Year over year, overprovisioning has increased across the board. Context from the Cloud Native Computing Foundation's annual survey, released in January, helps explain why: 82% of container users now run Kubernetes in production, and two-thirds of organizations hosting generative AI models rely on Kubernetes for inference. But only 7% deploy models daily. The implication is uncomfortable—most of this infrastructure sits waiting for workloads that arrive sporadically, if they arrive at all.

The economics have shifted, too. After selective GPU price cuts through much of last year, AWS raised prices for its EC2 Capacity Blocks for ML by roughly 15% in early January, particularly for H200 and P5e instances. When guaranteed capacity costs more and utilization lingers in the single digits, every idle hour compounds the misalignment. As VentureBeat observed in May, procurement FOMO has driven companies to pay for GPUs they don't actually use. And the prices just keep climbing.

Meanwhile, power has overtaken chip supply as the binding constraint. DataCenter Knowledge and TechTarget both reported in late April and early May that grid access and permitting delays are now gating new builds more than semiconductor availability. The International Energy Agency's reports this year have emphasized the scramble for electricity, grid connections, and the capital to secure both. CoreWeave's ambitious expansion will require navigating interconnection timelines measured in years, not months. Perhaps longer, depending on local politics.

What's Driving the Disconnect?

The root causes span technical, organizational, and—perhaps most frustratingly—economic domains. Start with something relatively fixable: resource requests are fundamentally guesswork. Most clusters still lack any pre-submission intelligence about how much memory, CPU, or GPU time a job will actually consume. Engineers over-allocate to avoid out-of-memory failures or queue rejections. Schedulers, for their part, accept those inflated requests at face value.

Microsoft's study from 2024 highlighted data pipeline bottlenecks as a leading cause of GPU idling. Training jobs starve GPUs while waiting for data preprocessing, shuffling, or augmentation. The fixes are often embarrassingly simple—adjusting batch sizes, parallelizing data loaders, caching frequently accessed datasets—but without visibility into what's actually blocking GPU cycles, teams don't know where to intervene. Or sometimes, they just don't bother.

Then there's organizational maturity, which hasn't kept pace with adoption velocity. The CNCF survey found that cultural and process challenges now surpass technical barriers as obstacles to effective Kubernetes deployment for AI workloads. Teams are deploying cutting-edge models on infrastructure they haven't fully learned to optimize. Only 7% of organizations deploy models daily; the rest run sporadic workloads on hardware provisioned for peak demand that rarely materializes. It's the enterprise equivalent of buying a sports car for a weekly grocery run.

Economic incentives haven't aligned to penalize waste—until recently, anyway. GPU availability mattered more than GPU efficiency. Companies stockpiled capacity wherever they could secure it. Utilization was an afterthought. But the January price hike for AWS Capacity Blocks signals a shift in the weather. When guaranteed access carries a premium and new capacity is constrained by power rather than chips, the cost of idling a reserved H100 or H200 becomes harder to explain during quarterly reviews.

Regulatory and environmental pressures are mounting, too. The EU's Energy Efficiency Directive requires data centers with 500 kilowatts or more of IT power to report energy and sustainability metrics annually. The first submissions arrived last September; subsequent cycles are due every May 15. Aggregate reporting across the EU showed 6.4 gigawatts of IT power by April. In the U.S., dozens of local jurisdictions have enacted data center construction bans or moratoriums. Maine passed the first statewide moratorium, effective until late 2027. Visual Capitalist, Axios, and TechRadar have tracked these actions through the spring. Operators facing public scrutiny on energy use and community impact now have fresh incentives to demonstrate they're maximizing the capacity they already possess.

Experiments at the Frontier

Digital illustration for article section "Experiments at the Frontier" in "The $725B GPU Waste Crisis: How Predictive Tools Unlock Idle Capacity" - A modern, abstract conceptual illustration representing dynamic resource fractioning and intelligent...

Nvidia's acquisition of Run:ai last December positioned the GPU giant to tackle utilization directly, rather than leaving it to customers. Run:ai's orchestration layer sits between workload schedulers and hardware, enabling dynamic GPU fractioning, memory oversubscription, and intelligent queuing. Technical blogs Nvidia published in March and April benchmarked the improvements: up to roughly 2x utilization gains and up to 1.4x higher throughput at heavy concurrency via dynamic fractions. A customer case study from 2024 documented one organization moving from 28% to 73% utilization after deploying Run:ai, though that predates current market conditions—and current desperation.

The integration reflects a strategic bet. Nvidia isn't content selling chips; it wants to own the orchestration layer that determines how efficiently those chips run. The company donated a GPU Dynamic Resource Allocation driver to the CNCF in March, aligning with Kubernetes 1.36's stabilization of DRA for GPU workloads in April. With DRA reaching general availability, GPU-aware scheduling becomes more standardized, and tools that plug into that layer gain traction.

Alibaba Cloud's Aegaeon system, presented at a symposium last fall and reported in October, took a more radical approach: token-level GPU pooling. Rather than allocating entire GPUs or fixed fractions to jobs, Aegaeon pools compute at the token generation level, dynamically routing inference requests to available GPU resources across a cluster. In internal trials, Alibaba reported up to 9x "goodput" and claimed an 82% reduction in required GPUs for their marketplace workloads. The system remains largely internal, and the data is now more than six months old. Still, it signals the outer bounds of what aggressive pooling strategies might achieve at hyperscale.

Open-source schedulers have proliferated in response. HAMi, a CNCF Sandbox project, enables GPU slicing and heterogeneous device virtualization. A case study involving NIO published in March noted utilization of 5 to 10% under full-GPU allocation before HAMi integration—a familiar refrain. Volcano, another batch scheduler, supports vGPU and Multi-Instance GPU configurations, with documentation updated through the middle of last year. Kueue, from the Kubernetes SIG Scheduling group, integrates with DRA and provides cluster-wide quota and queuing policies. Armada, from G-Research and also in the CNCF Sandbox, acts as a multi-cluster meta-scheduler, aiming to reduce fragmentation by pooling workloads across Kubernetes clusters in different data centers or clouds.

But schedulers alone don't solve the overallocation problem. They work with the resource requests they receive. If a user submits a job asking for 80GB of GPU memory when it only needs 40GB, even the smartest scheduler can't reclaim that waste. The job sits there, hogging resources, oblivious.

Enter predictive optimization—a newer approach that tries to get ahead of the problem. Expanse, a San Francisco startup founded last year and part of Y Combinator's Spring 2026 batch, is building what it calls a "pre-submit intelligence" layer. The four-person team, which claims lineage from building multimodal HPC resource predictors at EPCC and stints across quantitative funds and national supercomputing centers, focuses on predicting runtime, memory, CPU, and GPU requirements before a job ever hits the scheduler.

The value proposition is straightforward: right-size resource requests to unlock hidden capacity. Expanse's platform analyzes historical telemetry from a cluster—via integrations with Slurm, Kubernetes, Databricks, YARN, Cloud Batch, and Nomad—to predict what a job will actually consume. It also surfaces optimization suggestions based on local job and cluster history, and attempts to predict failures before they burn through resources. The company offers an open-source "Wastage Scanner" for Slurm and Kubernetes environments; as of May, the tool's live counter claimed analysis of 538,900 jobs, identifying 95.9 million wasted core-hours and $9.6 million in estimated waste across four scanned clusters. Whether those numbers hold up under broader scrutiny remains to be seen.

Expanse targets research labs and quantitative funds—environments where job diversity and resource heterogeneity make manual tuning impractical. Its positioning is complementary to schedulers like Kueue or Volcano, not a replacement. By feeding better requests into existing orchestration layers, it aims to reduce queue times, prevent out-of-memory failures, and surface code-level inefficiencies without requiring workflow changes. That last part matters—enterprises are wary of platforms that demand they rewrite everything.

Academic research supports the feasibility of prediction-based approaches, at least in theory. A study on malleable job scheduling published in February demonstrated utilization improvements of 5 to 52% depending on supercomputer architecture and policy. An April paper described a two-stage prediction model for GPU resource and power consumption based on Slurm and DCGM logs. Another, also from April, explored short-term power forecasting for AI data centers. An RL-based scheduler presented at a cloud computing conference last December claimed up to 20% utilization improvement and up to 81% reduction in queuing delay, though results were simulation-based using traces from Microsoft and Alibaba. Real-world deployments, as always, tend to be messier.

The pattern across these efforts is consistent: visibility unlocks optimization. Whether through dynamic allocation, intelligent fractioning, multi-cluster pooling, or predictive right-sizing, the common thread is replacing guesswork with data. It's not glamorous, but it works.

What Comes Next

Digital illustration for article section "What Comes Next" in "The $725B GPU Waste Crisis: How Predictive Tools Unlock Idle Capacity" - A dynamic, stylized illustration of an oversized, expressive performance gauge with its needle pushi...

Utilization is poised to become a primary KPI for AI infrastructure ROI over the next year or two. With GPU prices for guaranteed capacity rising, grid constraints slowing new builds, and CFO scrutiny intensifying, organizations can no longer afford to treat single-digit utilization as an acceptable baseline. Gartner has projected $2.5 trillion in global AI spending for 2026, with $401 billion earmarked specifically for infrastructure. TrendForce estimates that the combined capex of the top eight cloud service providers will exceed $710 billion. At current utilization rates, the industry is effectively writing off hundreds of billions in idle capacity. That's a difficult number to explain to a board.

Expect rapid adoption of rightsizing tools, GPU fractioning, and pre-submission prediction layers in the near term. The stabilization of Kubernetes DRA in version 1.36, released in April, provides a standard interface for vendors and open-source projects to build against. Nvidia's donation of a DRA driver in March signals the GPU incumbent is betting on standardized orchestration rather than proprietary lock-in for workload management. Tools like Expanse that integrate at the pre-admission stage should see increased uptake as enterprises look to cut queue times and reduce over-requesting without rewriting job submission workflows from scratch.

Multi-cluster schedulers and token-level pooling strategies, pioneered by hyperscalers like Alibaba, are likely to filter down to neoclouds and large enterprises. CoreWeave's rapid expansion—from 1 gigawatt today to a planned 5 gigawatts by decade's end—will only amplify the need for cross-cluster orchestration to avoid stranded capacity. Projects like Armada, which enable workload routing across geographically distributed Kubernetes clusters, address fragmentation by treating multiple sites as a single logical resource pool. It's the kind of solution that makes more sense the bigger you get.

Regulatory dynamics will reinforce the utilization imperative. The EU's Energy Efficiency Directive requires transparent reporting on data center energy use, with annual submissions due every May. U.S. state and local moratoriums on data center construction, documented by Visual Capitalist and others through May, push operators to justify expansions with clear efficiency metrics. Procurement policies are starting to favor operators that can demonstrate disciplined utilization over those simply racing to deploy raw capacity. "We have the most GPUs" is no longer the winning pitch it was eighteen months ago.

Export controls and geopolitical supply chain complexity add another layer of uncertainty. The U.S. government's licensing policy changes in January for chips like Nvidia's H200 and AMD's MI325X have created sourcing headaches. Companies that can extract more from existing hardware face less exposure to supply volatility than those relying on continuous hardware refreshes. That's an increasingly attractive position to be in.

The technical frontier is advancing, too. Nvidia's benchmarks with Run:ai show that dynamic fractioning and memory swap can nearly double utilization in the right configurations. Academic work on RL-based scheduling and malleable jobs suggests further headroom, though production deployment of those techniques remains limited. Alibaba's Aegaeon results, while from a controlled hyperscaler environment, hint at what becomes possible with token-level pooling at scale. But these are still early days.

Yet technology alone won't close the gap—it never does. The CNCF survey highlighted that cultural and process maturity, not technical capability, is the primary obstacle for many organizations deploying AI on Kubernetes. Teams need visibility, guidance, and tools that meet them where they are. Predictive systems that surface waste and suggest fixes without requiring users to rewrite job scripts or retrain on new platforms will likely see faster adoption than solutions demanding wholesale workflow overhauls. Humans are lazy; tools that respect that reality tend to win.

The question, then, is no longer whether the industry can build more GPU capacity. Power and permitting constraints have made that a long-lead, capital-intensive proposition fraught with local opposition. The question is whether the industry can optimize what it already has fast enough to keep pace with demand. At 5% utilization, there's more compute sitting idle in existing clusters than most companies will secure through new procurement in the next twelve months. That's not speculation. That's arithmetic.

The $725 billion buildout will proceed, barring some unforeseen collapse in the AI narrative. Grid constraints and moratoriums might slow it, but they won't stop it entirely. But the real unlock—the thing that might actually bridge the gap between what companies need and what they can get—may come from the billions already spent and largely unused.

Someone just needs to figure out how to turn it on.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Ardent Launches 6-Second Database Cloning for AI Agent Testing
  • From UX to AX: Why 40% of Enterprise Apps Are Redesigning for AI Agents
  • YC's Indexable Launches AI Agent Sandbox with Instant Environment Forks
  • YC-Backed ReasonBlocks Stops AI Agents From Burning Money Mid-Run
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.