Here's an inconvenient arithmetic problem for the AI boom: The semiconductor industry is racing toward $1.3 trillion in revenue by 2026, powered almost entirely by demand for AI accelerators. Global AI spending is projected to reach $2.59 trillion in 2026. And yet, according to data from tens of thousands of real-world Kubernetes clusters measured this past April, the average GPU sits idle 95% of the time.
Five percent. That's the baseline utilization figure before any optimization effort. Which means that for every dollar spent on the most expensive compute infrastructure in corporate history, ninety-five cents is evaporating into unused cycles.
This isn't a future problem that better software will eventually solve. It's happening now, eroding margins and forcing infrastructure teams into an exhausting game of resource arbitrage. The bottleneck, it turns out, has quietly shifted—from acquiring GPUs to actually using them. And the gap between what companies are spending on hardware and what they're getting out of it has never looked worse.
What the Data Shows
In April, Cast AI released findings from its State of Kubernetes Optimization report, a study covering tens of thousands of production clusters spanning AWS, Azure, and Google Cloud. That 5% GPU utilization number represents more than a missed opportunity; it's a structural inefficiency that compounds at every layer of the AI stack. For perspective, CPU utilization in those same environments averaged 8%. Memory hit 20%. GPUs, the crown jewels of the modern data center, lagged behind everything else.
Microsoft Research published a study at ICSE'24 that examined 400 internal deep learning jobs. All had GPU utilization of 50% or less. The researchers traced 46% of the issues to data pipeline operations—mundane stuff like batch sizes that don't align with GPU memory, data loaders that starve the accelerator, checkpoint strategies that idle entire clusters. Eighty-five percent of the problems, they found, could be fixed with modest code or script changes. Not exotic optimizations. Just basic operational hygiene.
The pattern holds up over time, too. Microsoft's older Philly cluster traces from 2019 showed mean GPU utilization hovering around 52% across job sizes. If anything, despite years of improvements in scheduler sophistication and workload orchestration, recent measurements suggest the situation has deteriorated.
Why It Matters Now

Three trends are converging to make GPU waste intolerable, and all of them have accelerated in the past year.
First: the economics have inverted. AWS raised prices for EC2 Capacity Blocks for ML by around 15% in several regions this January. Gartner, not typically prone to dramatic revisions, updated its 2026 AI spending forecast twice in four months—from $2.52 trillion in January to $2.59 trillion in May, a 47% year-over-year surge. When CoreWeave signed a $21 billion capacity agreement with Meta in April, followed the next day by a multi-year deal with Anthropic, the message to the broader market became impossible to ignore: GPU capacity is scarce, committed years in advance, and increasingly expensive.
Second: power is now the binding constraint. S&P Global reported in May that U.S. data center grid power to hyperscale, leased, and crypto operations grew 25% last year to approximately 64.4 gigawatts. Virginia remains the largest market, but the footprint is expanding everywhere. In March, the White House announced the Ratepayer Protection Pledge, securing commitments from Amazon, Google, Meta, Microsoft, OpenAI, Oracle, and xAI to build, bring, or buy their own generation capacity and cover delivery infrastructure costs. The subtext? Utilities and regulators won't let AI expansion shift costs onto households. Which means infrastructure teams must extract maximum value from every watt, or face constraints they can't negotiate around.
Third: operational complexity is scaling faster than the tools to manage it. Export controls on advanced semiconductors have injected planning uncertainty across GPU generations. NVIDIA's acquisition of Run:ai in late 2024 signaled the strategic importance of workload orchestration, but even sophisticated schedulers struggle when users systematically overestimate what their jobs need. The Cloud Native Computing Foundation published a scheduler plugin design in January for reclaiming underutilized GPUs through preemption, but the challenge sits upstream. Jobs arrive at the scheduler already misconfigured.
The Ecosystem Response

The industry's answer has been predictably fragmented, with each vendor addressing a different piece of the problem.
NVIDIA positions Run:ai and its KAI Scheduler as topology-aware orchestration for fairness and queuing, tightly coupled to NVLink and NVSwitch fabrics. Datadog announced general availability of GPU Monitoring in April, linking fleet health, cost, and performance to individual teams for troubleshooting and rightsizing. Cast AI added GPU utilization visibility to its Kubernetes FinOps platform around the same time. Scheduling, observability, cost allocation—each matters, but the root cause sits earlier in the workflow, before jobs even reach the scheduler.
That's the bet a four-person startup called Expanse is making. Based in San Francisco and part of Y Combinator's Spring 2026 batch, the company is attempting to solve GPU waste at submission time. Their platform predicts runtime, memory, CPU, and GPU requirements before a job hits the scheduler, flags likely failures, and suggests code or configuration changes. When failures do occur, their diagnostic tool returns root-cause analysis and the exact edits needed.
The architecture is deliberate. Expanse deploys on customer infrastructure, so training data and job metadata never leave the network. Integrations include Databricks, YARN, Cloud Batch, and Nomad—covering a broad swath of enterprise scheduling environments.
The company also launched a free utility called the OS Wastage Scanner that computes CPU, memory, and GPU waste locally on SLURM or Kubernetes clusters via a one-line script. It shows a utilization score, estimates cost, and maintains an opt-in leaderboard. The tool is, transparently, a lead generation mechanism. But it also surfaces something more fundamental: most teams don't know their baseline utilization until they measure it. And most don't measure it systematically.
Meta's engineering blog described their "Zoomer" performance tooling last November, emphasizing that "every percentage point of utilization improvement" yields material capacity gains at their scale. The post reflects what happens when GPU efficiency becomes a first-class operational metric, but building that capability in-house requires the kind of profiling infrastructure and ML ops depth that few organizations possess. Expanse and its peers are commercializing versions of that insight for teams that can't—or won't—invest in proprietary tooling.
Where This Goes

The trajectory, at least in the near term, points toward tighter integration between prediction, scheduling, and remediation. Research from late last year proposed dynamic multi-objective scheduling with size-aware batching that reached roughly 75% utilization in simulations. A paper published in May introduced Overall FLOP Utilization (OFU) as a fleet-scale metric combining on-chip tensor pipeline activity and SM clock data; the authors reported it caught 2.5x more performance regressions than traditional metrics. Another study from April examined execution-idle as a distinct state in GPU systems, targeting energy waste during non-productive periods.
Expect observability vendors and Kubernetes optimization platforms to converge over the next twelve to eighteen months, turning telemetry into automated rightsizing and remediation loops. The human role will shift from manual intervention to policy and guardrails. Submit-time prediction will likely become table stakes in job queuing systems like Kueue and Volcano. Scheduler plugins will handle preemption and reclamation with greater sophistication, though perhaps not as quickly as vendors promise.
For infrastructure leaders, the calculus is straightforward, if uncomfortable. If your fleet is averaging 5% utilization and you're locked into multi-year capacity commitments priced at current rates or higher, software that lifts effective utilization by even 10 percentage points is worth more than adding racks. The constraint is no longer chip availability. It's the discipline to measure, predict, and optimize before every job runs.
The companies that solve this operationally—that treat GPU efficiency as a first-order business problem rather than an engineering curiosity—will extract significantly more value from the same capital base. In a market where capacity is both scarce and expensive, that's the only competitive advantage that scales. The question is whether they'll solve it before the capital runs out.
