Veer Shah and Suryaa Rajinikanth have known each other since they were kids. Now they're trying to upend how developers pay for the most expensive resource in artificial intelligence.
The pitch, in its simplest form, sounds almost too obvious: Why rent an entire GPU by the hour when you're only using it for minutes at a time? It's the equivalent of paying for a full tank of gas every time you start your car, whether you're driving to the grocery store or cross-country. And yet that's essentially how most GPU cloud providers bill their customers today.
Cumulus Labs, the company the two founded last year, went through Y Combinator with a serverless GPU platform called Cumulus Cloud. The promise: cold starts in 12.5 seconds, billing based on actual compute cycles consumed, and the ability to scale to zero when your workload goes quiet. For AI engineers juggling intermittent inference tasks or training jobs that spike unpredictably, it's a value proposition that might finally align costs with reality.
But whether it works at scale—and whether developers will trust aggregated GPU capacity over the known quantities of established cloud providers—remains very much an open question.
Idle Compute, Wasted Money
The problem Cumulus is tackling isn't new. GPU clusters routinely run at 30 to 40 percent utilization, according to the company's launch materials. The rest of the time? You're still paying.
Under the hood, Cumulus aggregates idle GPU capacity from various sources and routes workloads through what it calls "AI scheduling agents." The platform currently supports NVIDIA's A100 and H100 GPUs—the workhorses of modern machine learning—with optimizations handled by Ion, a proprietary inference engine the team built specifically for NVIDIA's Grace Hopper architectures.
In a February blog post, the company detailed Ion's performance on NVIDIA's GH200 chips: 7,167 tokens per second on Qwen2.5-7B, 588 tokens per second on Qwen3-VL-8B. Those numbers, Cumulus says, come from techniques including coherent CUDA graphs, eager key-value cache writeback, and something they call "phantom-tile" attention scheduling.
Technical jargon aside, the cold start claim is what catches attention. Twelve and a half seconds to spin up serverless GPU inference represents a meaningful improvement—if it holds up. The company achieved this through memory snapshots and torch.compile() optimizations.
Worth noting: these remain vendor-reported benchmarks, not independently verified results. In the GPU cloud business, performance claims are easy to make and harder to prove at scale.
Two Products, One Bet on Underutilized Hardware

Beyond the cloud platform, Cumulus recently rolled out Cumulus OS—a Kubernetes operator designed for organizations that already run their own GPU clusters. The software handles fleet monitoring, intelligent bin-packing, and priority scheduling. Standard fare for modern container orchestration.
The twist? What Cumulus calls "one-click spillover compute" to its liquid GPU marketplace. The company says organizations can sell their idle capacity back into the Cumulus network, turning underutilized infrastructure into—at least in theory—a potential revenue stream.
There's also IonRouter, a separate inference service that multiplexes what the company claims are more than 50 models on a single GPU with millisecond-level swap times. Pricing varies: popular models like GLM-5 run $1.20 input and $3.50 output per million tokens, while custom models bill on a per-GPU-second basis.
Whether enterprises will trust their spillover capacity to a startup marketplace, or whether they'll stick with the devil they know, is another matter entirely.
From Space Force to Serverless GPUs
Shah studied computer science at the University of Wisconsin-Madison. Before Cumulus, he worked with NASA and the U.S. Space Force through Small Business Innovation Research programs—government work that likely shaped his thinking about resource optimization.
Rajinikanth took a different path. Georgia Tech, then lead engineer at TensorDock, followed by infrastructure work at Palantir. In a LinkedIn post announcing their Y Combinator acceptance, he laid out the technical approach: multi-cloud training distribution, execution state capture for inference optimization, and a diagnostic runtime for troubleshooting GPU workloads.
The childhood friends founded Cumulus in 2026. They went through YC with just the two of them, which is either admirably lean or a sign of how capital-efficient the initial build was. Perhaps both.
A Crowded, Fragmented Market

Cumulus enters a GPU cloud landscape that's become increasingly messy. CoreWeave, which went from crypto mining operations to AI infrastructure darling, has scaled rapidly—its investor presentations outline roadmaps for GB200 NVL72 deployments and federal business expansion. Lambda Cloud maintains steady on-demand and reserved pricing for H100 and B200 clusters. Vast.ai operates a marketplace where H100 rates from verified hosts can occasionally dip below $2 per hour, though quality and uptime vary wildly by provider.
The company frames its differentiation around utilization economics. By charging based on physical GPU fraction consumed rather than reserved time, Cumulus attempts to align cost with actual usage—a model that makes intuitive sense but depends entirely on how well the scheduling intelligence works in practice.
Preemptive optimization, automatic checkpointing, job recovery: these aren't nice-to-haves for a platform routing workloads across aggregated capacity. They're essential. Any friction in those systems could easily negate whatever cost benefits the billing model provides.
What We Know, What We Don't
Cumulus has published extensive documentation covering training jobs with automatic checkpointing, inference workloads with batch and live server options, resource configuration parameters. The platform supports dependency auto-detection in submitted Python scripts, per-job VRAM and streaming multiprocessor hints, priority-based scheduling.
What's missing: named customers. Production scale metrics. Third-party validation of those performance claims.
The company lists its backers as Y Combinator and the NVIDIA Inception Program. The latter is worth contextualizing—it's a self-reported membership program that couldn't be independently verified through NVIDIA's public channels. No additional funding rounds beyond Y Combinator have been announced, which means the company is either bootstrapping aggressively or keeping its fundraising close to the vest.
The team has been active on the technical side, publishing blog posts detailing Grace Hopper-native inference optimizations, multi-VLM single-GPU serving, and support for models including Qwen3.5. The documentation suggests a platform ready for production use.
But documentation and production reality are two different things.
The Proof Is in the Running

For now, Cumulus represents another experiment in GPU cloud economics—one that bets developers will pay for flexibility and granular billing over the simplicity of flat hourly rates. It's a reasonable bet. In a market where compute costs can make or break an AI product, even modest efficiency gains compound quickly.
The question isn't whether the idea makes sense. It does. The question is whether a two-person startup can execute on the operational complexity required to make serverless GPU routing work reliably at scale, compete against well-funded incumbents, and convince developers to trust aggregated capacity over the known quantities of established clouds.
Shah and Rajinikanth have the technical chops. They have Y Combinator's backing. They have a pitch that resonates with anyone who's ever looked at their GPU bill and wondered why they're paying for idle time.
What they don't have yet—at least not publicly—is proof that it works beyond the benchmarks and the blog posts. In the GPU cloud business, that's the only metric that ultimately matters.
