There's a peculiar anxiety that sets in when you're running AI workloads on someone else's hardware. You can pay through the nose to keep GPUs warm and ready, burning money while they sit idle between requests. Or you can scale to zero, save the cash, and then wait—sometimes minutes—while your model spins up from scratch each time a user hits your endpoint.
Neither option is particularly appealing. Which is why Cumulus Labs, a two-person outfit that emerged from Y Combinator's Winter 2026 batch, has planted its flag on what it claims is a third way.
The San Francisco startup went live in late January with a serverless GPU platform promising something the industry has been chasing for a while: sub-15-second cold starts, with billing tied exclusively to active compute. No idle charges. No waiting around. Just fast spinup times and pay-per-cycle pricing that supposedly cuts costs by 50% to 70%.
It's an audacious pitch in a space where the laws of physics—and economics—haven't been especially forgiving. But if Cumulus can deliver, it might actually matter.
The Problem Everyone Knows About
Ask any developer running inference workloads at scale about cold starts and you'll hear the same weary stories. Baseten's own documentation essentially admits defeat, warning that "for large models, cold starts can take minutes" and suggesting you just leave minimum replicas running to sidestep the issue entirely. RunPod sells "zero cold starts" by keeping instances pre-warmed, which sounds great until you realize those workers stay on the meter whether you're using them or not.
Together AI, Koyeb, and others have nibbled around the edges of the problem with per-second billing and various optimizations. But the fundamental tension remains: fast or cheap, pick one.
Cumulus—founded by childhood friends Suryaa Rajinikanth and Veer Shah—believes the answer lies in how models are loaded and managed at the chip level. Their platform supports both training and inference on NVIDIA's A100 and H100 GPUs, with autoscaling that can spin up from one to over 100 replicas. The company says onboarding requires roughly 20 lines of configuration, though that's the kind of claim that probably varies wildly depending on your existing stack.
For training jobs, there's automatic checkpointing and resumption if you get preempted—a feature that matters more than it might sound, given how often spot-like infrastructure can yank the rug out from under you.
The Ion Engine (No, Not That Kind)
The technical heart of the operation is something Cumulus calls "Ion," or occasionally "ionattention"—a proprietary inference engine that purportedly keeps more than 50 models "ready" on a single chip. The company published benchmarks in mid-February comparing its performance to Together AI across several vision-language models, claiming it can run five VLMs simultaneously on one GPU without collapsing under the weight.
According to those benchmarks, Cumulus claims to match Together AI's per-token pricing for popular models and offers per-GPU-second billing for custom or unsupported models. For custom or fine-tuned models, that per-second billing—rather than per hour—is a distinction that only becomes meaningful if you actually believe the 12.5-second cold start figure.
There are caveats, naturally. The company's own blog post acknowledges that at high concurrency, first-token latency can lag a dedicated H100 by one to two seconds. That's the price of multi-tenancy, the trade-off that makes the whole system work. Whether that latency penalty is acceptable depends entirely on your use case, and Cumulus isn't pretending otherwise.
Show Me the Numbers

Cumulus illustrates potential savings with a chart showing "45% billed, 55% idle—free," though a comprehensive public rate card remains elusive as of this writing. The comparison points are out there, at least: Together AI charges between $2.99 and $5.50 per hour depending on whether you're running H100s or the newer B200 chips. Baseten's per-minute rates start at about $0.06667 for an A100 with 80GB of memory.
Both claim they don't bill for idle time when scaling to zero. But as Cumulus sees it, the cold start penalty effectively forces you to over-provision anyway—so you end up paying either way.
The savings thesis hinges on that inefficiency being widespread enough that even a 50% reduction adds up to real money at scale. That's plausible, maybe even likely. But it's also the kind of claim that sounds better in a pitch deck than in production, where edge cases and unexpected load patterns have a way of surfacing.
Two Founders, No Customers (Yet)
Rajinikanth, who studied computer science at Georgia Tech, previously worked as lead engineer at TensorDock, a distributed GPU marketplace, before stints at Blackstone and Palantir. Shah graduated from the University of Wisconsin-Madison in December with a CS degree, having spent time at an aerospace startup and leading a Space Force SBIR contract for military satellite communications. He also worked on several NASA SBIR programs and, before college, captained a FIRST Robotics team.
The company incorporated in California on December 18 and went through Y Combinator's Winter 2026 batch. Its homepage displays an NVIDIA Inception Program badge.
What's missing, for now, is customers—or at least public ones. Cumulus is operating behind a waitlist and demo request system. There's no self-serve signup, no published case studies, no customer logos. Documentation mentions A100 and H100 availability. It's hard to say what the full hardware mix looks like at this stage.
The company also mentions something called "Cumulus OS," an on-premises offering, and notes that Ion is "coming soon"—though "soon" is doing some heavy lifting there given how early-stage everything appears to be.
Can They Pull It Off?

Here's the thing about specific performance claims: they're either true or they're not. The 12.5-second cold start benchmark that Cumulus advertises is concrete enough to be independently verified once developers start running actual workloads, though real-world results may vary based on specific conditions. The per-cycle billing model is similarly straightforward to test. Either the savings materialize or they don't.
What remains to be seen is whether Cumulus can maintain those numbers under real-world conditions—sustained high concurrency, diverse model types, the kind of unpredictable traffic patterns that characterize production AI applications. Lab benchmarks are one thing. Production is another.
The serverless GPU market is getting crowded, fast. Bigger players with deeper pockets and existing customer relationships are all chasing some version of this same problem. Cumulus has a narrow window to prove its technology works, sign customers, and build enough momentum that it doesn't get crushed by competitors with more resources.
For developers exhausted by the cold start versus idle cost dilemma, that window might be just wide enough. The question is whether two founders and a clever inference engine can move quickly enough to wedge it open before the market moves on.
