Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 3, 2026

DesignVerse raises $5.5M to automate enterprise software

DesignVerse raises $5.5M to automate enterprise software
Ai AutomationEnterprise Software+3
SaaS iconSaaSOctober 3, 2026

OSCP raises $6M for GPS-free navigation sensors

OSCP raises $6M for GPS-free navigation sensors
PhotonicsSensor Tech+3
SaaS iconSaaSMarch 28, 2026

YC's Sentrial Launches 'Datadog for AI Agents' to Catch Production Failures

YC's Sentrial Launches 'Datadog for AI Agents' to Catch Production Failures
YcAi Agents+3
Fintech iconFintechMarch 28, 2026

YC-Backed FullSeam Launches AI Employee to Automate Finance Tasks

YC-Backed FullSeam Launches AI Employee to Automate Finance Tasks
YcAi Agents+3
SaaS iconSaaS
March 28, 2026
YcGpu CloudAi InfrastructureCloud Infrastructure

YC's Cumulus Labs Launches GPU Cloud with Pay-Per-Second Pricing

Y Combinator Winter 2026 startup debuts performance-optimized GPU cloud platform with novel pay-per-use billing, claiming 50-70% cost savings and 12.5-second cold starts for AI workloads.

YC's Cumulus Labs Launches GPU Cloud with Pay-Per-Second Pricing

Every hour, somewhere in a cloud data center, a high-end NVIDIA GPU sits idle—or mostly idle—while its user pays the full freight. It's the open secret of AI infrastructure: companies reserve compute capacity they don't fully use, then absorb the cost because switching instances on and off takes too long to be practical.

Cumulus Labs thinks it has a better answer. The San Francisco startup, which graduated from Y Combinator's Winter 2026 cohort and emerged from stealth in early 2026, has built what it calls a serverless GPU platform that bills by the second for actual cycles consumed. No compute, no charge. The company claims this approach can slash GPU bills by half or more, provided your workload isn't running flat-out around the clock.

It's an appealing pitch in a market where compute costs are climbing and developers increasingly scrutinize every line item. Whether Cumulus can deliver on that promise—and whether its serverless model can compete with the scale and stability of entrenched players—remains an open question.

Paying for Idle Time

The premise is simple enough. Most GPU cloud platforms charge by the hour for reserved instances. Spin up an A100, and you're on the hook for 60 minutes whether you use it for six minutes or sixty. According to Cumulus, typical utilization rates hover somewhere between 15% and 30% for dedicated GPU workloads, meaning developers are effectively overpaying by a factor of three or four.

Cumulus's model flips that equation. Developers submit PyTorch training scripts or deploy inference servers through a Python SDK, and the platform handles the rest—scheduling, scaling, GPU allocation across a pool of NVIDIA A100 and H100 hardware. But instead of charging for reserved capacity, Cumulus bills for what it calls "actual GPU compute used." If your model sits idle for half the hour, you pay for 30 minutes.

The company illustrates this with a straightforward example: an hour of GPU time listed at $0.60 would cost just $0.27 if the hardware was only active 45% of the time. Scale to zero, and the bill goes to zero.

That kind of granular billing isn't entirely new—serverless compute has been around for years in traditional cloud infrastructure. But GPUs have resisted the serverless model, largely because cold starts (the time it takes to spin up a fresh instance) have been prohibitively slow for latency-sensitive workloads. Cumulus claims it has cracked that problem with 12.5-second cold starts, a figure the company positions as fast enough to make serverless GPU inference practical for production use.

Fractional Allocation and Fast Spin-Ups

The technical trick lies in how Cumulus allocates and manages resources. Rather than handing users full GPU instances, the platform supports fractional allocation. Developers can request specific amounts of VRAM or compute capacity—using parameters like vram_gb or sm_percent (streaming multiprocessor utilization)—and pay only for what they claim during active computation.

For inference workloads, Cumulus captures the live execution state of models and replicates it across what the company describes as a "global compute CDN." This state capture is what enables the sub-15-second cold starts: the platform can spin up a new instance with the model already loaded in memory, rather than initializing from scratch. For teams that need even lower latency, Cumulus offers "warm in memory" options with sub-second time-to-first-token, though those presumably come at a premium.

The company also touts dynamic migration for training jobs. Workloads can be moved mid-run to different hardware based on availability and cost, with automatic checkpointing and resume if a job gets preempted. It's an architecture designed to aggregate compute "from everywhere," as the team put it in an early YC Launch post—big cloud providers, trusted data centers, individual hosts—into a unified pool.

Whether that vision of seamlessly migrating jobs across heterogeneous infrastructure holds up under production load is another matter. Training runs are notoriously finicky, and any disruption—network hiccups, hardware inconsistencies—can derail progress. Cumulus argues its checkpointing and failover systems handle this, but the proof will be in how well it works at scale.

The GH200 Play

Digital illustration for article section "The GH200 Play" in "YC's Cumulus Labs Launches GPU Cloud with Pay-Per-Second Pricing" - A professional, conceptual illustration of a sleek, minimalist central engine smoothly channeling a ...

In mid-March, Cumulus launched IonRouter, an inference platform built around what the company calls IonAttention, a proprietary engine optimized for NVIDIA's GH200 (Grace Hopper) architecture. The headline number: 7,167 tokens per second on a single GH200 running Qwen2.5-7B, according to a February blog post.

Cumulus also claims it can multiplex five vision-language models on one GH200 GPU while matching or exceeding throughput from competitors like Together AI, based on internal testing with 2,700 video clips and concurrent users. IonRouter offers an OpenAI-compatible API with per-token pricing for models including GLM-5, Kimi-K2.5, and Qwen3.5-122B-A10B.

The founders have been candid about positioning IonRouter as a showcase for the underlying engine rather than a standalone business. One of them noted on Hacker News that it's "mostly to showcase our engine"—a telling admission that suggests Cumulus sees its real value in the infrastructure layer, not necessarily in competing with inference API providers head-on.

Still, the performance claims are notable, if unverified. All the throughput numbers—7,167 tok/s, five VLMs on one GPU—come from the company itself, not independent benchmarks. Cold-start figures are similarly self-reported, though Cumulus isn't alone in this; Modal, a competitor in the serverless space, reported improvements down to around 10 seconds in late 2025 using GPU memory snapshotting.

Who's Behind It

Cumulus was founded by Suryaa Rajinikanth and Veer Shah, childhood friends who reunited after careers in infrastructure and aerospace. Rajinikanth worked on custom GPU compute at TensorDock before joining Palantir as a forward deployed engineer, building infrastructure for U.S. government clients. Shah led a Space Force program and worked on machine learning for an aerospace startup supporting NASA missions.

The technical depth shows. IonAttention's optimizations for GH200—coherent CUDA graphs, phantom-tile attention scheduling—read like the work of engineers who've spent real time in the guts of GPU architecture. In a LinkedIn post from February, Rajinikanth framed the company's mission around eliminating vendor lock-in and making GPU compute "more liquid," a vision that aligns with the platform's emphasis on job migration and distributed scheduling.

Y Combinator partner Jon Xu backed the team through the Winter 2026 batch. Cumulus is also part of NVIDIA's Inception program, though that's a membership designation rather than an equity investment. As of March, the team size was still just two, with no disclosed funding beyond YC.

Early Signals

Digital illustration for article section "Early Signals" in "YC's Cumulus Labs Launches GPU Cloud with Pay-Per-Second Pricing" - A clean, minimal, and conceptual illustration representing a game asset creation and training pipeli...

Cumulus has at least one public customer. In March, the company published a case study detailing work with CoSprite, a visual generation platform for game asset creation. According to the blog post, Cumulus built CoSprite's training pipeline using supervised fine-tuning and online reinforcement learning, with render success rates improving from 88% to 94% after SFT. The post describes a production GRPO (Group Relative Policy Optimization) loop running on Cumulus infrastructure, though it's unclear how long the engagement has lasted or how much compute CoSprite is actually consuming.

The company has also rolled out Cumulus OS, a Kubernetes operator that brings the same optimizations—VRAM sharing, cold-start acceleration, intelligent bin-packing—to on-premises clusters. The pitch includes "one-click spillover compute" (burst to Cumulus's public marketplace when on-prem capacity fills) and the ability to sell idle on-prem GPUs back into the marketplace during downtime. It's an ambitious play to turn enterprise infrastructure into a two-way revenue stream.

The Market Context

Cumulus is entering a GPU cloud market that's both crowded and consolidating. CoreWeave secured a $2 billion investment from NVIDIA in January to expand AI compute capacity, layered on top of billions in contractual commitments from Meta and OpenAI secured the year before. Lambda, Crusoe, and other so-called "neoclouds" have collectively raised hundreds of millions to build dedicated infrastructure for hyperscale training workloads.

But Cumulus isn't trying to compete on raw capacity or long-term contracts with foundation model labs. The company is targeting a different segment: developers who need flexibility more than they need guaranteed uptime. Teams running inference with spiky traffic. Researchers fine-tuning models without paying for 24/7 reservations. Startups prototyping before they scale to dedicated clusters.

It's a sensible niche, assuming the economics hold. The 50–70% cost savings claim hinges on workload characteristics—how much idle time you're actually eliminating, how Cumulus's per-second pricing compares to competitors at volume, and whether the platform's cold-start speeds and migration capabilities deliver in practice.

The Open Questions

Digital illustration for article section "The Open Questions" in "YC's Cumulus Labs Launches GPU Cloud with Pay-Per-Second Pricing" - A minimalist and conceptual illustration representing the open questions of execution, scaling, and ...

The big unknowns cluster around execution. Can Cumulus maintain sub-15-second cold starts as the platform scales? Will dynamic job migration work reliably across heterogeneous hardware, or will edge cases and network latency erode the efficiency gains? How does pricing shake out for teams that need consistent throughput rather than bursty workloads?

Then there's the question of trust. CoreWeave and Lambda have built reputations (and raised billions) by offering stable, predictable infrastructure for high-stakes AI workloads. Cumulus is asking developers to bet on a two-person team with a novel architecture and limited production history. That's a harder sell for mission-critical deployments, even if the cost savings are real.

But perhaps that's the point. Cumulus isn't positioning itself as the platform for training GPT-5. It's positioning itself as the smarter, cheaper option for everything else—the long tail of AI workloads that don't need the guarantees of a CoreWeave contract but could certainly use a lower bill.

The documentation is live, the platform is shipping, and the team has moved fast. Whether Cumulus can turn that momentum into a sustainable business—one that scales beyond early adopters and withstands competition from better-funded rivals—will depend on how well the technology performs when the stakes get higher.

For now, developers tired of paying for idle GPUs have another option. Whether it's the right option depends on what they're building, and how much risk they're willing to take on unproven infrastructure.

More stories

  • DesignVerse raises $5.5M to automate enterprise software
  • OSCP raises $6M for GPS-free navigation sensors
  • YC's Sentrial Launches 'Datadog for AI Agents' to Catch Production Failures
  • YC-Backed FullSeam Launches AI Employee to Automate Finance Tasks
  • YC's Lucent Uses AI to Auto-Detect Bugs in Session Replays
  • YC-Backed Stilta Launches AI Platform to Automate Patent Drafting
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.