A San Francisco startup barely old enough to celebrate its first anniversary claims it can do what typically takes engineers six weeks: turn a warehouse full of bare-metal GPU servers into a production-ready Kubernetes cluster in under two days.
Aranya announced on September 1 it raised $11 million across seed and pre-seed rounds led by First Round Capital and Asylum Ventures, betting that AI companies desperate to deploy inference workloads will pay handsomely for orchestration software that eliminates the manual slog of configuring heterogeneous hardware. The company says it manages more than $500 million worth of GPUs, a company-provided metric without independent verification, though it founded only in 2025.
First Round Capital led the $9 million seed round, joined by Box Group, Vermilion Cliffs Ventures and Asylum Ventures. Asylum Ventures had previously led a $2 million pre-seed that included Founder Collective, Parable VC and Uncommon Ventures. An SEC filing from July listed 11 investors contributing $11.31 million, with an initial sale dated June 25.
The pitch is straightforward: companies acquiring GPU capacity no longer have time to waste on infrastructure headaches. Aranya's core product, clusterdOS, automates the orchestration layer that typically requires teams of DevOps engineers to babysit. The open-source, Kubernetes-native system handles everything from driver installation and kernel setup to multi-cluster federation, stitching together tools like Ceph, Cilium, Traefik, vLLM, ArgoCD and KubeVirt.
"We've matched that 48-hour timeline repeatedly, and no other operator does it consistently for custom architecture," CEO and co-founder Christian Bhatia Ondaatje told SiliconANGLE.
The company also offers a proprietary interface with natural-language controls for spinning up inference endpoints, adding nodes and catching GPU faults before they cascade into outages. Those faults might include ECC memory errors or thermal throttling, the kind of issues that can cripple production workloads if left undetected.

Aranya emerged from stealth on April 28 with a partnership already in hand: Hydra Host, a GPU infrastructure provider that later raised a $100 million Series A in June. At the time, the startup said it had deployed across more than 1,700 GPUs. By September, it claimed to be managing hundreds of millions of dollars in compute for unnamed inference providers.
One joint customer running more than 1,700 GPUs saw setup time collapse from six weeks to under 48 hours, according to company materials and a LinkedIn post from Hydra Host in August. The same deployment experienced a 90 percent reduction in cluster outages. Ondaatje wrote separately on LinkedIn that one of the world's three largest inference providers is running production workloads on Aranya, though the company declined to name the customer.
Ondaatje co-founded Aranya with Sasivarnan Kanaghasalam Sathyapriya and Aryamika Bhatia Ondaatje. He says he joined Crusoe as the third engineer seven years ago and has spent a decade in distributed GPU infrastructure, credentials that matter in a market where credibility still counts for something.

The funding will expand engineering and go-to-market teams. Job postings on Standout list openings for Kubernetes DevOps engineers and design engineers at salaries between $150,000 and $200,000 plus equity. LinkedIn showed the company employed between two and 10 people as of early September. Aranya also plans to launch what it calls an AI-native multi-cluster interface, though details remain vague.
The urgency around inference orchestration reflects a broader shift in how the AI industry spends money. GPU clouds and bare-metal providers are racing to capture inference workloads as the focus pivots away from training-dominated budgets. Gartner projected late last year that 55 percent of AI-optimized infrastructure-as-a-service spending in 2026 would support inference, climbing above 65 percent by 2029. A VentureBeat survey from August found 66 percent of enterprises running AI workloads in production, using an average of three infrastructure platforms. That fragmentation creates demand for orchestration layers capable of managing fleets cobbled together from different vendors.
Hydra Host, CoreWeave, Lambda and Voltage Park all compete for GPU capacity. Hydra Host positions itself as an NVIDIA Cloud Partner focused on so-called AI factory deployments. Lambda raised more than $1.5 billion late last year to expand what it calls superintelligence cloud infrastructure, a term that may or may not survive contact with reality.

"In less than a year, Aranya is already managing hundreds of millions of dollars in GPUs for some of the most demanding inference workloads in AI," Todd Jackson, a partner at First Round Capital, said in the press release. Ondaatje, for his part, said the company aims to deliver an accessibility jump for clusters comparable to what personal computers achieved in the nineties.
Whether a 48-hour deployment window will prove durable as a competitive advantage remains to be seen. But for now, at least, Aranya has convinced investors that speed matters more than ever.
