The problem, as Akshay Budhkar sees it, isn't that AI agents can't do sophisticated work. It's that we can't see them doing it.
His startup, Spine AI, is wagering that the future of multi-agent collaboration looks less like a chat window and more like a whiteboard—one where every calculation, every delegation, every misstep is visible in real time. The company's latest product, Spine Swarm, treats AI orchestration as a spatial problem: agents work across a shared canvas made up of discrete blocks, each representing a task or artifact. Users can watch the work unfold, intervene when things drift, and steer without blowing up the entire workflow.
It's an unusual pitch in a market increasingly dominated by slick chatbots and opaque agent frameworks. Whether it gains traction remains an open question. But the underlying friction Spine is targeting—how do you actually manage a team of AI agents?—is real enough.
Show Your Work
Spine Swarm, which launched on Product Hunt on March 10, builds on the company's earlier product, Spine Canvas. The core mechanic is straightforward: instead of conversing with a bot or watching agents manipulate files in some hidden directory, users arrange workflows visually. Blocks—Chat, Deep Research, Memo, Image, Slides—serve as modular units. Agents spawn new blocks, pass context between them via explicit links, and execute tasks in parallel.
Every step is editable. If an agent veers off course mid-run, you can intervene. The canvas doubles as an audit trail, showing exactly how an agent reasoned through a problem and what sources it leaned on. Ashwin V. Raman, a co-founder who previously worked at NVIDIA, Instacart, and Amazon, describes this as "swarm intelligence"—not replacing human judgment, but making it easier to manage distributed AI labor.
The company supports more than 300 models. Users can assign different models to different steps and swap them mid-execution. This multi-model orchestration—combined with what Spine calls "tiered orchestration" and "dynamic ensembling"—appears central to the platform's design.
Bold Claims, Pending Scrutiny
Spine has made a provocative assertion: that its platform outperforms Perplexity, Google's Gemini Deep Research, and OpenAI's models on complex research benchmarks.
According to company-reported numbers from late February, Spine scored 61.5% accuracy on GAIA Level 3, a notoriously difficult question-answering benchmark. (The company adjusted that to 76.0% after filtering what it deemed "ambiguous or mislabeled" questions.) By comparison, it cited Genspark at 58.8%, Manus at 57.7%, and OpenAI Deep Research at 47.6%.
On DeepSearchQA—a newer benchmark from Google DeepMind introduced in January—Spine reported 87.6% across the full 900-prompt set. The company positioned this above Perplexity's 79.5%, Claude Opus 4.5 at 76.1%, and Gemini Deep Research at 66.1%.
These numbers have not been independently reproduced. Whether they hold up under outside scrutiny remains to be seen. Benchmark gaming is a well-worn tradition in AI circles, and self-reported figures should be approached with caution. Still, the underlying product addresses something tangible: the need for transparency when agents start doing complex, multi-step work.
How the Mechanics Play Out

Each block on the canvas represents a discrete task. A Deep Research block might pull data from the web, analyze it, and spit out findings. A Memo block synthesizes those findings into something structured. Agents pass context explicitly via visible links rather than relying on some internal black box.
Users can branch conversations, explore alternative paths, and run parallel analyses without torching existing work. Spine calls this "safe branching"—a design choice that matters when workflows involve iteration and judgment calls. The company targets use cases like competitive research, strategy development, pitch decks, and technical documentation. Its pricing page estimates a comprehensive competitive analysis at roughly 5,000 credits; an investor pitch deck at around 6,000.
CEO Budhkar, who previously led generative AI initiatives at Georgian Partners and worked at YC alumni Scribd and FarmLogs, frames the problem as "context blindness"—agents losing track of what they're doing as tasks sprawl. The canvas, in theory, prevents that.
Crowded Territory

The timing is tricky. Spine went through Y Combinator's Summer 2023 batch and operates with a team of eight in San Francisco (as of the March launch). But the landscape has shifted considerably since then.
OpenAI released AgentKit in October 2025, complete with a visual Agent Builder for creating and versioning multi-agent workflows. Its native canvas offers similar orchestration inside the OpenAI ecosystem. Salesforce has been rolling out Agentforce, autonomous agents deeply integrated with MuleSoft, Slack, and Data Cloud. Asana announced AI Teammates last fall. Perplexity launched its freemium Deep Research product early last year.
In other words, Spine is competing in a space where the model providers themselves are building native orchestration tools. The bet hinges on whether independent orchestration—full model flexibility, transparent auditability—offers enough differentiation from vertically integrated alternatives.
There's also the open-source question. Developer frameworks like LangGraph, CrewAI, and AutoGen offer programmatic control over multi-agent systems. Spine's wager is that visual orchestration lowers the barrier for non-engineers while preserving the flexibility engineers want. Maybe. Or maybe developers prefer code and non-technical users prefer simpler, more opinionated interfaces.
The Price of Entry

Spine uses a credit-based pricing model. The free tier includes 2,000 credits per month with access to all 300+ models and unlimited canvases. Pro, at $16 monthly when billed annually, offers 30,000 credits plus a daily refresh of 1,500. The Max tier—$80 per month annually—includes 200,000 credits, a 5,000-credit daily refresh, and priority support. Early users received a 3,000-credit bonus.
The product is live, with documentation at docs.getspine.ai and access at swarm.getspine.ai. Initial reception on Product Hunt was solid if not overwhelming: #9 for the day with 149 upvotes.
What Actually Matters
Benchmark numbers make for good headlines. But whether Spine Swarm gains traction will depend less on leaderboard rankings and more on whether the visual canvas model resonates with teams actually building agent workflows.
The company has laid out specific use cases: battlecards, strategy rooms, content systems, learning hubs. These are collaborative, iterative workflows where seeing the work unfold might actually matter. If agents are going to handle more of that work, perhaps the interface needs to show the reasoning, not just the result.
Human-in-the-loop isn't a feature here so much as the entire architecture. The canvas isn't just for watching agents work—it's for working alongside them, steering when necessary, auditing when skeptical.
That might matter more than any single benchmark score. Then again, it might not. The AI tooling market is littered with elegant interfaces that solved problems users didn't know they had. Spine's bet is that transparency and control—seeing your AI team think—is a problem people do have, even if they haven't articulated it yet.
Time, as ever, will tell.
