When large language models try to manipulate spreadsheets, something curious happens: they fail. A lot.
It's a problem documented in academic research over the past few years, and now a London-based startup called Witan Labs thinks it has built the infrastructure to solve it. The company has announced backing from Positive Sum, Angular Ventures, and LocalGlobe, according to investor logos displayed on its website, though no details on funding amounts or round timing have been disclosed.
The pitch is narrow but pointed: Witan is building what it calls "headless Excel for agents," a specialized engine designed to let AI systems read, edit, and calculate spreadsheets with fewer hallucinations and faster execution than the native tools shipped by OpenAI or Anthropic. Whether that bet pays off depends on a somewhat unglamorous question: Can you turn spreadsheet accuracy into venture-scale infrastructure?
The Money Suggests Someone Thinks So
Witan's backers aren't alone in betting that spreadsheets represent a chokepoint for AI adoption in enterprises. Earlier this year, a startup called Meridian raised $17 million at a reported $100 million post-money valuation for what it describes as an agentic spreadsheet IDE. Last year, OpenAI's venture arm led a $14 million round into Endex, which embeds AI agents directly inside Excel itself.
The flurry of activity reflects a stubborn reality: spreadsheets may be unglamorous, but they're also ubiquitous. And for all the talk of AI agents automating enterprise workflows, the models still struggle with basic formula generation and reasoning tasks when confronted with actual .xlsx files.
Academic papers published in recent years have cataloged the failures. Both Anthropic and OpenAI have introduced what they market as "Excel skills" for their models, but performance remains inconsistent enough that startups like Witan see an opening.
Vendor Claims, Real Gaps
Witan has run its own benchmarks—internally, on a dataset the company says contains 250 questions—and the results, at least as the startup presents them, suggest meaningful improvement. The company claims its engine achieves roughly 73% accuracy compared to around 59% for Anthropic's implementation and 62% for what it labels OpenAI's Excel capability. Execution speed, according to Witan, is also faster: 115 seconds at the 90th percentile versus 185 seconds for the Anthropic baseline.
These are vendor-reported numbers, self-reported by Witan with no independent replication published. But the gap Witan describes—even if the real-world delta is smaller—is apparently credible enough that the company dedicated a four-month engineering sprint to the problem. A public GitHub repository maintained by the team chronicles that work, with research logs spanning late last year into early this year.
Infrastructure, Not an App

Witan isn't trying to replace Excel or Google Sheets. It's building the layer underneath—a CLI-first toolchain that lets AI agents manipulate workbooks server-side without launching a graphical interface. The idea is to eliminate the fidelity loss and performance drag that comes from, say, running LibreOffice in headless mode.
The company's pitch centers on a concept it calls "WASIWYG"—what the agent sees is what you get. The engine renders specific cell ranges to avoid wasting vision tokens, returns structured text from calculations, and includes what Witan describes as a semantic linter (think ESLint, but for formulas). All of this happens server-side, allowing agents to "proof before they ship," as the website puts it.
Witan Labs Ltd was incorporated in the UK on October 3, 2025. Nuno F. Campos, who co-authored a book on LangChain and maintains a package on PyPI, appears connected to the project. LinkedIn lists the team size at somewhere between 2 and 10 employees, which means this is still a very early operation.
The company's product tiers, outlined recently, include a no-signup personal option, a free cloud API with rate limits, and an enterprise self-hosted version. No customers or partnerships have been publicly disclosed.
What Distinguishes This From Everything Else
Witan's infrastructure focus sets it apart from competitors like Meridian, which is building a full-fledged spreadsheet environment for financial modeling, or Sourcetable, which raised capital earlier this year for what it calls an autonomous spreadsheet interface. Witan is closer to plumbing—the kind of unglamorous but critical layer that other companies' agents might call.
Whether that positioning works depends on several things. First, do the benchmark improvements hold up under real-world conditions, with actual enterprise workloads? Second, will developers building AI agents adopt a CLI tool from a startup, or will they default to the native skills baked into Claude, GPT, or whatever model they're already using?
Then there's the architectural question. Witan's GitHub log suggests the team pivoted midway through development toward what it describes as a REPL-based execution model. That kind of iteration is normal for early-stage startups, but it also signals the product isn't fully settled yet.
The Bet Behind the Bet

Positive Sum, Angular Ventures, and LocalGlobe all specialize in enterprise infrastructure and early-stage deep tech. Their backing suggests they believe the spreadsheet-agent accuracy problem is real, and large enough to support a venture-scale business.
Witan now has investor backing to test that thesis. The company will need to prove not just that its engine is faster or more accurate—benchmarks are easy to game—but that enterprises building AI agents see enough value in the abstraction layer to adopt it.
That's a harder sell than it sounds. Integrating directly with OpenAI or Anthropic is straightforward. Convincing a developer to route through a third-party CLI requires the kind of performance gap that can't be dismissed.
For now, Witan is placing a very specific bet: that as AI agents proliferate, the companies building them will need better spreadsheet infrastructure. Whether that bet pays off—or whether model providers simply close the accuracy gap themselves—remains to be seen.
