The problem with enterprise AI agents isn't that companies don't want them. According to Gartner's Hype Cycle for Agentic AI, steep near-term adoption intent is evident among tech executives. The problem is getting them to work.
Current adoption hovers at just 17%. Between ambition and execution sits what might generously be called a reliability gap—and a dawning recognition that the industry has been measuring the wrong things. When your AI agent aces a synthetic benchmark but can't successfully book a calendar meeting across three corporate systems, something fundamental is broken. And it probably isn't the model.
Ooak Data, a Paris-and-Delaware startup that joined Y Combinator's Summer 2026 batch, is betting that gap represents an opportunity. The company is building Alexandria—a library of anonymized, real-world enterprise workflows designed to train and evaluate AI agents before they encounter actual employees, actual data, actual consequences. The pitch sounds almost obvious in hindsight: if you want agents that function in Slack, Jira, Notion, and SharePoint, you need training data drawn from those environments. Not synthetic proxies. Not academic datasets. The messy, multi-step reality of how work actually happens.
Whether that thesis holds—and whether Ooak can execute on it before larger players crowd the space—remains an open question. Y Combinator's Demo Day will offer the first public signal. Until then, the company's existence mostly underscores how far the enterprise AI agent market still has to go.
When Benchmarks Stop Telling the Truth
Ooak Data laid out its core argument in a March 2026 research post that didn't mince words: synthetic benchmarks are saturated, contaminated, and disconnected from the operational realities of enterprise software. The academic community has been circling the same conclusion, if more cautiously.
In January 2026, researchers released APEX-Agents, a benchmark built from investment banking, consulting, and law firm tasks that emphasizes long-horizon workflows spanning multiple applications. ServiceNow's WorkArena and WorkArena++ benchmarks, constructed around actual enterprise software tasks, revealed uncomfortable gaps between what agents promise and what they deliver end-to-end. Mind2Web and its online variants track over 2,000 tasks across 130 real websites. Most agents still score below 50% success in realistic settings.
Then came EnterpriseClawBench in June 2026: 852 reproducible tasks extracted from real enterprise agent sessions. The researchers couldn't release the raw data, which rather proved the point. Real workflows contain sensitive business context—customer names, contract terms, internal communications. You can't just publish them. You also can't train reliable agents without them.
The Catch-22 has opened a lane for companies willing to build the infrastructure that solves it.
Digital Twins and the Anonymization Challenge
Ooak Data's approach revolves around what it calls "digital twins"—anonymized, multimodal representations of enterprise work. The company sources real enterprise data, strips identifying information, and reconstructs it as reinforcement learning environments with multi-step, multi-tool tasks. The target customers: frontier AI labs, enterprise AI teams, and startups "building agents, not chatbots," in the company's phrasing.
Ooak operates through two entities—Ooak Data Inc. in Delaware and Ooak Data SAS in Paris—a dual structure that may prove strategic as the EU AI Act comes into force. General application started August 2, 2026, with transitional deadlines extending through 2027 and 2028. Data minimization and de-identification requirements will directly affect anyone training models on enterprise workflows, which is to say: everyone trying to build production-ready agents.
Founder Pierre-Louis Vouteau, alongside co-founders Thomas Aubry (previously at Samsung AI and PayLead) and Grégoire L. (formerly at Epsor), has positioned Alexandria as infrastructure for what the company calls the "agentic AI stack." LinkedIn posts from July 2026 referenced an "APEX Explorer" tool for comparing agent performance across office workflows, evaluating Claude, Gemini, GPT, and other model families. The company describes itself as an applied AI research lab—pointedly, one built for agents rather than chatbots.
The distinction matters less as marketing copy than as a signal of where Ooak thinks the real demand lives. Chatbots answer questions. Agents are supposed to do things. The difference is everything.
The Enterprise Platform Race (And Why Governance Shipped First)

While startups like Ooak build evaluation infrastructure, enterprise platforms have been racing to ship production agent features. Microsoft announced computer-using agents in Copilot Studio on May 13, 2026, and launched its Agent 365 governance plane alongside the Foundry Agent Service. KPMG disclosed in June that it's scaling Agent 365 globally for audit and client deployments—a vote of confidence, or at least a bet that enterprise customers will demand agent capabilities whether or not they're fully baked.
Salesforce has been productizing Agentforce with monthly feature updates through mid-2026, including outcomes-based pricing for customer service agents. ServiceNow went further. The company launched its Autonomous Workforce in February and expanded its AI Control Tower in April to discover and govern every AI system, agent, and workflow across an enterprise—regardless of where it runs.
The pattern is consistent: vendors are shipping governance, observability, and evaluation tools faster than raw agent capabilities. That's not an accident. It's a tell.
Enterprise buyers aren't asking whether agents will break. They're asking how to contain the damage when they do.
The Reliability Gate No One's Quite Cleared
Benchmarks keep surfacing the same issue, just from different angles. UI-CUBE, published in November 2025, stressed operational reliability beyond single-click accuracy and exposed brittleness across full workflows. A June 2026 framework for evaluating agentic skills at scale tested 19 agent-model configurations and found skill-level performance wildly inconsistent—excellent at one task, catastrophically bad at a nearly identical one. An August 2026 paper on architectural implications of agentic AI workflows characterized agent workloads as crossing LLM inference, tool orchestration, and compute boundaries in ways current infrastructure simply wasn't built to handle.
ServiceNow claimed in February 2026 that it autonomously handles 90% of its own L1 IT tickets internally. Salesforce reported that its agent on Salesforce.com handled over 100,000 conversations since February 2025. These are vendor-reported success stories, carefully chosen. They also suggest the gap between controlled internal deployments and general-purpose agents remains wider than the marketing materials care to admit.
The challenge isn't purely technical. Security researchers at Lakera and the Cloud Security Alliance have cataloged a rising wave of prompt-injection threats as agents consume untrusted documents and web content. Studies from May and June 2026 explored automated injection attacks in agentic settings and proposed context-aware defenses, though with varying degrees of success. Regulations are catching up too. NIST's AI Risk Management Framework and ISO/IEC 42001 are showing up in enterprise RFPs, demanding traceability and auditable AI governance before the first agent touches production.
The Race for Real Data (And Real Differentiation)

Ooak Data isn't alone in trying to corner real-world agent datasets. Mercor positions its APEX-Agents collection as license-ready, peer-reviewed data that's "hard to game." BrowserGym Foundry offers commercial datasets and simulations for enterprise web agents with expert annotations. Third Origin is building multimodal human workflows for robotics and embodied agents. Alexandria.so—confusingly, a different company with a nearly identical name—captures actions and intents across employee workflows to build institutional "experience" layers.
The differentiation comes down to anonymization pipelines, environment fidelity, and task diversity. Ooak's site explicitly lists Slack, Gmail, Notion, Jira, SharePoint, and Teams as target environments and emphasizes reinforcement learning compatibility. The bet is that frontier labs and enterprise teams will pay for curated, multimodal, evolving task libraries over static academic benchmarks.
It's a reasonable bet. Whether Ooak can execute on it before competitors with deeper pockets move in is another question entirely.
What the Next Eighteen Months Look Like

The market is moving fast, perhaps faster than the underlying technology. Grand View Research estimates the enterprise agentic AI market reached $2.6 billion in 2024, projects $5.3 billion in 2026, and forecasts $24.5 billion by 2030—a 46% compound annual growth rate. IDC predicts roughly half of enterprises will use AI agents by 2027, with agentic capabilities embedded across more than 40% of enterprise applications.
But IDC also emphasizes data foundations and governance as the adoption gate, which is another way of saying: don't confuse aspiration with deployment. LangChain's June 2026 State of Agent Engineering report found organizations in production increasingly formalizing evaluations and adopting frameworks like LangGraph for long-running, stateful processes. Databricks updated its agent documentation on July 24, 2026, with end-to-end support for building, evaluating, and deploying agents on the lakehouse. Automation Anywhere's 2026 platform enhancements include design-time and runtime AI evaluations, with a Context Intelligence Graph targeted for Q3 2026.
The pattern suggests a normalization around agent evaluation is underway. Success won't be measured by benchmark leaderboard position but by operational reliability metrics: repeatability, guardrail adherence, cost per completed task. High-value workflow libraries—capturing document formats, app states, tool calls, and communication patterns—may become proprietary differentiators for both model training and go-live evaluation.
If that thesis holds, Ooak's timing could be right.
The Unanswered Questions
Ooak Data is early. The company hasn't disclosed customers publicly as of late July 2026, and funding details remain limited to Y Combinator's standard investment. Demo Day will offer the first real signal of enterprise traction—or lack thereof. But the timing may work in Ooak's favor. If agents are moving from pilots to production at the scale IDC projects, demand for trusted evaluation infrastructure could accelerate faster than the models themselves. Especially infrastructure that satisfies EU data protection and AI Act obligations, which is not a trivial consideration.
The question isn't whether enterprises will deploy agents. Gartner's Hype Cycle makes clear the intent is there. The question is whether they'll work reliably enough to justify the operational risk—and whether the data libraries that enable that reliability become the next AI infrastructure moat.
Ooak Data is betting yes. By September, we'll have a better sense of whether the market agrees.
