The demonstrations looked flawless. Then came production.
It's a storyline that has unfolded over the past year: AI agents that post impressive scores on standardized benchmarks stumble badly when turned loose on the messy, multi-step workflows that define actual business operations. The culprit, according to a growing chorus of researchers and founders, isn't the models—it's what those models learned from.
While the AI industry sprinted to build ever-larger models, a more mundane question lingered in the background. Where, exactly, do you find the kind of realistic business workflow data that could teach an agent how companies truly operate?
That gap has given birth to a new category of infrastructure startups racing to capture, anonymize, and package real company workflows into datasets that might actually prepare agents for the wild. Among them: Ooak Data, a five-person team that emerged from Y Combinator with an ambitious pitch—build what they're calling "the world's largest library of real-world business workflow datasets." Based in Paris and incorporated on September 9, 2024, by Thomas Aubry, Grégoire Lamy, and Pierre-Louis Vouteau, the startup is betting that synthetic data and simulations aren't enough.
They're hardly alone in that conviction.
When Benchmarks Meet Reality
The numbers suggest companies aren't waiting around for perfection. According to Gartner's April 2026 research, 17% of organizations have already deployed AI agents in production, with more than 60% expecting to follow suit within two years. In an August 2025 press release, Gartner had projected that 40% of enterprise applications would feature task-specific agents by the end of 2026, up from less than 5% in 2025.
But deployment is one thing. Reliability? That's where things get complicated.
Consider OSWorld 2.0, a benchmark released on June 28, 2026, that tests long-horizon computer use across 108 workflows requiring up to 500 steps. The best-performing model managed just 20.6% task completion. These aren't exotic edge cases, either—they're the kind of cross-system, multi-tool workflows that employees handle daily.
The disconnect between lab performance and field results has become hard to ignore. Forrester noted in a June 2026 analysis that "companies are chasing, few are catching" when it comes to agentic AI. IDC went further, forecasting that companies could face a 15% productivity loss by 2027 without proper AI-ready data foundations. (Whether that specific prediction holds up remains to be seen, but the sentiment reflects widespread anxiety about the current state of agent reliability.)
Meanwhile, the infrastructure has arrived. Microsoft shipped computer-using agents as a generally available feature in Copilot Studio in May. Google launched its "Agentic RAG" framework for enterprise agents in June. ServiceNow, Salesforce, and a parade of other platforms have rolled out agent capabilities.
What's missing is training data that reflects how businesses actually work—with all the exception handling, legacy system quirks, and undocumented workarounds that define real operations.
Three Converging Pressures

Several forces have aligned to turn this data gap into a market opportunity.
First, there's the technical reality: synthetic benchmarks don't reliably predict production performance. Ooak Data's research page puts it plainly, focusing on "the gap between benchmark performance and real-world capability." Snorkel AI, another player in this space, reported that its domain-expert simulated environments improved insurance underwriting agent performance from 10.9% to 42.0% pass@1—meaningful progress, sure, but still hitting a ceiling when working with simulated rather than actual company workflows.
Second, regulatory pressure is accelerating demand for datasets with clear provenance. The EU AI Act's general-purpose AI obligations began in August 2025, with most rules applying from August 2026, requiring training data summaries, copyright policies, and data governance documentation. Article 10 mandates appropriate data governance for high-risk AI systems. Companies building agents for European deployment need datasets with documented lineage and privacy controls—precisely what anonymized real-world workflow data promises to provide.
Third—and perhaps most telling—the process intelligence market is booming. ResearchAndMarkets estimates the process mining sector reached $3.82 billion last year, growing to $4.64 billion this year, with projections hitting $15.20 billion by 2032. UiPath, a bellwether for the automation platform category, reported $1.43 billion in revenue for fiscal 2025. These companies generate enormous workflow event logs—data that could theoretically feed agent training, if properly structured and anonymized.
The Digital Twin Approach

Ooak Data's strategy centers on what they call "digital twins"—anonymized replicas of entire company data ecosystems. Their pipeline sources real enterprise data, anonymizes it, then generates reinforcement learning environments with expert-level, multi-step, multi-tool tasks. The company frames its offering as "built for agents, not chatbots," and has stated plans to acquire and anonymize over 300 full data ecosystems from real companies within six months.
Whether they can execute on that timeline—and navigate the technical and legal complexity of handling PII, maintaining data utility while stripping identifying information, and securing rights—remains an open question.
The competitive landscape is filling out quickly. Snorkel AI offers its "Enterprise Environments" series, featuring domain-expert simulated companies for agent training. Mercor markets "off-the-shelf data" including APEX Agents with 2,620 tasks across eight domains. InfoBay positions its Operational Dataset—which claims 500,000+ messages, 500+ senders, and 1,000+ conversations across Slack, Notion, Salesforce, HubSpot, SAP, Workday, and Zendesk—as compliant with EU AI Act Article 10 requirements.
Then there's Horyx, which converts "real B2B work to versioned training and eval data" with agent trajectories. Rise Data Labs provides services to build structured task design and behavioral annotation for agent training. Third Origin focuses on physical AI workflows with multimodal temporal datasets.
The benchmark community is responding, too. Enterprise-Bench, launched in early July by DevRev and the Laude Institute, evaluates 14 production-like enterprise tasks with an emphasis on reliability, cost, and access constraints at production data scale. IBM released its VAKRA benchmark in March to evaluate multi-hop, multi-source agent reasoning. The Harbor framework provides adapters for Terminal-Bench and Enterprise-Bench, maintaining full trajectories and action logs for reproducibility.
Real deployments hint at what's at stake. In an April 2026 blog post, Microsoft highlighted Coca-Cola Beverages Africa using Copilot Studio agents with Dynamics 365 to autonomously run planning cycles, reportedly saving planners one to one-and-a-half hours daily. That kind of operational integration demands training data that captures actual planning workflows, decision points, and exception handling—not simplified simulation scenarios.
What Happens Next

The market is shifting toward production-fitness metrics that extend beyond single-answer accuracy. Enterprise-Bench explicitly evaluates reliability across multiple runs, cost per task, and compliance with access controls. OSWorld 2.0's low completion rates at 500 steps reveal just how far models remain from handling complex, long-horizon workflows.
Platform vendors are standardizing primitives that amplify dataset demand. Microsoft's generally available computer-using agents and Google's Agentic RAG patterns create new integration points that need workflow-grounded training data. When agents can orchestrate multi-tool interactions and navigate computer interfaces, the quality and realism of training workflows becomes the bottleneck.
Founders and product leaders should watch several indicators in the months ahead.
How fast can startups like Ooak Data actually acquire, anonymize, and package hundreds of company ecosystems? The technical and legal complexity of that pipeline remains unproven at scale. What pricing models emerge? Will this data sell as one-time licenses, subscriptions, or compute credits? The economics will determine whether workflow datasets become infrastructure commodities or defensible moats.
And how do evaluation benchmarks evolve? If Enterprise-Bench and similar frameworks gain traction as industry standards, they create strong demand signals for datasets that reproduce production-like constraints during training.
The regulatory timeline matters, too. As EU AI Act compliance requirements firm up, companies will increasingly need datasets with documented provenance and privacy controls—advantages that real-company digital twins offer over scraped or synthetic alternatives.
Perhaps the most critical question is whether the benchmark-to-production gap narrows or persists. If training on realistic workflow data closes that gap substantially, the companies controlling high-quality workflow datasets gain strategic leverage. If the gap remains stubbornly wide despite better data, the problem may lie elsewhere—in model architectures, evaluation frameworks, or the fundamental difficulty of reliable multi-step reasoning.
For now, the race is on. Enterprise leaders are deploying agents, sometimes haltingly. Platform vendors are shipping orchestration tools. And a cohort of startups is betting that the missing piece—the data that captures how work actually happens, in all its unglamorous complexity—is worth building a company around.
Whether they're right may depend less on the sophistication of their anonymization pipelines than on whether the industry can agree on what "production-ready" actually means.
