Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSJuly 25, 2026

YC-Backed Rekursiv.ai Claims 10,000x Cost Cut in AI Research

YC-Backed Rekursiv.ai Claims 10,000x Cost Cut in AI Research
YcAgi Research+3
Healthtech & Biotech iconHealthtech & BiotechJuly 25, 2026

SecurITe Raises Double-Digit Million Seed for Healthcare AI Security

SecurITe Raises Double-Digit Million Seed for Healthcare AI Security
HealthtechCybersecurity+3

Founders Mentioned

Grégoire Lamy

Ooak Data

saas icon
SaaS

Pierre-Louis Vouteau

Ooak Data

saas icon
SaaS

Grégoire Lamy

Ooak Data

saas icon
SaaS

Pierre-Louis Vouteau

Ooak Data

saas icon
SaaS
SaaS iconSaaS
July 25, 2026
Ai AgentsEnterprise AiTraining DataB2b Saas

Why AI Agents Fail at Work: The Race to Build Real Workflow Data

As AI agents stumble in real business tasks despite benchmark success, startups race to build datasets from actual company workflows. Inside the emerging market reshaping enterprise AI.

Why AI Agents Fail at Work: The Race to Build Real Workflow Data

The demonstrations looked flawless. Then came production.

It's a storyline that has unfolded over the past year: AI agents that post impressive scores on standardized benchmarks stumble badly when turned loose on the messy, multi-step workflows that define actual business operations. The culprit, according to a growing chorus of researchers and founders, isn't the models—it's what those models learned from.

While the AI industry sprinted to build ever-larger models, a more mundane question lingered in the background. Where, exactly, do you find the kind of realistic business workflow data that could teach an agent how companies truly operate?

That gap has given birth to a new category of infrastructure startups racing to capture, anonymize, and package real company workflows into datasets that might actually prepare agents for the wild. Among them: Ooak Data, a five-person team that emerged from Y Combinator with an ambitious pitch—build what they're calling "the world's largest library of real-world business workflow datasets." Based in Paris and incorporated on September 9, 2024, by Thomas Aubry, Grégoire Lamy, and Pierre-Louis Vouteau, the startup is betting that synthetic data and simulations aren't enough.

They're hardly alone in that conviction.

When Benchmarks Meet Reality

The numbers suggest companies aren't waiting around for perfection. According to Gartner's April 2026 research, 17% of organizations have already deployed AI agents in production, with more than 60% expecting to follow suit within two years. In an August 2025 press release, Gartner had projected that 40% of enterprise applications would feature task-specific agents by the end of 2026, up from less than 5% in 2025.

But deployment is one thing. Reliability? That's where things get complicated.

Consider OSWorld 2.0, a benchmark released on June 28, 2026, that tests long-horizon computer use across 108 workflows requiring up to 500 steps. The best-performing model managed just 20.6% task completion. These aren't exotic edge cases, either—they're the kind of cross-system, multi-tool workflows that employees handle daily.

The disconnect between lab performance and field results has become hard to ignore. Forrester noted in a June 2026 analysis that "companies are chasing, few are catching" when it comes to agentic AI. IDC went further, forecasting that companies could face a 15% productivity loss by 2027 without proper AI-ready data foundations. (Whether that specific prediction holds up remains to be seen, but the sentiment reflects widespread anxiety about the current state of agent reliability.)

Meanwhile, the infrastructure has arrived. Microsoft shipped computer-using agents as a generally available feature in Copilot Studio in May. Google launched its "Agentic RAG" framework for enterprise agents in June. ServiceNow, Salesforce, and a parade of other platforms have rolled out agent capabilities.

What's missing is training data that reflects how businesses actually work—with all the exception handling, legacy system quirks, and undocumented workarounds that define real operations.

Three Converging Pressures

Digital illustration for article section "Three Converging Pressures" in "Why AI Agents Fail at Work: The Race to Build Real Workflow Data" - A sleek, minimalist conceptual image visualizing three converging pressures turning a data gap into ...

Several forces have aligned to turn this data gap into a market opportunity.

First, there's the technical reality: synthetic benchmarks don't reliably predict production performance. Ooak Data's research page puts it plainly, focusing on "the gap between benchmark performance and real-world capability." Snorkel AI, another player in this space, reported that its domain-expert simulated environments improved insurance underwriting agent performance from 10.9% to 42.0% pass@1—meaningful progress, sure, but still hitting a ceiling when working with simulated rather than actual company workflows.

Second, regulatory pressure is accelerating demand for datasets with clear provenance. The EU AI Act's general-purpose AI obligations began in August 2025, with most rules applying from August 2026, requiring training data summaries, copyright policies, and data governance documentation. Article 10 mandates appropriate data governance for high-risk AI systems. Companies building agents for European deployment need datasets with documented lineage and privacy controls—precisely what anonymized real-world workflow data promises to provide.

Third—and perhaps most telling—the process intelligence market is booming. ResearchAndMarkets estimates the process mining sector reached $3.82 billion last year, growing to $4.64 billion this year, with projections hitting $15.20 billion by 2032. UiPath, a bellwether for the automation platform category, reported $1.43 billion in revenue for fiscal 2025. These companies generate enormous workflow event logs—data that could theoretically feed agent training, if properly structured and anonymized.

The Digital Twin Approach

Digital illustration for article section "The Digital Twin Approach" in "Why AI Agents Fail at Work: The Race to Build Real Workflow Data" - A conceptual, minimalist visual representation of a "digital twin" ecosystem, featuring two perfectl...

Ooak Data's strategy centers on what they call "digital twins"—anonymized replicas of entire company data ecosystems. Their pipeline sources real enterprise data, anonymizes it, then generates reinforcement learning environments with expert-level, multi-step, multi-tool tasks. The company frames its offering as "built for agents, not chatbots," and has stated plans to acquire and anonymize over 300 full data ecosystems from real companies within six months.

Whether they can execute on that timeline—and navigate the technical and legal complexity of handling PII, maintaining data utility while stripping identifying information, and securing rights—remains an open question.

The competitive landscape is filling out quickly. Snorkel AI offers its "Enterprise Environments" series, featuring domain-expert simulated companies for agent training. Mercor markets "off-the-shelf data" including APEX Agents with 2,620 tasks across eight domains. InfoBay positions its Operational Dataset—which claims 500,000+ messages, 500+ senders, and 1,000+ conversations across Slack, Notion, Salesforce, HubSpot, SAP, Workday, and Zendesk—as compliant with EU AI Act Article 10 requirements.

Then there's Horyx, which converts "real B2B work to versioned training and eval data" with agent trajectories. Rise Data Labs provides services to build structured task design and behavioral annotation for agent training. Third Origin focuses on physical AI workflows with multimodal temporal datasets.

The benchmark community is responding, too. Enterprise-Bench, launched in early July by DevRev and the Laude Institute, evaluates 14 production-like enterprise tasks with an emphasis on reliability, cost, and access constraints at production data scale. IBM released its VAKRA benchmark in March to evaluate multi-hop, multi-source agent reasoning. The Harbor framework provides adapters for Terminal-Bench and Enterprise-Bench, maintaining full trajectories and action logs for reproducibility.

Real deployments hint at what's at stake. In an April 2026 blog post, Microsoft highlighted Coca-Cola Beverages Africa using Copilot Studio agents with Dynamics 365 to autonomously run planning cycles, reportedly saving planners one to one-and-a-half hours daily. That kind of operational integration demands training data that captures actual planning workflows, decision points, and exception handling—not simplified simulation scenarios.

What Happens Next

Digital illustration for article section "What Happens Next" in "Why AI Agents Fail at Work: The Race to Build Real Workflow Data" - A minimalist, conceptual composition featuring a sleek, abstract staircase of smooth, modern pillars...

The market is shifting toward production-fitness metrics that extend beyond single-answer accuracy. Enterprise-Bench explicitly evaluates reliability across multiple runs, cost per task, and compliance with access controls. OSWorld 2.0's low completion rates at 500 steps reveal just how far models remain from handling complex, long-horizon workflows.

Platform vendors are standardizing primitives that amplify dataset demand. Microsoft's generally available computer-using agents and Google's Agentic RAG patterns create new integration points that need workflow-grounded training data. When agents can orchestrate multi-tool interactions and navigate computer interfaces, the quality and realism of training workflows becomes the bottleneck.

Founders and product leaders should watch several indicators in the months ahead.

How fast can startups like Ooak Data actually acquire, anonymize, and package hundreds of company ecosystems? The technical and legal complexity of that pipeline remains unproven at scale. What pricing models emerge? Will this data sell as one-time licenses, subscriptions, or compute credits? The economics will determine whether workflow datasets become infrastructure commodities or defensible moats.

And how do evaluation benchmarks evolve? If Enterprise-Bench and similar frameworks gain traction as industry standards, they create strong demand signals for datasets that reproduce production-like constraints during training.

The regulatory timeline matters, too. As EU AI Act compliance requirements firm up, companies will increasingly need datasets with documented provenance and privacy controls—advantages that real-company digital twins offer over scraped or synthetic alternatives.

Perhaps the most critical question is whether the benchmark-to-production gap narrows or persists. If training on realistic workflow data closes that gap substantially, the companies controlling high-quality workflow datasets gain strategic leverage. If the gap remains stubbornly wide despite better data, the problem may lie elsewhere—in model architectures, evaluation frameworks, or the fundamental difficulty of reliable multi-step reasoning.

For now, the race is on. Enterprise leaders are deploying agents, sometimes haltingly. Platform vendors are shipping orchestration tools. And a cohort of startups is betting that the missing piece—the data that captures how work actually happens, in all its unglamorous complexity—is worth building a company around.

Whether they're right may depend less on the sophistication of their anonymization pipelines than on whether the industry can agree on what "production-ready" actually means.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • YC-Backed Rekursiv.ai Claims 10,000x Cost Cut in AI Research
  • SecurITe Raises Double-Digit Million Seed for Healthcare AI Security
  • India Targets Dorsey's Bitchat—App Migrates to Censorship-Proof Platform
  • Buz Emerges as Zig-Based Bun Fork with Sub-1s Incremental Builds
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.