Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 3, 2026

DesignVerse raises $5.5M to automate enterprise software

DesignVerse raises $5.5M to automate enterprise software
Ai AutomationEnterprise Software+3
SaaS iconSaaSOctober 3, 2026

OSCP raises $6M for GPS-free navigation sensors

OSCP raises $6M for GPS-free navigation sensors
PhotonicsSensor Tech+3
SaaS iconSaaSSeptember 17, 2026

Kepler emerges with $468M to fix AI memory bottleneck

Kepler emerges with $468M to fix AI memory bottleneck
Ai MemoryAi Hardware+3
Fintech iconFintechSeptember 17, 2026

Tare raises $13.25M for on-chain private credit exchange

Tare raises $13.25M for on-chain private credit exchange
Blockchain InfrastructurePrivate Credit+2

Founders Mentioned

Dev Mandal

Markov Studios

saas icon
SaaS

Harish Ashok

Markov Studios

saas icon
SaaS

Dev Mandal

Markov Studios

saas icon
SaaS

Harish Ashok

Markov Studios

saas icon
SaaS
SaaS iconSaaS
September 17, 2026
YcTraining DataAi AgentsB2b SaasFoundation Models

Markov Studios sells 33k+ hours expert data to AI labs

YC-backed startup supplies specialized training data for computer-use agents as frontier labs race to build AI that can operate desktop software and browsers.

Markov Studios sells 33k+ hours expert data to AI labs

Markov Studios has sold more than 33,000 hours of expert computer-use training data to leading AI labs, according to the company's Y Combinator directory page. The figure suggests that frontier model builders are willing to pay for something increasingly hard to find: footage of how professionals actually use software.

The startup, which emerged from YC's Summer 2026 batch, specializes in synchronized screen recordings that capture professionals navigating CAD programs, architecture platforms, and financial applications. These are the workflows OpenAI, Anthropic, and Google need as they try to build AI agents capable of operating desktop software and browsers. But scraping the web doesn't yield that kind of data.

"The biggest bottleneck is the lack of high-quality training data," co-founder Dev Mandal wrote in the company's launch post in August. Mandal spent time at Sarvam and studied at IIT Madras. His co-founder, Harish Ashok, previously started Zenith before turning to Markov.

What Markov targets are long, multi-step desktop workflows: an engineer designing a bridge in AutoCAD, a trader reconciling positions across Bloomberg terminals, a designer assembling renders in Blender. The kind of sequences that general web-scraping misses entirely.

What the Datasets Contain

Markov's datasets typically include synchronized mouse movements, keyboard inputs, and annotations tied to individual frames. The company's CAD 1000 Hours dataset contains just over 1,021 hours across 597 workflows spanning 10 CAD and building-information-modeling tools, according to its Hugging Face page. Downloads last month reached 148,941. A broader collection, Computer-Use Large, holds approximately 12,300 hours across AutoCAD, Blender, Excel, Photoshop, Salesforce, and Visual Studio Code. That one logged 48,687 downloads in the same period.

Markov reported 200,000-plus total Hugging Face downloads and said it sold 33,000-plus hours to labs under commercial licenses. At launch in August, the company cited 15,000 hours sold and 150,000 downloads, which suggests the business doubled its paid volume in about six weeks. Whether that pace holds is another question.

The Agent Race Is Real, But Performance Isn't There Yet

OpenAI introduced its Computer-Using Agent on January 23, 2025, calling the capability "the next step in AI development, allowing models to use the same tools humans rely on daily." The company equipped its Responses API with a computer environment and shell tool for select developers. Anthropic shipped computer-use features in October 2024 and expanded them to a 17-tool "computer_toolset" by August 2026. Google unveiled Project Mariner, a browser agent, in December 2024 and integrated computer-use previews into its Gemini Enterprise Agent Platform this year.

Despite the product announcements, performance remains underwhelming. OSWorld, a benchmark covering 369 real computer tasks across operating systems and applications, reported that the best models achieved roughly 12 percent success when the benchmark debuted at NeurIPS in 2024. Anthropic wrote in an April 2026 blog post that "computer use is still early compared to Claude's ability to code or interact with text."

Labs need training data that mirrors how experts actually work. WindowsWorld, a 2026 benchmark preprint, found that 78 percent of realistic tasks span multiple applications. A financial analyst pulls data from Excel, transforms it in Python, and uploads results to Salesforce. Markov's datasets capture these cross-application sequences with frame-level detail, which matters when you're trying to teach a model to behave like a human operator.

The Legal Landscape Is Shifting

Digital illustration for article section "The Legal Landscape Is Shifting" in "Markov Studios sells 33k+ hours expert data to AI labs" - A minimalist, conceptual representation of shifting legal landscapes and AI regulation, featuring a ...

Markov entered a market shaped by tightening legal and regulatory constraints. The European Union's AI Act imposed general-purpose AI obligations on August 2, 2025, requiring providers to publish summaries of training content and comply with copyright law, according to the European Commission's GPAI fact page. An OECD report published in February 2025 flagged intellectual-property, database-rights, and trade-secret risks for AI models trained on scraped data, concerns that have continued to shape industry practices through 2026. Law firms highlighted rising DMCA anti-circumvention theories in 2026 litigation over downloaded platform content.

Buyers increasingly demand rights-cleared data with documented contributor consent. Appen, an enterprise data vendor, marketed "enterprise data to advance frontier AI" with defined licensed uses this year. YC company Datoric released Computer-Use Agent Traces 250k in mid-2026, offering 250,000 traces across roughly 15,000 hours under commercial license and positioning itself as a "security-first training data R&D engine for frontier labs," according to its YC profile and Hugging Face repository.

ServiceNow published VideoCUA at an ICLR 2026 workshop, contributing roughly 55 hours of continuous 30-frame-per-second expert desktop video across 10,000 tasks, plus GroundCUA with more than 5 million UI annotations. PatronusAI released FigmaTrace in August 2026, covering 3,469 Figma trajectories across more than 200 hours and claiming out-of-distribution transfer gains in a paper published August 20. Microsoft's synthetic-computers-at-scale dataset, updated in 2026, generated synthetic personas, filesystems, and tasks for agent training.

The provenance of training data now carries a premium. Labs can't afford the reputational or legal risk of scraped datasets with unclear lineage.

Governance Can't Keep Pace with Deployment

Enterprise adoption is accelerating, but governance lags. McKinsey reported in August 2026 that 44 percent of organizations now scale AI enterprise-wide, up from 38 percent year-over-year. Deloitte's 2026 State of AI report, based on surveys fielded in August and September 2025, found that roughly 80 percent of organizations lack mature governance for agentic AI. Only 5 percent said business processes are "highly prepared" for agents.

ITPro reported in September 2026 that agents have "hit the mainstream in software engineering" but security and governance practices lag incident rates. The gap between deployment speed and oversight has driven interest in auditable, replay-capable datasets that labs can trace and regulators can review.

Markov competes on coverage of niche professional software and synchronization quality. The company's workflow samples page, accessed September 16–17, 2026, advertised more than 20 hours across 21 workflows with frame-level narration and "native Scale-CUA action trajectories," totaling 5,700 narration annotations and 10.7 million input events. Scale offers "Computer Use Environments" for training and evaluation, according to a September 2026 product page.

Benchmarks Are Getting Harder

Digital illustration for article section "Benchmarks Are Getting Harder" in "Markov Studios sells 33k+ hours expert data to AI labs" - A clean, minimalist conceptual image representing the escalating complexity of software benchmarks, ...

Benchmarks are maturing from single-application tests to multi-app, process-centric evaluations. WindowsWorld covers 181 tasks. MacOSWorld, released in June 2025, spans 202 multilingual tasks across 30 macOS applications. GUI-Odyssey, presented at ICCV 2025, includes more than 8,300 episodes across 212 apps and maintains active leaderboards through 2026.

OS-Harm, a safety benchmark published in June 2025, tests 150 tasks for misuse, prompt injection, and agent misbehavior. It's now standard in enterprise requests for proposals, according to the research paper.

OSWorld-Human, a June 2025 study, found that agents take 1.4 to 2.7 times more steps than necessary. Planning and reflection dominate latency. Labs will need data showing efficient expert paths, not just successful outcomes.

Model providers continue expanding managed computer-use environments. OpenAI's Responses API now includes a computer environment. Anthropic formalized its toolset schema. Google exposed preview computer-use access via Gemini's agent platform. These hosted, auditable environments increase demand for datasets that map cleanly to tool schemas and include logs that buyers can replay and audit.

Can Markov Sustain the Growth?

Digital illustration for article section "Can Markov Sustain the Growth?" in "Markov Studios sells 33k+ hours expert data to AI labs" - A minimalist, conceptual representation of rapid business growth and the accumulation of time, featu...

Markov's 33,000 hours sold and six-week doubling signal early traction. The company is riding two trends: labs racing to ship computer-use agents and buyers demanding clean, licensed data. Whether the startup can sustain growth depends on how fast labs productionize computer-use agents and how much governance pressure forces them to replace scraped data with licensed alternatives.

The market for expert workflow data is still young. But if agents are going to work as advertised, they'll need to learn from people who know what they're doing. That's what Markov is betting on.

More stories

  • DesignVerse raises $5.5M to automate enterprise software
  • OSCP raises $6M for GPS-free navigation sensors
  • Kepler emerges with $468M to fix AI memory bottleneck
  • Tare raises $13.25M for on-chain private credit exchange
  • Ryft raises $27M Series B to scale marketplace payment platform
  • QuBeats raises $15M to build quantum navigation sensors
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.