There's an unglamorous bottleneck slowing the march toward AI agents that can actually navigate a computer screen, and it has nothing to do with transformer architectures or GPU clusters. It's screen recordings. Specifically: the painstaking, mind-numbing work of capturing, verifying, and labeling what a human being does when they click through Excel macros, debug in VS Code, or wrestle with Salesforce.
That bottleneck is what a small, almost anonymous startup called Markov Studios thinks it can solve.
The company—just two people, backed by Y Combinator—released what it bills as "the world's most advanced computer-use dataset" in March 2026: 48,478 screen recording videos covering some 12,300 hours across six professional applications. AutoCAD, Blender, Excel, Photoshop, Salesforce, VS Code. Released under a CC-BY-4.0 license on Hugging Face, the dataset has reportedly drawn more than 100,000 downloads, a figure claimed by Dev Mandal in LinkedIn posts—a co-founder and former engineer at Sarvam who studied at IIT Madras.
Most companies chasing the AI agent dream want to build the agents themselves—the superintelligent assistants that will someday book your meetings, reconcile your spreadsheets, maybe even design your infrastructure. Markov isn't interested in that race. It's selling the raw material. Call it data infrastructure rather than model development, a bet that in the gold rush for autonomous software, the real money might be in picks and shovels.
Scale Without the Glamour
What makes Markov's release notable isn't novelty exactly. It's the combination of breadth and volume. Most public computer-use datasets released so far skew either small or synthetic. XLANG's OpenCUA, presented at NeurIPS, offers 22,600 human-annotated trajectories—a foundational reference, but limited in scope. Microsoft put out 98 synthetic desktop environments in a May research drop. Both useful, neither close to the hours of real human behavior Markov claims to have bottled.
The timing, at least, makes sense. Gartner projected in August 2025 that 40 percent of enterprise applications would feature task-specific AI agents by the end of 2026, up sharply from less than 5 percent in 2025. The broader AI agents market is forecast to balloon from $7.63 billion in 2025 to $182.97 billion by 2033, per Grand View Research. Training data—the substrate that teaches a model how to actually use Salesforce or find a bug in Visual Studio Code—represents its own distinct market, valued in the billions and climbing fast, according to Fortune Business Insights.
But forecasts are one thing. Deployment is another. And deployment, it turns out, is running ahead of the infrastructure needed to support it.
OpenAI's Operator launched in research preview early last year, then integrated into ChatGPT by summer. Anthropic's Computer Use capability went live for API customers. Google consolidated agent tooling under its Gemini Enterprise Agent Platform in the spring; Microsoft's Agent Framework hit general availability around the same time. Salesforce's Agentforce, by some industry tallies, crossed 29,000 deals and approached $800 million in annual recurring revenue late last year.
Yet a KPMG AI Pulse survey conducted early this year found that 65 percent of organizations cited scaling difficulties as a top barrier. Another 62 percent flagged skills gaps. The technical challenge, when you drill down, is subtle but stubborn: agents need to learn not just what to do, but how interfaces actually work. Where buttons live. How dropdowns behave. What happens when you right-click in Photoshop versus Excel. That kind of implicit knowledge doesn't come from text corpora or image-caption pairs. It comes from watching humans work—frame by frame, click by click.
Crowded, But Not Settled

Markov isn't alone in spotting the opportunity. Scale AI, the $13 billion data labeling incumbent, has productized "Agent Data" specifically for computer-use training, complete with reinforcement learning environments. Paradigm Shift AI offers Captr, a recorder that captures real human traces with DOM snapshots and event logs—50GB per session, in some cases. Microsoft's synthetic dataset, while not as voluminous, provides procedurally generated desktop environments designed to stress-test generalization.
The competitive moat here isn't purely about volume, though. It's about rights, quality, and verifiability. A paper published in March by researchers at AISI and Oxford analyzed over 177,000 tools across nearly 20,000 Model Context Protocol servers and raised flags about action-oriented tools lacking proper oversight. Another study—ProCUA-SFT, released in June—demonstrated that naively fine-tuning on existing datasets like AgentNet can produce negative transfer. In other words, models that perform worse, not better. The authors released 3.1 million step-level samples distilled from synthetic trajectories as a corrective, a hint that not all training data is created equal.
Then there's the regulatory backdrop. The EU AI Act's transparency requirements for general-purpose AI models took effect last August, with enforcement beginning this year. Screen recordings, by their nature, often capture personally identifiable information—requiring GDPR-compliant redaction and lawful processing. In the U.S., session-replay litigation around wiretap claims and consent has turned data provenance into a diligence checkpoint, not a nice-to-have.
Markov's founders—Mandal and co-founder Harish Ashok, who previously founded Zenith Robotics and worked at Boldmoon Cloud—are young, but they appear to have picked their lane carefully. The company's sparse website describes "RL environments for computer-use AI" and "high quality tasks and data to train the next generation of computer use agents." That phrasing—RL environments, not just datasets—suggests the roadmap includes interactive sandboxes where models can practice, fail, and improve, not just learn from passive observation.
Whether 12,300 hours of screen recordings translate into meaningfully better agents remains an open question.
The Quality Question

Epoch AI warned earlier this year that public text data will likely be exhausted before the decade's end, driving demand for new data modalities: interaction logs, video, proprietary captures. But quantity alone doesn't solve for quality, a lesson the industry has learned repeatedly. OpenComputer, a paper from May, highlighted generalization gaps when models trained on verified benchmarks like OSWorld face open-ended real-world tasks. HiL-Bench from Scale AI in April showed that teaching agents when to escalate to humans—a form of safe judgment that's perhaps more valuable than autonomous execution—requires purpose-built training loops, not just more footage.
For now, Markov remains mostly in stealth beyond its Hugging Face releases and Y Combinator directory listing. The two-person team, based in San Francisco, hasn't disclosed funding amounts, customers, or revenue. But the download counts, whatever their exact figure, suggest traction. And the timing—releasing a massive computer-use dataset just as enterprises scramble to deploy agents—suggests they've read the room correctly.
As Meta's AI research chief Dawn Song put it in a recent interview, the goal is "AI agents that are economically valuable"—agents that can actually do the boring, repetitive work humans currently pay SaaS companies to enable. Those agents need training wheels. Markov is building them, one screen recording at a time. Whether that's enough to carve out a durable business in a crowded, fast-moving market is the bet they're making. For a two-person team with a dataset and a hypothesis, it's not the worst bet in the world.
