Twelve thousand three hundred hours. That's how much screen-recorded human activity a small Y Combinator-backed startup decided the world needed to train the next generation of computer-using agents—software capable of clicking through Excel, navigating Salesforce, or editing in Photoshop without human hands on the keyboard.
Whether that's anywhere close to enough is another question entirely.
Markov Studios released what it describes as the world's largest open-source dataset of computer-use recordings in March 2026: 48,478 videos spanning six professional applications, all under a Creative Commons license. The timing speaks to a bottleneck the industry would rather not discuss. As Microsoft, Google, and OpenAI barrel toward deploying agents that can autonomously operate software—Microsoft's Copilot Studio went generally available on May 13, 2026, Google integrated computer-use features into Gemini 3.5 Flash the following month—they've hit a constraint that has nothing to do with compute budgets or transformer architectures.
It's data. Specifically, the vast, tedious, expensive-to-generate mountains of high-quality human demonstrations needed to teach models how to actually use a computer.
When Copilots Stop Suggesting and Start Doing
Computer-use agents represent something of a leap from the assistive AI most enterprises have cautiously adopted. A copilot suggests code snippets or summarizes meeting transcripts. A computer-use agent—industry shorthand: CUA—books your travel, manipulates your spreadsheets, navigates your enterprise software stack end to end. The difference between autocomplete and autonomy.
Gartner published its Hype Cycle for Agentic AI in April 2026, finding that only 17% of organizations had deployed agents to date. But more than 60% expected to within two years, a gap that suggests either aggressive optimism or corporate FOMO at scale. The consultancy projects Fortune 500 companies will collectively run something on the order of 150,000 agents by 2028—a forecast that hinges on technical maturity catching up to executive enthusiasm.
Market projections reflect the momentum, if not always the rigor. BCC Research estimated in January 2026 that the AI agents market stood at $8 billion in 2025, projected to hit $48.3 billion by 2030—a compound annual growth rate of 43.3%. The methodology, sourced from a vendor-sponsored synopsis, probably warrants a raised eyebrow. But the directional trend aligns with what enterprises are actually doing. McKinsey, in an October 2025 report, suggested agentic commerce alone could orchestrate up to $1 trillion in U.S. B2C revenue by 2030, a projection based on the consultancy's proprietary methodology.
Yet adoption is throttled by a stubborn technical reality. The OSWorld 2.0 benchmark, published in June 2026, reports that the best models complete only 13–21% of long-horizon tasks at fixed step budgets. Real-world enterprise deployment demands far higher reliability. And that reliability, it turns out, hinges almost entirely on training data.
The kind Markov is now giving away for free.
The Volume Problem, and the Quality Problem
Training a model to perceive graphical user interfaces and execute tasks via mouse and keyboard requires enormous volumes of what researchers call "trajectory data"—recordings of humans actually doing things on computers, step by step. That data simply doesn't exist in the wild at the scale needed.
Microsoft Research's Fara1.5 computer-use agent family, detailed in papers from May and June 2026, trained on approximately 2 million samples. The mix was dominated by web trajectories (60%), supplemented by synthetic environment data, form-filling interactions, and GUI grounding tasks. Building such a corpus requires scraping vast quantities of human demonstrations, filtering for quality at the step level, and often applying reward models to grade individual actions—distinguishing a productive click from a confused one.
OpenAI's foundational work illustrates the constraint. The company's Operator agent, introduced in a January 2025 research preview and later integrated into ChatGPT by July 2025, relied on reinforcement learning trained against human trajectories. Anthropic added computer-use features to Claude in 2026. Google's approach with Gemini 3.5 Flash emphasized adversarial training for live environments. Each effort circles back to the same fundamental problem: high-quality GUI interaction data is scarce, expensive to produce, and maddeningly difficult to label at scale.
The academic community responded with a proliferation of benchmarks in 2026—OSWorld 2.0, MyPCBench, WeaveBench, GUIDE (presented at CVPR 2026), WebSTAR (ACL 2026). Most focused on evaluation rather than large-scale training corpora. Several research groups turned to mining unlabeled screen-recorded videos from the internet, an approach that feels a bit like panning for gold in YouTube tutorials.
VideoAgentTrek, presented as an ICLR 2026 poster, reported a 70% relative improvement on OSWorld-Verified by extracting GUI actions from videos via inverse dynamics modeling. WebSTAR synthesized 13,300 trajectories and 267,000 graded steps from OpenAI's computer-use preview outputs, introducing step-level reward models as quality control. Progress, in other words, but still measured in thousands of samples when the industry likely needs millions.
This is the landscape Markov Studios stepped into.
Professional Software, Not Just Browsers

Markov's "Computer Use Large" dataset, hosted on Hugging Face, offers 12,300 hours of recordings distributed across Blender (3,624 hours), Photoshop (2,060 hours), AutoCAD (2,149 hours), Excel (2,336 hours), Salesforce (2,002 hours), and VS Code (127 hours). The total file size approaches 856 GB. According to Agent Wars, a third-party news site covering the agent ecosystem, the dataset drew 45,000 downloads within its first month; as of July 2026, the Hugging Face page listed 1,142 downloads in the most recent monthly period—a dropoff that could signal either market saturation or the dataset's limitations becoming apparent.
The company was founded by Dev Mandal (CEO, formerly at Sarvam AI, an IIT Madras alum) and Harish Ashok (previously at Zenith Robotics). Mandal has positioned the dataset as a natural application of "training AI from video," a theme that aligns with broader academic trends toward mining screen recordings at scale. The mission, as framed on the company's website: "data + environments for training computer-use AI."
How does Markov's offering compare to competitors? Datoric released a dataset on Hugging Face in July 2026 titled "computer-use-agent-traces-250k," claiming 250,000 real-world interaction traces with screen captures, action logs, and DOM/accessibility tree data where available—a different approach emphasizing structured metadata over raw video volume. Paradigm Shift AI's "computer-use-data-psai" dataset offers multimodal HCI traces, though it appears smaller in scale and carries 2025 timestamps. Nxtscape Data markets proprietary browser interaction data but provides sparse public documentation. Turing.com, a services vendor, claims to have built more than 10,000 supervised GUI tasks for pretraining and alignment work—a fundamentally different model centered on human annotation rather than raw trajectory collection.
What potentially differentiates Markov is the volume focused on professional desktop applications rather than just web browsing. Web-centric datasets like WebChain (31,725 real-world web interaction traces, published in March 2026) or Online-Mind2Web (dynamic web tasks updated in May 2026) serve complementary use cases but don't address the enterprise software landscape where much agent deployment is expected to occur. Salesforce, AutoCAD, Excel—these are the connective tissue of Fortune 500 workflows. Precisely the environments Gartner predicts will shift toward outcome-based agentic platforms by 2028.
Still, questions linger.
The dataset card on Hugging Face provides limited detail on step-level annotations or quality controls beyond "segment trimming" and metadata. Microsoft's Fara1.5 relied on step-level filtering and reward models. WebSTAR introduced grading signals to distinguish high-quality actions from noise. It's unclear whether Markov's open dataset incorporates similar quality mechanisms, or if the raw video recordings require significant downstream processing by users. The company's website teases a "largest computer-use data collection platform" beyond the open artifacts, hinting at proprietary environments or labeling services available via contact, but public technical documentation remains sparse. Which raises the possibility that the open release is as much marketing as public good.
The Reliability Gap

The inflection point is arriving faster than infrastructure is maturing—a pattern Silicon Valley investors might recognize from prior hype cycles.
Microsoft's Copilot Studio reached general availability globally in May 2026. Google integrated computer use into a flagship model in June. Gartner's April 2026 forecast anticipates that by 2028, most enterprises will abandon assistive AI in favor of platforms that commit to workflow outcomes—a shift that demands agents capable of complex, multi-step task completion. Legacy applications with bolt-on AI capabilities, Gartner warns, face margin compression by 2030.
Yet the technical gaps remain wide. Perhaps wider than the breathless announcements suggest.
OSWorld 2.0's sub-25% success rates on long-horizon tasks underscore how far models have to go. A University of Washington study published June 30, 2026, identified significant cybersecurity risks in agentic browsers, including same-origin policy bypass scenarios enabled by prompt injection attacks—the kind of vulnerability that makes enterprise IT departments very nervous. The MIT/Princeton AI Agent Index, presented at FAccT 2026, documented transparency and safety reporting gaps across 30 deployed agent products. Some browser agents, the researchers noted, actively market capabilities to "bypass anti-bot systems," a practice with murky compliance implications and obvious ethical questions.
Regulatory pressures are mounting in parallel. The EU AI Act entered full applicability on August 2, 2026, requiring general-purpose AI model providers to publish training data summaries and maintain copyright policies for web-mined content. Singapore's IMDA launched the first agent-specific governance framework in January 2026, updated in May with case studies and best practices around autonomy bounds, human-in-the-loop controls, and risk mitigation. Platforms that can demonstrate robust data provenance, contributor consent workflows, and compliance posture may gain a trust advantage as enterprises navigate these requirements. Or they may simply find themselves tangled in bureaucratic knots.
Training Regimes, Not Just Training Data
For founders and AI labs building computer-use agents, the data bottleneck presents both constraint and opportunity. Markov's open dataset—alongside academic efforts like VideoAgentTrek and WebSTAR—signals a maturing ecosystem where high-quality trajectory data is becoming more accessible. But volume alone won't solve the reliability problem, and anyone who tells you otherwise is selling something.
The next phase will likely emphasize step-level grading, adversarial robustness, and verified benchmarks that can differentiate agents capable of production deployment from those that falter on edge cases or hallucinate destructive actions. The Fortune 500 agent fleets Gartner envisions by 2028 won't materialize without datasets measured not just in hours, but in verified, grounded, reproducible task completions. Raw screen recordings are a start. They're not a finish line.
Whether Markov and its competitors can bridge that gap—transforming 12,300 hours of video into training regimes that close the 80% failure rate on complex tasks—will determine whether the current hype cycle delivers on enterprise expectations or stalls in pilot purgatory. The pattern is familiar: promising technology, aggressive timelines, infrastructure racing to catch up.
For now, the data infrastructure race is on. The stakes, if the projections hold, are measured in trillions of orchestrated workflows. And 12,300 hours of screen recordings, however impressive, might just be the opening bid.
