The screen recordings arrive without context: two invisible artificial intelligence systems, each trying to book a dinner reservation or fill out a tax form, clicking through identical browser windows. Human judges watch the replays, blind to which company's software is which, and vote for the winner. Within two weeks of launching this digital gladiator pit, a two-person San Francisco startup called CoArena pulled in $60,000 in revenue.
That early traction speaks to a peculiar tension gripping the AI industry right now. Tech giants and startups alike are racing to build computer-use agents that can navigate the internet, operate desktop software, and carry out multi-step tasks on behalf of humans. Anthropic shipped its computer-use capability in public beta in October 2024; OpenAI integrated Operator into ChatGPT this past July. Marketing claims pile up. But when academics test these systems against benchmark tasks, the pass rates hover somewhere between dismal and slightly less dismal.
CoArena, which emerged from Y Combinator's YC S26 batch, has turned that measurement gap into a business model. Instead of static test suites that go stale the moment they're published, the company runs live head-to-head battles on real user tasks and maintains a crowdsourced leaderboard. The top spot has changed hands three times in three days, according to the company's launch announcement. What's selling isn't just the rankings—it's the trajectory data underneath, which buyers can license to train their own models.
The market opportunity looks enormous, at least on paper. McKinsey estimated in a January 2026 report that computer-use agents could mediate between $3 trillion and $5 trillion in consumer commerce by decade's end. Gartner pegged enterprise spending on cross-functional AI agents at $23 billion by 2030. Yet Forrester's June report, colorfully titled "The Great AI Agent Rationalization," found that security leaders rank agentic AI among their top worries. A late-August McKinsey survey noted that while 62 percent of organizations are experimenting with agents, only 23 percent are actually deploying or scaling them.
The gap between hype and reality shows up starkly in the academic literature. OSWorld, a benchmark dataset presented at NeurIPS 2024, included 369 real desktop and web tasks in its initial release. Human testers scored 72 percent. The best AI model managed just over 12 percent. WebArena, published by Carnegie Mellon in mid-2023, showed GPT-4-based agents completing between 10 and 14 percent of realistic browser tasks. Even OSWorld 2.0, released this past June with longer workflows, found top agents still completing fewer than a quarter of tasks when allowed 500 steps. Claude Opus hit 20.6 percent.
Anthropic described its computer-use feature as "experimental—at times cumbersome and error-prone" when it launched. OpenAI's Operator relies on a model it calls Computer-Using Agent, combining GPT-4o's vision capabilities with reinforcement learning to interact with graphical interfaces. Both companies have lined up early partners: Anthropic named Asana, Canva, Cognition, DoorDash, and Replit; OpenAI listed DoorDash, Instacart, OpenTable, Priceline, StubHub, Thumbtack, and Uber.
But vendor claims muddy the water. UiPath announced in January that its Screen Agent ranked first on OSWorld-Verified at 53.6 percent, citing October research. Coasty Systems, CoArena's parent company, markets a model called coasty-v5 with claims of "85.6 on OSWorld; 82.8 independently verified" on its website. Those numbers sit far above the academic baselines, and without independent confirmation, they register more as marketing assertions than settled science.
The broader benchmark ecosystem carries its own credibility questions. A 2025 paper called "The Leaderboard Illusion," accepted to NeurIPS, audited the dynamics of LMSYS Chatbot Arena and raised concerns about provider access, overfitting, and private variant testing inflating ranks. LMSYS published a response contesting several claims and outlining mitigations, though the debate continues.
CoArena's approach sidesteps some of those pitfalls through controlled conditions. Users submit real tasks. Two agent runs execute under identical constraints: same prompt, same 1280×800 viewport, same sandbox image, same 100-action and 20-minute limits. Every competitor is a third-party model reached through its vendor's public API—Anthropic, OpenAI, Google, Meta, xAI, and open-weight models hosted by Fireworks AI, all driven by one testing harness. Judges watch screen replays without identifying information and vote. The system calculates ratings using a maximum-likelihood Bradley-Terry model with uncertainty intervals, weighted 60 percent on a fixed test set and 40 percent on arena comparisons.

The platform defines 13 browser actions and eight desktop actions. Desktop runs use a real X display; no terminal access is exposed. A constrained "command" verb invokes allowlisted programs with sandboxing. Matchmaking weights unknown or low-sample agents higher to keep the comparison graph connected, the company said.
One wrinkle: model providers receive raw, unmasked screen captures on every step while battles run, before any detector has touched the frames. "A computer-use agent acts by looking at the screen, so there is no version of this product in which that does not happen," the company wrote on its governance page. Screenshot bytes never ship in dataset exports; licensees receive pairwise preferences and trajectory logs—actions, errors, reasoning text—but not images.
The founders, Prateek Jannu and Nitish Kovuru, met at Purdue before scattering to Stanford and Columbia for graduate work. Jannu holds a Stanford master's in electrical engineering and built an agent called Coasty that the Y Combinator profile claims ranked first on OSWorld. Kovuru studied computer science at Columbia, and the YC listing credits him with "SOTA CUA framework on OSWorld at 83% accuracy," though those claims lack verification against published academic baselines.
CoArena sells dataset licenses in tiered packages. Buyers get pairwise preferences and trajectory logs, with consent collected at task submission. The company corrected its terms on September 4 after discovering that between early August and early September, the terms had overstated desktop privacy protections. Roughly 9,600 desktop battles were pooled and 4,500 judged by others during that window. The frames were of CoArena's sandbox desktop, not user machines, and screenshot bytes were never delivered to licensees.
Early performance data from the platform paints a sobering picture. Only 66 percent of agent runs finish their tasks at all. On identical assignments, the best frontier agent completes roughly 87 percent of runs; the worst manages around 42 percent. In about 7 percent of battles, human judges ruled both agents failed. User growth has been doubling weekly since launch, according to the company.
The platform's Metrics API defines ten measurement families, including Outcome, Route, Recovery, Precision, Tempo, Cost, Expression, Output, Head-to-Head, and Judgment. But the API documentation carries a warning: "Completion" is self-reported by the agent, not externally verified success. "A demo is not a measurement and a self-reported success rate is not a result," the company wrote. "A static benchmark score is real the day it is published and decays from then on."

CoArena arrives as the field pivots from static test sets to live evaluation, following the model pioneered by LMSYS Chatbot Arena for large language models. LMSYS has published peer-reviewed research at ICML 2024, NeurIPS 2023, and NeurIPS 2025, grounding its rankings in tens of millions of user votes. Computer Agent Arena, an open crowdsourced platform from xlang-ai accepted to ICLR 2026, shares the live-benchmark philosophy.
New academic benchmarks continue to proliferate. MacArena, published this past June, covers 421 macOS tasks on Apple Silicon. WindowsAgentArena, led by Microsoft, targets multi-modal Windows environments. BrowserGym, described in a December 2024 paper, integrates WebArena and AgentLab for large-scale web-agent experiments. All emphasize long-horizon workflows, error recovery, and multi-app processes.
Regulatory and platform friction add another layer of complexity. Microsoft's Recall feature faced immediate backlash over screen capture defaults and now requires opt-in, the company said last September. macOS Sequoia began prompting monthly for screen-recording permission in August 2024. The EU AI Act, consolidated this past July, details obligations for general-purpose AI systems, and an FTC proposed policy statement from mid-2026 emphasizes accountability for deceptive AI marketing and accuracy claims.
Gartner forecast worldwide AI spending to reach $2.59 trillion this year, citing agentic automation in multistep processes as a key driver. Forrester's 2026 predictions cautioned that production scarcity of true multi-agent systems would shift spending into 2027, with governance and security as obstacles. The number that matters now: whether CoArena's leaderboard converges with academic results or vendor claims, and whether enterprise buyers actually use it to cut through the marketing noise.
The OECD.AI tool catalog listed CoArena last week as "a free computer-use utility and live evaluation platform for frontier AI agents." Free for users submitting tasks, perhaps. But for companies desperate to know which agent actually works better than the marketing deck claims it does, the data license is starting to look worth the price.
