Anyone who has tried to automate their desktop knows the frustration: wire up an API that breaks in the next update, build a workflow that chokes when a button moves three pixels left, or maintain a brittle script that your future self will curse. What if, instead, you could just show a piece of software what you want, once, and have it figure out the rest?
That's the premise behind Understudy, an open-source agent platform that surfaced on GitHub in early March with a teach-by-demonstration model. Rather than recording clicks and coordinates—the bread and butter of most automation tools—it attempts to extract underlying intent from your actions. You demonstrate a task, and it builds what the project calls a "reusable skill" capable of handling variations later.
The project arrived under an MIT license and drew early attention from the local-AI crowd: over two hundred GitHub stars within days, a dozen forks, the usual signs of developer curiosity. It runs entirely on your machine, no cloud dependency, orchestrating tasks across GUI applications, browsers, shell commands, file systems, and messaging platforms from a single runtime.
Whether it delivers on that promise—and whether the teach-once-run-many workflow holds up under real use—remains an open question. But the ambition is clear enough.
Show, Don't Script
At its heart, Understudy is trying to solve a familiar problem. Most desktop automation asks you to think like a programmer even when you aren't one. You map out steps, anticipate edge cases, handle exceptions. Understudy flips that: you work through a task normally—booking a meeting, pulling data from three apps, formatting a report—and the agent watches, recording both screen video and what its documentation describes as a "semantic event stream."
From there, it constructs an "evidence pack" (scene changes, timelines, keyframes) and runs an analysis pass to tease out intent. You get a chance to validate the recording. If it looks right, you publish it as a skill artifact stored in a SKILL.md file. The CLI workflow is terse: /teach start begins recording, you perform the task, /teach stop ends it, /teach confirm --validate for review, /teach publish <draftId> to finalize.
Everything stays local by default, though certain images or keyframes may be sent to whatever model provider you've configured—Anthropic's Claude, OpenAI's GPT models, Google's Gemini, MiniMax, others—for analysis. The system is model-agnostic, which is to say it doesn't care which AI vendor you pick, as long as you configure one.
Beyond the Browser
Where Understudy diverges from browser-only tools like Skyvern or Stagehand is its execution model. A single planner decides, step by step, whether to use native GUI control, browser automation (via Playwright or an optional Chrome extension relay), shell commands, web requests, messaging adapters, or memory queries. The platform ships with 13 native GUI action primitives—click, type, drag, scroll, screenshot, wait—and can switch between them mid-task.
The GUI grounding mechanism uses a dual-model architecture: one model plans the action, another selects the target and coordinates. The project claims a benchmark of "30/30 targets resolved," using HiDPI-aware screenshot analysis with refinement passes for small UI elements. An optional overlay validation mode lets the agent verify its targeting across multiple rounds. How reliably that works in practice, outside controlled tests, is harder to say.
Out of the box, the agent supports 47 built-in skills covering apps like Obsidian, GitHub, Slack, Spotify, 1Password, Trello, Bear, and Things. It connects to eight messaging channels—Web, Telegram, Discord, Slack, WhatsApp, Signal, LINE, and iMessage—and includes scheduling for cron jobs or one-shot timers. It's a broad surface area, perhaps more than the project's small team expected to maintain at launch.
Privacy First, macOS Only

The local-first architecture is deliberate. Screenshots, recordings, traces—all of it stays on your machine. No telemetry to a central service, no data leaving your desktop except when the agent sends selected frames to your configured model provider for grounding. A March 13 update (version 0.1.3) clarified privacy and contributor documentation, suggesting early users had questions.
The tradeoff, at least for now, is platform support. The full teach-by-demonstration and GUI automation capabilities are macOS-only, requiring Accessibility and Screen Recording permissions plus Xcode Command Line Tools for a native Swift helper. Windows and Linux GUI backends are listed as planned, but there's no timeline. For a project aiming at broad desktop automation, that's a narrow opening.
Installation is straightforward if you're in the npm world: npm install -g @understudy-ai/understudy gets you the CLI. The runtime also spins up a gateway server (HTTP, WebSocket, JSON-RPC) on port 23333 by default, with terminal, webchat, and dashboard interfaces. You'll need Node.js 20.6 or later, and optional dependencies include Playwright, Chrome, ffmpeg, tesseract for on-device OCR, and signal-cli. A doctor --deep diagnostic command helps troubleshoot permissions.
Alpha Snapshot, Ambitious Roadmap

The documentation frames the current release as an alpha. Layers 1 and 2—native GUI capability and learning from demonstrations—are fully implemented, along with the daemon, gateway, terminal and web UIs, channels, scheduling, subagents, and the skills system. Layers 3 (Crystallized Memory) and 4 (Route Optimization) are partially built out.
Layer 5, labeled "Proactive Autonomy," sketches a more ambitious future: isolated workspaces or a second desktop, and a progressive trust model that moves from manual approval through suggest mode to auto-confirm and, eventually, full autonomy. The vision is there. How much of it ships, and when, is another story.
The maintainers behind the "understudy-ai" GitHub organization haven't disclosed their identities publicly; the org page lists no public members. There's no mention of funding, investors, or corporate backing. Just an MIT-licensed runtime installable from npm and a handful of contributors adding issues and PRs.
A Crowded Field
Understudy enters a space already populated by competing approaches. Browser-focused tools like browser-use and Stagehand have gained traction for web workflows. Microsoft's UFO and UFO2 research projects explore OS-level agent orchestration on Windows. DesktopCommanderMCP offers terminal and filesystem control through Claude Desktop's MCP protocol.
What sets Understudy apart—at least in theory—is its unified runtime spanning GUI, browser, shell, files, and messaging in one loop, paired with teach-by-demonstration that aims to capture intent rather than just replay coordinates. Research projects like LearnAct have explored one-shot and few-shot GUI agents in mobile contexts; Understudy applies a similar philosophy to desktop workflows.
Community discussion around OS-level versus browser-only agents has highlighted familiar tradeoffs. OS-level control sidesteps anti-bot DOM friction and generalizes beyond web apps, but it requires deeper system permissions and platform-specific implementations. Understudy bets on the former, accepting macOS-first constraints in the short term.
Early Days

For developers and technical founders experimenting with local-first agent platforms, Understudy offers an early look at demonstration-driven automation without vendor lock-in. Whether the teach-by-demonstration flow proves reliable enough for production use—whether it can truly distinguish intent from incidental clicks, whether it gracefully handles UI changes—those are questions that only sustained use will answer.
And whether the project can expand beyond macOS quickly enough to retain momentum, or whether Windows and Linux support will lag long enough to cede ground to competitors, remains to be seen. For now, it's an intriguing experiment in rethinking how we automate the work we do every day. Sometimes the most interesting projects are the ones that ask, with genuine curiosity, whether there's a better way.
