The pitch sounds almost too good to scrutinize seriously: a two-person startup, working out of San Francisco with what amounts to pocket change in the world of foundation models, says it has built an AI system that learns to use computers by simply watching—no labels, no instructions, just raw video.
Induction Labs made that claim in late July, and the specifics are striking enough that they've managed to cut through the noise. The company's debut model, Photon-1, allegedly outperforms Google's Gemini 3.1 Flash-Lite on internal benchmarks for computer use. The kicker? They say it took roughly 30 times less compute to train.
For a team of two, that's a bold assertion. But boldness and Y Combinator—Induction Labs is part of the accelerator's summer cohort—have always had a certain affinity.
What They're Actually Building
Jonathan Li and David Li, the founders, describe their work as "building intellectually curious AI." Strip away the branding, and what emerges is a technical approach that departs from the supervised-learning playbook most of the industry still follows. Instead of feeding a model explicit action labels or step-by-step instructions, Photon-1 learns by predicting what happens next in a learned representation of the world—what the team calls an "imagination model."
The architecture is a sparse mixture-of-experts transformer, 106 billion parameters, with a context window of 32,000 tokens. It compresses video frames into 960 discrete latent tokens using finite scalar quantization, pulling from a codebook with more than 390 million possible states. The model trained on 575 million frames—roughly 18 years of video if you sample at one frame per second—culled from screen recordings buried in an index of 2 billion public videos.
All of that, according to Induction Labs, required about 30,000 H200 GPU-hours. In FLOP terms, that's around 4.4×10^22. Not trivial, but a far cry from the compute budgets rumored to power the latest models out of OpenAI or Anthropic.
In demos the company released alongside its announcement, Photon-1 simulates desktop environments from a single screenshot—plausible interactions with VS Code, Gmail, ChatGPT. After reinforcement learning, the model apparently learned to use ChatGPT "like a human," whatever that means in practice. It played checkers after finetuning on 20,000 tournament games. It even simulated billiard physics.
Impressive, perhaps. But these are company-provided demonstrations, not independent benchmarks.
The Compute Story Everyone Wants to Believe

The efficiency claim is where Induction Labs really leans in. They say Photon-1 bests Gemini 3.1 Flash-Lite—a production model from Google DeepMind, released earlier this year and designed for agentic tasks—on internal computer-use benchmarks. And they claim it does so using 30 times less pretraining compute and delivering three times lower serving costs.
There's just one catch: those benchmarks are internal. As of early August, there are no peer-reviewed papers, no OSWorld scores, no standardized leaderboard entries. The comparison rests entirely on Induction Labs' own methodology, which hasn't been scrutinized by outsiders.
That doesn't mean the claim is wrong. It does mean it's unverified.
Still, timing matters. Foundation model builders are acutely aware of compute costs right now—perhaps more than at any point since the GPT-3 era. If the efficiency story holds, even partially, under independent review, it would represent a meaningful data point in the ongoing debate over how to pretrain models on video and build what the research community loosely calls "world models."
A Tiny Team in a Crowded Field

Induction Labs remains a two-person operation, working out of San Francisco and actively recruiting according to their Launch YC post. The company is part of Y Combinator's summer batch, which Demo Day records suggest includes more than 100 startups. The cohort presents on September 10, where around 1,500 investors and media types typically show up.
YC's standard deal is $500,000: $125,000 on a post-money SAFE for 7 percent equity, plus another $375,000 on an uncapped most-favored-nation SAFE. Beyond that, Induction Labs hasn't disclosed additional funding. Harj Taggar, a YC partner, is listed as their primary contact.
Two people. Half a million dollars. A model they say rivals Google's.
The computer-use domain, meanwhile, is anything but uncrowded. Anthropic's Claude has offered computer-use capabilities for months. OpenAI introduced its Computer-Using Agent earlier this year. The field is moving quickly, and the window for differentiation narrows every quarter.
The Curiosity Gambit

Induction Labs frames its work around "intrinsic curiosity," a concept with deep roots in reinforcement learning research. The idea: an agent can learn useful behaviors by seeking out novel states rather than optimizing for explicit rewards. It's been explored in academic labs for years, but rarely deployed at the scale of a commercial foundation model.
Whether that framing translates to a durable competitive edge is harder to say. Curiosity-driven learning has theoretical appeal, but the market rewards results—specifically, models that perform reliably on tasks customers actually care about.
For now, Photon-1 is a research preview. A technical demonstration from a tiny team with an ambitious hypothesis about how foundation models ought to learn. Demo Day in September will clarify where the founders intend to take this, and whether investors believe curiosity-driven pretraining can challenge the incumbents who have spent orders of magnitude more to get where they are.
The claim is audacious. The team is small. The compute budget, if accurate, is strikingly lean.
What remains to be seen is whether any of it scales.
