Induction Labs, a tiny San Francisco outfit with just two founders, says it has trained an AI model that rivals a production Google system for controlling computers, using 30 times less computing power. That's according to benchmarks the company released in July, a claim that arrives at a moment when the AI industry is wrestling with diminishing returns from traditional text-based training methods.
The approach sounds almost heretical in an era of billion-dollar training runs. Jonathan Li and David Li, the founders behind Induction, call their technique "imagination models." They train on raw screen recordings scraped from billions of internet videos, teaching the system to predict what happens next without any labeled actions or captions. The model learns by watching, not by being told.
Their first model, Photon-1, uses a sparse mixture-of-experts architecture with 106 billion total parameters. It processes 32,000 tokens of context at once. The training diet consisted of 575 million frames pulled from 2 billion videos, roughly equivalent to 18.2 years of footage if played at one frame per second. The whole endeavor took about 30,000 hours on Nvidia's H200 GPUs and cost far less than the multi-hundred-thousand GPU-hour runs that other labs typically require for similar benchmarks, according to a research post the company published July 23.
That efficiency gap, if it holds up, matters more than the technical details might suggest.
When Text Runs Out
Foundation model labs have mostly moved on from pure chatbots. Computer-use agents are the new prize, systems capable of navigating desktops, clicking buttons, filling forms, and chaining together workflows across multiple applications. OpenAI launched a Computer-Using Agent research preview in early 2025, reporting 38.1 percent success on OSWorld and 58.1 percent on WebArena at debut, though performance has evolved since. Google folded computer-use capabilities into Gemini 3.5 Flash in June and shipped desktop Gemini apps for Windows and macOS on September 11. Anthropic's Claude Opus and Sonnet models held the top spots on OSWorld-Verified through much of this year, according to system cards the company published.
But even the best systems still stumble on roughly a third of standard desktop tasks. OSWorld 2.0, released in July, introduced 108 multi-hour workflows that span multiple applications. The leading model, Claude Opus 4.8 with maximum thinking enabled, completed just 20.6 percent of tasks under a 500-step limit, according to the July arxiv paper. Human completion time for these tasks averages 1.6 hours and requires an average of 318 tool calls with Claude Opus 4.7. That gap has kept enterprise adoption measured. Gartner reported in April that only 17 percent of organizations had deployed AI agents at that time, though over 60 percent expected to do so within two years.
The industry's standard recipe has grown expensive. Train on vast text corpora, then add inverse dynamics modeling or action-conditioned finetuning using labeled screen interactions. Goldman Sachs Research estimated in May that AI capital expenditure could reach roughly $765 billion this year, with some scenarios requiring $700 billion or more to sustain prior cycle peaks. Mordor Intelligence projects the foundation model market will grow from $31.2 billion in 2026 to $119.29 billion by 2031, according to a forecast updated in January.
Scaling text alone has slowed. OpenAI's o3 model broke SWE-bench Verified in December 2024, but further gains have demanded exponential compute increases, according to academic leaderboards tracked through mid-year.
Three forces are pushing labs toward observation-based architectures. The stockpile of high-quality text is finite. Synthetic data can extend training but introduces model collapse risks documented in multiple papers published last year. Desktop and physical-world tasks require spatial and temporal reasoning that text encodes poorly. OSWorld and robotics benchmarks reward models that understand pixels, not paragraphs. And video is abundant: billions of hours of screen recordings, gameplay footage, and webcam streams sit unlabeled on the internet, offering a training signal richer than any caption dataset.
Curiosity-driven learning has roots in 2017 work by Pathak and colleagues and 2019's Dreamer model by Hafner's group. Recent adaptations for large language models include Meta's Self-Rewarding Language Models from January 2024 and variants published through last year. Induction Labs applies the idea to world modeling. Its "intrinsic discovery" method, detailed in an August 27 post, uses reinforcement learning to push an agent toward unexplored terminal states, then trains a world model on the self-generated rollouts. The company said the approach produced five times more diverse terminal states than a frozen policy and that its resulting world model, Terminal-35B-A3B, outperformed GPT-5.6 Sol on AgentWorldBench-Terminal-V2 in internal evaluations.
The technical innovation lies partly in how Induction compresses video. Finite scalar quantization, the vision-encoding technique the company uses, squeezes each frame into 960 discrete tokens represented as eight-dimensional vectors with values in the set {−1, −½, 0, ½, 1}. That achieves roughly 100-times compression at 2.2 kilobytes per frame. A differential latent encoder captures inter-frame deltas. The architecture sidesteps the pixel-space diffusion common in earlier world models and the action-label bottleneck of inverse dynamics approaches, the founders said.
Testing the Claims

Induction Labs benchmarked Photon-1 against Gemini 3.1 Flash-Lite, a lightweight production model Google released. The company said Photon-1 matched or exceeded Flash-Lite's performance with 30 times less pretraining compute and lower serving cost, according to the July 23 post. After finetuning, the model learned to simulate full desktop sessions from a single screenshot, play checkers, model billiard physics better than language-model baselines of similar size, and "use ChatGPT like a human." Those capabilities emerged from observing unlabeled video. No action annotations were used during pretraining.
The caveat is that these benchmarks remain internal and have not been independently verified. OSWorld, WebArena, and AgentWorldBench results have not been independently reproduced by third parties.
Jonathan Li, Induction's founder and CEO, joined Cohere at 17 as the company's youngest researcher, working on reasoning, long-context models, and reinforcement learning infrastructure, according to his personal site. Co-founder David Li shares authorship on the lab's research posts. Neither founder's LinkedIn profile is fully accessible. The team is based in San Francisco and is hiring, according to the company site.
They are not alone in this pivot. Standard Intelligence, a Sequoia-backed startup, announced in April that it is training "general intelligence in pixel space" from raw computer-use video. Prentis, co-founded by Reid Hoffman and Mark Pincus, told TechCrunch in July it is training models to learn office workflows and reported wins on WindowsAgentArena and ScreenSpot-v2. Nvidia released Cosmos, a family of "world foundation models," in January 2025 and positioned the platform for agentic and physical AI at its annual developer conference. Forbes wrote in June that world-model startups raised over $3 billion in the first half of this year, citing venture-capital sources.
Markov, another company from the same Y Combinator batch, is building "expert computer-use data for frontier AI labs," according to its YC page. That's a potential data partner for action-conditioned finetuning, complementing rather than competing with observation-based pretraining. AWS added computer-use tools for Claude to Bedrock Agents in beta. Microsoft described Agent 365 and agentic security frameworks at Build earlier this year and released Copilot+ PCs with on-device Phi-Silica small language models, according to a Windows blog post from June 2.
Headwinds and Open Questions

Observation-based architectures face three technical barriers. Video pretraining at Induction's scale remains small compared to text pretraining runs that span trillions of tokens and hundreds of thousands of GPU-hours. Whether video scales as favorably as text when budgets grow tenfold is unproven. The benchmarks require independent verification. And the models still require action-conditioned finetuning or reinforcement learning to move from passive observation to active control. Photon-1's "use ChatGPT like a human" behavior emerged only after RL, the company wrote.
Regulatory frameworks are evolving. The EU AI Act's Article 50 transparency obligations for general-purpose AI providers took effect August 2, with the AI Office's enforcement authority beginning the same day. Providers training models with systemic-risk potential must notify the Commission. The U.S. has no comparable federal mandate, though NIST published an AI Risk Management Framework and a concept note for critical infrastructure in April. China issued guidance in May to regulate and boost "intelligent agents," defining their integration with cyberspace and the physical world. A U.S. advisory issued September 9 alleged that Chinese firms including DeepSeek, Alibaba, and Moonshot AI distilled capabilities from American labs' models since late last year.
Funding beyond Y Combinator has not been publicly disclosed. Crunchbase lists Induction as pre-seed with one to 10 employees. Dealroom shows a $125,000 line item from Y Combinator in August 2025, though the company has not announced a formal seed round. The company declined to disclose revenue, customer names, or next milestones.
If Induction's efficiency claims hold under independent scrutiny, the implications extend beyond computer-use agents. Robotics, autonomous vehicles, and digital twins all depend on models that understand physical dynamics from observation. The company's August intrinsic-discovery work suggests a path toward self-improving agents that generate their own curricula, an echo of DeepMind's AlphaGo Zero, which surpassed human champions by playing itself.
Whether two founders in San Francisco can sustain that research trajectory against labs with billion-dollar budgets remains an open question. Perhaps the more interesting question is whether 2027's foundation models will learn from the world or from each other's text. The answer may determine who leads the next wave of AI development, and how much it will cost.
