Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Healthtech & Biotech iconHealthtech & BiotechJuly 24, 2026

Juno Bio Raises $3.8M, Opens First Women's Health Sequencing Lab

Juno Bio Raises $3.8M, Opens First Women's Health Sequencing Lab
Womens HealthFemtech+3
Healthtech & Biotech iconHealthtech & BiotechJuly 24, 2026

SiftMed Lands $5M CAD Seed to Cut Insurance Claims Review by 70%

SiftMed Lands $5M CAD Seed to Cut Insurance Claims Review by 70%
Ai AutomationClaims Processing+3

Founders Mentioned

Jonathan Li

Induction Labs

saas icon
SaaS

David Li

Induction Labs

saas icon
SaaS

Jonathan Li

Induction Labs

saas icon
SaaS

David Li

Induction Labs

saas icon
SaaS
SaaS iconSaaS
July 24, 2026
YcAgi ResearchWorld ModelsReinforcement LearningAi Infrastructure

Induction Labs Rethinks AI Foundation Models Through Curiosity

YC-backed two-person team builds 'imagination models' that learn from video observation, not labeled actions—claiming 30× efficiency gains over traditional approaches.

Induction Labs Rethinks AI Foundation Models Through Curiosity

In a narrow office somewhere in San Francisco, two engineers are nursing what might be either a contrarian insight or an expensive delusion: that the entire foundation model industry has been going about AI training the hard way.

While the hyperscalers—the Googles and Microsofts and OpenAIs of the world—pour upward of $700 billion into AI infrastructure this year and compete to meticulously label every mouse click and keystroke in their training datasets, Induction Labs has built what it's calling "imagination models." These systems, the founders claim, learn to navigate computers by watching unlabeled video. Much the way a child learns by observation, absorbing patterns and possibilities long before anyone asks them to explain what they're doing.

The assertion, laid out in a technical note published July 23, cuts against the grain. Their Photon-1 model, they say, outperforms a production-grade large language model trained with 30 times more compute. It costs a third as much to run. And it learned most of what it knows from simply predicting what happens next in screen recordings—no labeled actions required during the pretraining phase. Just video, curiosity, and later a dose of reinforcement learning to bridge understanding and execution.

It's the sort of claim that invites skepticism. Induction Labs, backed by Y Combinator and founded by Jonathan Li and David Li, hasn't yet subjected its results to independent benchmarking. No peer review. No public leaderboard entries. But the approach itself arrives at a moment when the economics of agentic AI are shifting rapidly, and the industry—restless, capital-rich, yet oddly constrained—is hungry for architectural breakthroughs that don't just scale compute but fundamentally rethink how models learn.

When Tokens Add Up Faster Than Expected

The foundation model market, pegged at $31.2 billion in 2026 and projected to hit $119.3 billion by 2031 according to Mordor Intelligence, is grappling with what you might call a cost paradox. Token prices are falling—Nvidia's Blackwell platforms promise fivefold software-only cost reductions within months of launch, and the Global Compute Price Index documents rapid declines in inference pricing across providers. Good news, on its face.

Yet for agentic systems that execute long, multi-step workflows, those tokens accumulate with surprising speed. A McKinsey analysis published earlier this year noted that while per-token costs drop, the sheer volume of interactions in agent-driven tasks changes the underlying math. You're not just generating a chatbot response anymore. You're orchestrating sequences of actions, each one requiring the model to observe, reason, and act—over and over.

Adoption is accelerating. Gartner reported in April that 17% of organizations have deployed agents, with more than 60% planning rollouts within two years—the fastest uptake curve among emerging technologies. But Forrester, in a June analysis that felt almost cautionary, warned that while three-quarters of companies are exploring agentic AI, few are operating at scale. The firm predicts that over 40% of agent projects could be canceled by the end of 2027. A Deloitte survey found that 80% of organizations scaling agents lack mature guardrails, the kind of governance infrastructure that separates experiments from production systems.

The bottleneck isn't just oversight or risk management, though those matter. It's efficiency. Current computer-use models—Google's Gemini Computer Use, OpenAI's Codex variants with Windows control, Microsoft's digital employee frameworks—are trained on labeled demonstrations. A human performs an action on screen; the model learns to map observation to behavior through supervised fine-tuning on millions of annotated examples. It works. But the data collection is expensive, the models are large, and latency creeps upward as task horizons extend.

OSWorld 2.0, released June 28, introduced 108 long-horizon workflows specifically to test whether agents can handle streaming interactions and cross-application reasoning without breaking down. The benchmark's existence signals where the field is headed: toward tasks too complex and open-ended for brittle, narrowly trained systems.

Observation Before Action

Induction Labs starts from a different premise. What if models learned to understand the world first, and figured out actions second?

Their architecture—what they're calling an "imagination model"—predicts future states in a latent space, trained initially on video alone with no action labels. Photon-1 is a sparse 106-billion-parameter Mixture-of-Experts transformer with 32,000 tokens of context. It was pretrained on 575 million frames extracted from roughly two million screen recordings, filtered from a pool of two billion public videos. That translates to something like 18 years of video at one frame per second, pretrained for a single epoch using an estimated 30,000 H200 GPU-hours—about 4.4×10²² FLOPs, according to the company's note.

The vision encoder compresses each frame into 960 discrete tokens using Finite Scalar Quantization, an eight-dimensional vector space with five values per dimension. Induction claims this achieves roughly 100 times better compression than multimodal OCR baselines while preserving text, layout, and state information. That matters in desktop environments where pixel-level detail is less important than semantic structure—what's happening, not just what's visible.

During pretraining, the model learns only to predict what comes next. No mouse clicks annotated. No keyboard inputs logged. Just: given this sequence of screen states, what's likely to appear in the next frame?

It's a world model in the tradition of research from Meta's V-JEPA 2 (released June 2025) and the DINO-world family, which demonstrated that self-supervised video prediction can enable zero-shot planning and physical reasoning. Yann LeCun argued in a May 29 talk at ETH Zurich that world models represent the key to AI's next breakthroughs. Andrej Karpathy echoed the theme in blog posts throughout 2026: understanding precedes action.

But Induction adds a step. After pretraining, they introduce actions through a modest fine-tuning phase and scale up with online reinforcement learning across diverse Linux desktop environments. The model already knows what screens tend to do; now it learns how its actions make those predicted futures real. The company reports that Photon-1 "learns to use ChatGPT like a human" after reinforcement learning and transfers to games like checkers and physics simulations with limited fine-tuning.

All of these results, however, remain internal. Unverified.

The Numbers That Matter—If They Hold

Digital illustration for article section "The Numbers That Matter—If They Hold" in "Induction Labs Rethinks AI Foundation Models Through Curiosity" - A conceptual macro photography image representing a massive 30-fold reduction in AI compute training...

The headline claim—30 times fewer FLOPs than a comparable production LLM—sits at the heart of Induction's thesis. They compare Photon-1's ~4.4×10²² FLOPs of training compute to an unnamed baseline and report superior performance on internal computer-use benchmarks. They also claim serving costs three times lower, a figure that would matter substantially in production environments where inference economics determine whether deployment makes financial sense.

These are bold assertions for a team of two with no peer-reviewed validation. OSWorld's public leaderboard doesn't list Photon-1 as of late July. No third-party evaluations have surfaced. An earlier product, Axiom 1, announced in August 2025, claimed 60.2% success on OSWorld-Verified with 16 times lower latency per action than GPT-5 in their internal reproductions. But again: company-published numbers without independent confirmation.

Still, the approach aligns with broader research trends, which is perhaps what makes it worth taking seriously. Meta's CLAW paper, published June 2, demonstrated continuous latent-action world models learned from action-free videos. DINO-WM, presented at ICML 2025, showed that world models built on pretrained visual features enable zero-shot planning. Microsoft Research's CORPGEN framework, released February 26, tackled multi-horizon corporate tasks and found that hierarchical planning with memory isolation improves multi-task performance—a different architecture, but solving the same fundamental problem of making agents reason over long sequences without degrading.

The pattern is clear enough. The field is moving toward models that can predict, plan, and act across extended time horizons, and doing so efficiently enough to deploy at scale. Induction's architecture, if the numbers hold under scrutiny, offers one possible path forward.

Curiosity as Infrastructure

The language Induction uses—"intellectual curiosity," "intrinsically motivated to learn"—might sound like startup marketing. But it draws from a lineage of reinforcement learning research dating back to Intrinsic Curiosity Modules (Pathak et al., ICML 2017), Random Network Distillation (Burda et al., 2018), and Go-Explore (published in Nature, 2021). All explored how agents can learn to navigate environments without external rewards, driven instead by novelty and information gain.

Jonathan Li, who previously worked on post-training and reinforcement learning at Cohere and has publications in ACL 2023 and legal AI workshops, appears to be applying those principles to computer-use pretraining. His co-founder David Li, a University of Waterloo software engineering graduate, brings the systems background to make it scale. The company's website, updated in July, describes their mission as building intelligence "intrinsically motivated to learn," scaling "curious" foundation models that learn broadly from observation and act in the world—especially on computers.

The curiosity framing matters because it reframes the data problem. Instead of needing millions of labeled demonstrations—expensive to collect, narrow in coverage—you need video that captures the dynamics of what happens on screens. Those two million screen recordings didn't require human annotation. The model learns the statistical regularities of how applications behave, how users navigate, what states tend to follow what states. Then, when actions are introduced via fine-tuning and reinforcement learning, the model already has a rich prior about how the world works.

Whether this generalizes beyond internal benchmarks to real production environments remains very much an open question. But the efficiency claim, if validated, would matter in an industry where inference costs and agent reliability are the twin gates to deployment at scale.

Regulation Meets Innovation

Induction's technical note arrived just as the European Union's AI Act general-purpose AI provisions came into force on August 2. GPAI models now face transparency, cybersecurity, incident-reporting, and evaluation obligations if they're deployed in Europe. The UK's AI Safety Institute is pushing harder on evaluations, and NIST's AI Risk Management Framework is evolving toward sector-specific profiles for critical infrastructure.

For a two-person team with a novel architecture and no independent evaluations, that regulatory environment poses challenges—perhaps more than the founders initially anticipated. Deloitte's survey found that organizations scaling agents rapidly often lack mature governance. Forrester's June analysis warned that budget and oversight gaps could derail projects before they reach production. TechRadar Pro noted in July that agent cost models differ from typical SaaS pricing, creating budgeting surprises for enterprises unaccustomed to usage-based inference economics.

At the same time, the market is maturing around them. Benchmarks like OSWorld 2.0 emphasize long-horizon workflows, streaming interaction, and cross-application reasoning—exactly the scenarios where unified, low-latency policies shine. Microsoft's Agent Framework, OpenAI's adoption of the Model Context Protocol, and the rise of infrastructure providers like Browserbase signal that the ecosystem is consolidating around standards. Google's Gemini Computer Use documentation encourages builders to use Playwright or cloud VMs. ServiceNow claims 37% automation of support workflows in internal blog posts from earlier this year. UiPath's Autopilot releases through late April target agentic orchestration.

The shift from chat to action is accelerating, and the companies that crack the efficiency problem—building models that can handle complex, multi-step tasks without ballooning costs or latencies—stand to gain an edge in a crowded field.

What's Missing

Digital illustration for article section "What's Missing" in "Induction Labs Rethinks AI Foundation Models Through Curiosity" - A minimalist, conceptual visualization of unverified claims and missing financial data, featuring a ...

None of this obscures the fact that Induction Labs' claims remain unverified. Y Combinator is listed as a backer, with a $125,000 seed entry dated August 2025 appearing on Dealroom, and the company's announcement from August 19, 2025 explicitly states "Born in SF, backed by Y Combinator." But the official YC Startup Directory doesn't list Induction Labs as of late July, and the batch year remains unconfirmed. The founders appear in informal LinkedIn mentions among Canadian YC alumni, but that's anecdotal at best.

More importantly, the performance numbers are self-reported. No public leaderboard. No third-party benchmark runs. No peer review. The claims about outperforming models with 30 times more training compute and serving at a third of the cost would be transformative if true, but they require independent validation before the broader AI research community can assess them with any confidence.

That's not unusual for early-stage startups publishing technical notes rather than academic papers. But it does mean investors, enterprise buyers, and developers need to treat the numbers with appropriate caution. A paper published May 18 on "Agent Meltdowns" demonstrated how injected errors cause cascading failures in GPT, Grok, and Gemini stacks, underscoring that even well-tested systems can behave unpredictably under real-world conditions. For a novel architecture with no public track record, the unknowns are considerably larger.

The Road Ahead

Digital illustration for article section "The Road Ahead" in "Induction Labs Rethinks AI Foundation Models Through Curiosity" - A conceptual miniature scene representing the road ahead for foundation models, featuring a clean, a...

If Induction Labs can deliver on its efficiency claims and scale beyond internal benchmarks, the implications ripple outward. A 30× training efficiency gain would lower the barrier for new entrants to build competitive foundation models for computer use. A 3× serving cost reduction would make long-horizon agent workflows economically viable for a broader set of applications. And a pretraining approach that learns from observation rather than labeled actions could accelerate data collection and domain adaptation across industries.

But the path from technical note to production deployment is, as always, littered with obstacles. The EU AI Act's GPAI obligations took effect this month, and companies deploying models with systemic risk must meet rigorous evaluation and transparency standards. The UK's AISI is pushing toward more formalized model assessments, and NIST is developing sector-specific profiles. For a two-person team, navigating that regulatory landscape while scaling a novel architecture will require resources and institutional know-how they may not yet possess.

There's also the question of market positioning. Google, Microsoft, OpenAI, and Anthropic have poured billions into computer-use agents and maintain extensive benchmark leads. Startups like Cognition's Devin have raised substantial funding and built enterprise traction. For Induction to carve out a niche, it needs either to outperform incumbents on public benchmarks, partner with larger players to embed its models in their stacks, or target use cases where efficiency and adaptability matter more than absolute capability.

The two founders are hiring, according to their website, signaling intent to scale beyond the current team. Whether they can translate a compelling research direction into a sustainable business remains uncertain. The foundation model market is projected to grow at 30.8% annually through 2031, and the autonomous agents segment is expanding even faster—estimates range from 19% to 51% CAGR depending on which analyst you ask. There's room for new architectures, especially ones that challenge the prevailing wisdom about how models should learn.

For now, Induction Labs represents a bet that observation teaches more than instruction, that curiosity can replace annotation, and that rethinking the fundamentals of pretraining can unlock step-change efficiency gains. Whether a two-person team in San Francisco can make imagination scale is a question the industry will be watching closely. The technical note is out there. The real test begins when the numbers face independent scrutiny.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Juno Bio Raises $3.8M, Opens First Women's Health Sequencing Lab
  • SiftMed Lands $5M CAD Seed to Cut Insurance Claims Review by 70%
  • Living Compute: Startups Train Human Brain Cells to Run AI
  • Precision Biopesticides Promise Resistance-Free Future for Agriculture
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.