The bill arrives quickly when you're training AI agents to navigate software systems. Frontier language models charge thousands of dollars per million tokens. Agentic systems learn through trial and error, each mistake another line item. For startups building in this space, the meter is always running.
Two founders from Y Combinator's latest cohort think they've spotted an arbitrage opportunity in an unlikely place: the autonomous vehicle industry. Their pitch? Borrow the simulation playbook that AV companies used to stress-test robo-taxis in virtual worlds, and apply it to AI agents learning to work with databases, APIs, and enterprise software.
Experiential Labs claims it can cut AI agent development costs by 90 percent while matching the performance of far more expensive frontier models. It's the kind of promise that raises eyebrows—the claim lacks independent verification and comes from the company's own statements. But the timing is hard to ignore. As enterprises race to deploy AI agents and McKinsey estimates that the combined 2026 capex for leading hyperscalers exceeds $700 billion, with the majority directed towards AI infrastructure, the cost-per-interaction math is becoming impossible to ignore.
An Unusual Pedigree
Kion Fallah and Silen Naihin don't fit the typical AI agent founder template.
Fallah spent years at Waabi, the Toronto-based autonomous vehicle company, where he led mixed-reality simulation work. Waabi's MixSim platform became known in AV circles for blending real sensor data with generated traffic scenarios—a way to test edge cases without putting cars on actual roads. Naihin, meanwhile, was involved with AutoGPT and co-founded Stackwise, another YC company from an earlier batch.
Their shared hypothesis: the autonomous vehicle industry figured out how to train safety-critical systems in simulation over the course of a decade. Those same techniques—reconstructing environments from sensor traces, generating plausible what-if scenarios, running thousands of cheap iterations—could work for AI agents navigating digital environments instead of physical ones.
It's not an obvious connection. Self-driving cars operate in three-dimensional space with physics constraints and real-time sensor fusion. Software agents parse JSON responses and query databases. But both face a similar problem: real-world testing is expensive, slow, or dangerous. Simulation offers a faster, cheaper alternative, assuming you can build accurate-enough models.
Experiential Labs has released an open-source repository called "world-model-harness" on GitHub. The code describes a runtime and optimizer that constructs world models from collected execution traces, then uses those simulated environments to test and refine agent behavior without burning API tokens. The repo has attracted over 100 stars. The company also operates a hosted platform at platform.experientiallabs.ai, though it hasn't disclosed customer names or pricing.
In June, Fallah and Naihin published two research papers—one accepted as an ICML spotlight on sparse autoencoders for interpretability, the other titled "CLaaS: Continual Learning as a Service," proposing methods for online learning in deployed language model agents. The company has been involved in publishing research papers, which could be seen as contributing to both infrastructure and the academic space. Whether that dual identity helps or distracts remains to be seen.
World Models Go Mainstream

World models used to be confined to robotics labs and reinforcement learning theory. Not anymore.
The concept has expanded to cover video-prediction networks, physics simulators, generative models that synthesize sensor data—pretty much any system that learns to predict what happens next. And in the past year, the idea has jumped from academic papers to product launches.
Waymo introduced its "Waymo World Model" in early February, leveraging Google DeepMind's Genie 3 architecture to generate controllable camera and lidar simulations. The system can take dashcam footage and convert it into multi-sensor rollouts, adjust weather conditions via text prompts, and test counterfactual driving decisions. Waymo emphasized an "efficient variant" for longer scene generation—a tacit acknowledgment that these models are computationally expensive.
A week later, World Labs—founded by Fei-Fei Li and backed by a $1 billion funding round as of February 2026 that included $200 million from Autodesk—announced its first product, Marble. The software generates editable 3D environments and is exploring integration with Autodesk's CAD workflows. Li's framing positioned world models as infrastructure not just for robotics but for digital design, entertainment, and beyond.
Nvidia, unsurprisingly, has been aggressive. Its Cosmos world models, Omniverse digital twin libraries, and Isaac robotics tools all received major updates in mid-2026. In June, the company released an open-source collection of agent skills and reconstruction tools. In July, it expanded its Agent Toolkit with Omniverse libraries for building simulation-ready environments. The pitch to developers: stop reinventing simulation infrastructure and focus on application logic instead.
A June research preprint from Nvidia, titled "OmniDreams," claimed real-time generative world modeling for closed-loop autonomous vehicle simulation using Cosmos priors and 21,000 hours of driving data. It's early work, not yet independently validated, but it signals how rapidly academic concepts are being commercialized.
The Economics of Simulated Worlds
The cost case is compelling, at least in theory.
McKinsey reported in late June that custom inference silicon can reduce cost per token by 70 to 80 percent for specific workloads. That's on the inference side. World models promise savings on the training and evaluation side by replacing expensive real-world interactions—or expensive API calls—with cheaper simulated ones.
The synthetic data market offers a rough parallel. Multiple reports from the year project strong growth, with Global Industry Analysts forecasting the sector will reach $3.9 billion by 2032, up from what they estimate was a $291 million baseline in 2025. Vendor forecasts should always be taken with skepticism, and the methodologies vary widely. But the directional signal is clear: companies want to generate training and test data computationally rather than collecting it in the wild.
Applied Intuition, which builds simulation platforms for autonomous systems, raised a $250 million Series E at a $6 billion valuation in March 2024, then a Series F that pushed its valuation to $15 billion in June 2025. Those numbers are more than a year old now, but they illustrate the capital flowing toward companies promising to de-risk development cycles through simulation.
Foretellix, a verification platform, claimed in January 2024 that partnerships with companies like Nuro can "slash development time in half and save hundreds of millions" through scenario-based testing. These are vendor assertions, not peer-reviewed findings, but the narrative resonates: in industries where real-world testing is slow, dangerous, or prohibitively expensive, simulation becomes a strategic advantage.
Enterprise Adoption and Growing Pains

Gartner's Hype Cycle for Agentic AI, released this year, reports that 17 percent of organizations have deployed AI agents, while more than 60 percent plan to deploy within two years. That's a steep adoption curve from a low base. IDC's FutureScape predicts half of enterprises will use AI agents by next year, with agentic systems approaching half of all AI spending by 2029.
Forrester's research from mid-year highlights both enthusiasm and friction. Nine in ten U.S. marketing agencies are using AI to cut costs, according to one survey. But governance challenges, integration complexity, and security concerns remain primary obstacles. Half of security leaders cited agentic AI as a worry in recent research.
The regulatory landscape is tightening in parallel. The EU AI Act's general application date arrived in early August, with obligations for general-purpose AI rolling into enforcement cycles through 2028. NIST updated its Generative AI Profile in April and launched a GenAI Evaluation Program in May. Enterprises deploying agentic systems or building general-purpose AI tooling will need documentation, transparency mechanisms, and safety processes aligned with these frameworks.
For startups like Experiential Labs, this creates both opportunity and obligation. The opportunity: enterprises want testbeds to validate agents before production deployment. The obligation: those testbeds need to support compliance reporting, auditing, and risk assessment—not just performance optimization.
Other startups are staking out adjacent territory. One Robot, a YC company from an earlier batch backed by Accel, is hiring for "founding machine learning world models" roles to build action-conditioned video and dynamics models for robot evaluation. Chronicle Labs is building "staging environments for enterprise AI agents." Instance focuses on "automated evals for robot policies." The YC portfolio alone suggests world models and agent evaluation infrastructure are becoming a recognized category, not just a research curiosity.
Open Questions
A survey paper titled "The Cost of Dreaming," published in mid-July, catalogs computational bottlenecks in generative and latent world models. The authors argue for hierarchical architectures, dynamics-aware quantization, and standardized efficiency reporting to push toward sub-100-millisecond on-device models. It's a research community acknowledgment that current world models are computationally expensive—and that efficiency gains are necessary before these tools can be widely deployed.
Another July paper, "Imagined Rollouts are Kinematic, Not Dynamic," diagnoses failure modes in long-horizon world-model rollouts. The authors found that many models can predict plausible short-term sequences but lose physical fidelity over longer timescales. It's a useful reminder that simulation is not reality, and that performance gaps between simulated and real-world behavior remain an unsolved problem.
Sequoia Capital published an essay in late April arguing for "training general intelligence in pixel space" as a path to agentic systems that operate computers from raw visual input. It's thought leadership, not peer-reviewed research. But it reflects a broader industry conviction: that agents will learn to navigate complex environments not through hand-coded interfaces but through observation and interaction with simulated—and eventually real—worlds.
The infrastructure stack is converging quickly, perhaps more quickly than anyone expected a year ago. Nvidia's open libraries, Waymo's multi-sensor world model, World Labs' 3D environment generators, and startups like Experiential Labs betting on AV-grade simulation for digital agents—all suggest a future where simulation isn't a niche tool for robotics but a standard layer in the AI development process.
The Proof Is Still Out

Whether Experiential Labs' 90 percent cost reduction holds up under independent scrutiny remains an open question. The company is early-stage. The claim is bold. The details are sparse.
But the problem it's addressing—how to train, test, and refine AI agents without incurring prohibitive compute and API costs—is undeniably real and growing more urgent as enterprises race to deploy agentic systems at scale.
The next year will clarify which approaches deliver on their promises, which claims were overblown, and which companies successfully bridge the gap between simulation and production deployment. For now, the bet is straightforward: if you can simulate it cheaply, you can afford to fail fast and iterate often.
In an industry where one botched deployment can cost millions in compute or irreparable reputational damage, that's a proposition worth paying attention to. Even if the numbers don't quite add up to 90 percent.
