Bernt Børnich saw the future of humanoid robots not in a laboratory, but on the internet. When his company 1X launched its World Model Lab last June, the Norwegian startup joined what has become a peculiar arms race: teaching machines to understand the physical world by watching billions of hours of video that already exists online.
"Neo can now learn from internet-scale video and apply that knowledge directly to the physical world," Børnich told TechCrunch in January. His company's bet mirrors a broader conviction taking hold across the robotics industry—that the same scaling laws that turbocharged large language models can work for machines that need to navigate factory floors, assemble car parts, and manipulate objects in three-dimensional space.
The shift represents more than a technical pivot. It signals an industry-wide gamble that the answer to robotics' data problem sits in plain sight, scattered across YouTube cooking tutorials, factory livestreams, and dashcam footage. Where chatbots trained on trillions of tokens of text, robotics companies now pretrain on millions of hours of human activity, then fine-tune on comparatively tiny volumes of actual robot teleoperation data—sometimes just tens of hours.
Action-conditioned video models, architectures that predict what happens next when a robot or human takes a specific action, emerged as the state-of-the-art approach for manipulation and mobile robotics over the past two years. A January survey published on arXiv catalogued the explosion of activity. Google DeepMind released Genie 2 in December 2024, followed by Genie 3 in August 2025, then linked the system to Street View in May 2026 to simulate real streets and sidewalks, according to a company blog post. NVIDIA promoted its Cosmos world foundation model line and what it calls the World-Action Model taxonomy in June and July, positioning the technology as the backbone for its Isaac GR00T 2 humanoid initiative.
Runway, better known for its video-generation tools aimed at filmmakers, introduced GWM-1 on December 11 with variants including GWM-Robotics for synthetic training data and agent simulations. "To build a world model, we first needed to build a really great video model," Anastasis Germanidis, Runway's chief technology officer, explained in livestream remarks that day. "Predict pixels directly, at sufficient scale and with the right data."
The money flowing into the space suggests investors believe the approach will work. Skild AI, a Pittsburgh-based foundation-model company, closed a $1.4 billion Series C led by SoftBank on January 14, reaching a valuation above $14 billion, TechCrunch reported. The company's model pretrains heavily on human video plus simulation before transferring to physical robots. Xiaomi Robotics released a preprint in July claiming state-of-the-art embodied video generation and manipulation planning with its U0 "World Foundation Model."
Three technical advances unlocked the current wave, though the breakthroughs arrived in sequence rather than all at once. Transformer and diffusion architectures scaled to video generation at over 20 frames per second on consumer GPUs, making real-time or near-real-time robot planning feasible. Researchers demonstrated that models pretrained on passive human video could transfer to robot arms and mobile platforms with minimal robot-specific data—sometimes zero-shot, sometimes with just a handful of examples. The cost of labeling and curating video for what engineers call "physical plausibility" dropped as data vendors including Annotera and 1to1 built pipelines to tag contact events, object permanence, and scene-state changes at scale.
"Large world models will lay the foundation for robotics," NVIDIA's Jim Fan told QbitAI in a February interview. Yann LeCun, speaking at ETH Zürich in a talk posted last May, emphasized learned, action-conditioned predictive models in latent space for planning and counterfactuals. Academic labs published methods that keep a pretrained video planner intact and add embodiment-specific inverse dynamics modules, according to arXiv preprints. One framework, dubbed VERA, surfaced in multiple papers. Another, GE-Sim 2.0, offered a closed-loop simulation approach.
The economics helped, too. Morgan Stanley projected the humanoid robot market could reach roughly $5 trillion in annual revenue by 2050 in an April 2025 report. Goldman Sachs forecast a $38 billion market by 2035 in its January 2025 update. MarketsandMarkets, in a July 2025 release, projected the embodied AI market at $4.44 billion, growing to $23.06 billion by 2030 at a 39 percent compound annual growth rate. The International Federation of Robotics reported an installed base of roughly 4.66 million industrial robots by the end of 2024.

Whether the technology works in practice remains an open question, though early deployments offer clues. Figure AI deployed its Figure 02 humanoid at BMW's Spartanburg plant, where the robot supported production of more than 30,000 X3 vehicles over an 11-month deployment, BMW wrote in a press release. Figure 03 began work at the same facility late last June, according to BMW. Apptronik announced its first commercial deployment with Mercedes-Benz, production scaling with manufacturing partner Jabil, and a 90,000-square-foot "Robot Park" data facility on June 30, 2026, Forbes reported. Agility Robotics shipped Digit units to Amazon and other partners starting in 2024, with broader availability in 2025.
A cluster of Y Combinator companies now builds infrastructure around world models, though their scope varies considerably. One Robot, from the winter 2026 batch, focuses on what it describes as "world models for robot evals and training" and is hiring for a founding machine-learning role dedicated to world models, according to its YC profile. Vision Lab, from the summer 2026 batch, pairs first-person factory video with standard operating procedures to train foundation models. Lucid builds interactive, action-conditioned diffusion video models at over 20 frames per second on an NVIDIA 4090 to close simulation gaps. Asimov collects real-world human motion data for humanoids. Instance generates automated evaluations for robot policies from video. Verne Robotics and General Trajectory train robot arms and humanoid grasping models that learn new skills in hours rather than weeks, the companies claim.
Pantograph, a world-model startup, trained an interactive Minecraft simulator on internet-scale video and plans to transfer the approach to robots, the company wrote on its site in July. AWS, in a blog post, highlighted zero-shot manipulation claims from labs using internet-scale video priors plus tens of hours of unlabeled robot video, though the post noted industry claims vary widely.
The race now splits into three layers: pretraining scale, embodiment transfer, and what might be called enterprise hardening. Companies with capital—Skild AI, Figure, 1X, backed by SoftBank, BMW, and OpenAI respectively—are betting they can amortize the cost of billion-parameter models across multiple robot form factors. Startups including One Robot and Vision Lab are building the data curation, evaluation, and task-specific loops required to turn those big models into reliable factory workers.
Two headwinds loom, both with potential to slow or derail the current trajectory. Legal exposure from training on internet video without consent escalated over the past year and change. YouTubers filed class actions against Snap in January and February 2026, Runway on February 27, 2026, and Amazon on April 7, 2026, TechCrunch reported. Suits referencing NVIDIA Cosmos date to August 2024. YouTube introduced a "third-party training" opt-in setting in December 2024, allowing creators to block AI training use. The EU AI Act's transparency obligations for general-purpose AI models took effect August 2, 2025, with provider and deployer guidance published in July, the European Commission wrote. Training compute thresholds and disclosure requirements now apply to large video and world models shipped in the EU.
Safety standardization also tightened. UL published UL 3300:2024 for service robots on April 16, 2025, and the U.S. Consumer Product Safety Commission referenced the standard in guidance issued this year, according to ANSI's webstore. ISO 13482 for personal and service robotics saw updates, and ISO 10218 for industrial robots received similar revisions.

The technical roadmap ahead focuses on compute efficiency and latency. Research teams published methods in June to accelerate world-action inference to real-time on NVIDIA L40S and H100 GPUs, including Flash-WAM and GE-Sim acceleration techniques. "Robotics can't industrialize without an evaluation layer," One Robot wrote in a job posting for world-model engineers. That evaluation layer—task-specific simulations that reduce the need for expensive robot hours—will determine which startups scale and which stall.
The companies that solve data provenance, build defensible datasets, and ship models fast enough to guide a robot arm through a factory shift will capture the first wave of enterprise contracts. The rest will train models no one deploys. Børnich and his peers are wagering that the internet holds enough knowledge to teach machines how the physical world works. Whether that knowledge transfers cleanly from pixels to the factory floor remains the industry's defining question.
