The robot grips the part. Executes the sequence. Fails. In simulation, it tries again. In the lab, it tries again. But on the factory floor—where the lighting shifts, where parts arrive with real-world variance, where the workflow bears only passing resemblance to training data—it stumbles.
This gap, the maddening distance between laboratory promise and industrial reality, has shadowed robotics for years. Now, as foundation models begin making their way onto production lines—BMW deploying humanoids at its South Carolina plant, Toyota signing service agreements, installation projections climbing toward a million units annually by decade's end—the industry confronts a bottleneck that might surprise engineers accustomed to racing for compute power: data.
Not just any data. Real factory data, captured at scale, structured for machines that need to learn what human hands already know.
Enter Vision Lab, a startup founded in 2025 and operating out of San Francisco, which announced a $6 million seed round on June 3, 2026 to attack what it calls the "physical AI data wall." The company reports working with more than 2,000 factories spanning 50-plus industries across 27 countries, recording first-person industrial workflows and pairing them with standard operating procedures—building, in essence, the training sets that robotics foundation models will need to function beyond the lab.
Race Capital led the round. Y Combinator, Foothill Ventures, 500 Global, and a handful of angels from Google DeepMind, xAI, OpenAI, and NVIDIA came along for the ride.
"Physical AI will require a different kind of data infrastructure than language models," a Race Capital general partner observed when the deal was announced. "Robotics teams need access to real industrial environments, real workflows, and real human task execution at scale."
Whether Vision Lab becomes that infrastructure—or merely one player in what's shaping up to be a crowded field—depends on execution, factory relationships, and a bet that the next era of robotics will be gated less by algorithmic breakthroughs than by access to ground truth.
The Numbers Tell Part of the Story
Industrial robot installations hit their second-highest year on record in 2024, according to the International Federation of Robotics. China alone accounted for roughly 54 percent of those installations, pushing its operational stock past 2 million units. The United States installed about 34,200 robots in 2024—a 9 percent drop from 2023, though still a significant volume.
Global robot density reached 162 units per 10,000 manufacturing workers in 2023, double the 74 recorded in 2016.
Manufacturing job openings in the U.S. climbed to approximately 474,000 in April 2026, up 26 percent year-over-year. Labor gaps, in other words, remain wide—and automation interest follows those gaps. Wissen Research projects the global industrial robotics market will grow from $34 billion in 2024 to $70.6 billion by 2030, a 13 percent compound annual growth rate.
Yet the vast majority of industrial robots still run pre-programmed routines. Fixed. Brittle. The leap to adaptive, learning-capable systems—machines that handle variance, understand context, generalize across tasks—depends on foundation models trained on diverse, messy, real-world data.
And that data remains remarkably scarce.
Three Forces Converge
The first is the maturation of vision-language-action models. Google DeepMind's RT-2, announced in July 2023, demonstrated that multimodal transformers could translate vision and language into robot control. OpenVLA, an open-source project from June 2024, gave researchers a foundation for building cross-embodiment policies. NVIDIA's Project GR00T, unveiled at GTC 2024 and expanded with the open GR00T N1 model in 2025, offered a platform for training humanoids and general-purpose robots.
The second force: deployments are accelerating. BMW began piloting Figure AI's humanoid robots at its Spartanburg facility in 2025, moving north of 90,000 parts across more than 1,250 operating hours, supporting production of over 30,000 vehicles. In February 2026, the automaker announced a European pilot in Leipzig and plans for a "Center of Competence for Physical AI in Production."
Toyota Motor Manufacturing Canada signed a robotics-as-a-service agreement with Agility Robotics that same month, following earlier pilot programs. Mercado Libre announced a commercial deal with Agility in December 2025 for warehouse operations in Texas, with potential expansion across Latin America.
The third driver, perhaps less visible but no less critical, is regulatory convergence. The European Union's AI Act entered force in August 2024, with staged applicability rolling out through 2026 and 2027. The EU Machinery Regulation, set to replace the prior directive on January 20, 2027, addresses AI and cyber risks in industrial equipment. In the United States, ANSI/RIA published the revised R15.06-2025 standard—the most significant update to U.S. industrial robot safety requirements in over a decade—bringing alignment with ISO 10218-1 and -2.
These frameworks impose documentation, risk management, and human oversight obligations that demand transparency in model behavior. Which, in turn, requires representative training data—the kind that reflects not just laboratory conditions but the chaotic realities of a working factory floor.
From Lab to Line

BMW's deployment offers a case study in what happens when the data infrastructure actually supports the model. "Motion sequences trained in the laboratory could be quickly transferred into stable shift operation," a BMW production executive noted in the company's February announcement. The Spartanburg pilot ran up to 14 hours per day across multiple shifts, suggesting the humanoid's task repertoire aligned reasonably well—perhaps more than the founders expected—with real factory conditions.
Vision Lab's approach differs from the general-purpose robotics datasets emerging from research labs. The company doesn't collect data in isolation. It embeds capture infrastructure directly in operating factories, pairing egocentric and exocentric video with process knowledge at the standard operating procedure level. Dense temporal labeling. Human-in-the-loop quality checks. Anonymization protocols for competitive intelligence concerns.
The company claims its proprietary temporal vision-language model achieves 2.5 times better performance than Gemini on temporal understanding in industrial settings. That metric is self-reported, of course.
Vision Lab says it serves "multiple frontier AI labs, including three of the Magnificent Seven tech giants," though client names remain under NDA. Partner factories have reportedly earned more than $1 million in cumulative data-licensing revenue, according to company statements cited by Techsauce in June 2026. The eight-person team is led by Tanachart (James) Kujareevanich, a former McKinsey operations consultant with an MBA from MIT, who went through Y Combinator's Spring 2025 batch.
"Factories represent where a large share of future robots will be deployed," Kujareevanich told Techsauce, "yet industrial environments remain vastly underrepresented in existing datasets."
Other players are tackling pieces of the same puzzle. Covariant announced its RFM-1 robotics foundation model in March 2024, building on multimodal fleet data from warehouse and e-commerce deployments. A German distributor, Obeta, has run a Covariant-powered KNAPP Pick-It-Easy robot in production for over two years, achieving 600 objects per hour at 99 percent accuracy with 70 percent SKU coverage—solid numbers in a high-throughput environment.
Amazon crossed the one-million-robot milestone and is developing a generative AI foundation model for its warehouse fleet, though details remain closely held.
Meanwhile, venture capital is pouring into robotics AI infrastructure. Physical Intelligence was reportedly in talks at an $11 billion valuation in March 2026, per TechCrunch, following a $5.6 billion valuation round reported by Bloomberg in November 2025. Skild AI hit $14 billion in January 2026. Apptronik raised $935 million at a valuation above $5 billion in February.
The money is moving fast. The question is whether the data will keep pace.
Public Datasets, Private Gaps
The Open X-Embodiment dataset, released in October 2023 and updated through 2024 and 2025, provides over one million robot trajectories across 22 embodiments. Ego4D and Ego-Exo4D offer large egocentric datasets with multimodal annotations, useful for understanding human manipulation and hand-object interactions. MVTec AD 2, published in 2026, advances industrial anomaly detection benchmarks for inspection tasks.
These public datasets are valuable for research. They're not, however, what industrial deployments require. They don't capture the SOP-level, production-context workflows that a robot needs to navigate a real shift at a real factory with real tolerances and real consequences for failure.
Vision Lab believes that the next wave of robotics foundation models will be constrained not by algorithmic innovation but by access to factory data at scale. The company's planned expansion into Southeast Asia aligns with regional manufacturing growth and the ongoing diversification of global supply chains.
Whether Vision Lab becomes the "default industrial data layer," as it aims to, depends on execution: maintaining factory partner relationships, scaling capture infrastructure, ensuring data quality, and navigating intellectual property and privacy concerns in competitive industrial settings.
The Risks Are Real

Factory partners may balk at extended data capture if competitive intelligence leaks become a concern—or even a perceived risk. Frontier AI labs may prefer to build proprietary pipelines rather than rely on third-party providers. Regulatory frameworks, particularly the EU AI Act's requirements around data governance and transparency, could complicate cross-border data flows in ways that are difficult to predict.
And then there's the simple question of whether the data will generalize. Industrial environments are notoriously heterogeneous. What works in automotive may not translate to electronics assembly. What works in electronics may fall apart in food processing. The company will need to prove that its datasets—and the models trained on them—can bridge those gaps.
But the opportunity, at least in outline, is clear.
Deloitte's 2026 Technology, Media, and Telecommunications Predictions projects conditional growth in robot installations, contingent on AI model maturity, data integration, and security infrastructure—precisely the terrain where companies like Vision Lab are staking claims. The firm suggests that industrial robot installations could add roughly 100,000 units per year from 2027 through 2030, potentially reaching one million annual installations by 2030, though these projections depend on infrastructure readiness.
BCG's May 2026 Factory of the Future report, based on a survey of 1,000 manufacturers, found that adoption speed varies by local labor markets, energy costs, and capital expenditure constraints. Countries with high robot density, such as South Korea, already show advanced trajectories. For laggards, the challenge isn't just acquiring robots—it's ensuring those robots can learn and adapt in environments that deviate from training conditions.
As humanoids move from pilots to production, as vision-language-action models transition from academic papers to deployed systems, and as manufacturers face the twin pressures of labor shortages and quality demands, the companies that control high-fidelity industrial training data may hold leverage comparable to what compute providers held in the language model era.
Vision Lab is placing an early bet on that thesis. The $6 million seed round gives the team runway to test whether factories will pay for adaptive robots—and whether robotics builders will pay for the data to train them. The answer, as with most bets in emerging markets, won't be clear for a while.
