Picture a warehouse in Sacramento on a Tuesday afternoon. Fluorescent lights hum overhead. Safety signs from last year partially block the shelf labels. A worker reaches for a part, flips it, sets it down. Nothing remarkable happens. Except that somewhere, a robotics company is desperately wishing it had video of exactly that moment.
The problem haunting the physical AI industry isn't whether humanoid robots can execute tasks in pristine lab conditions. It's whether they can handle the messy, cluttered, unpredictable environments where actual work happens. And the solution, it turns out, might require turning every factory floor and distribution center into a data collection operation—filming workers as they go about their shifts, capturing the mundane choreography of human labor that no simulation can fully replicate.
This is spawning a small but growing cohort of startups betting that industrial workplaces represent the next great training corpus. Not code repositories or image databases, but footage of people doing their jobs.
Among them is Praxis, a team that emerged from Y Combinator claiming to have embedded data collection infrastructure across tens of thousands of workers on five continents. The founders—Dev Karpe, Tommy Li, and Rohan Seelamsetty—describe their work as capturing egocentric video, 3D scans, and other multimodal data from inside factories, warehouses, and residential work sites. The pitch: transform every company into a data vendor for the robots that will someday replace, or at least augment, their human workforce.
Whether that vision materializes remains an open question. But the underlying premise reflects a genuine bottleneck in an industry awash with capital and optimism.
A Market Searching for Its Training Set
The humanoid robotics sector is in a peculiar moment. Forecasts are stratospheric. Bank of America Institute projected that annual shipments could reach 1.2 million units by 2030, climbing to 10 million by 2035, with costs per unit falling from roughly $35,000 to under $17,000 within the decade. Grand View Research estimated the market at $2.4 billion in 2025, forecasting growth to $40.5 billion by 2033—a compound annual rate of 38.2%.
IDC reported around 18,000 humanoid units shipped globally in 2025, a year-over-year jump of more than 500%, driven largely by Chinese manufacturers. Goldman Sachs analysts suggested in a June 2026 report that Korean supply chains alone might represent 30% of global production by 2035.
The numbers sound staggering. Yet nearly every forecast circles back to a single constraint: data. Or more precisely, the lack of it.
Goldman Sachs flagged this in research dating back to early 2024, and the observation has only become sharper since. Unlike large language models, which gorged on the internet's text, or vision models trained on vast image databases, robotics has no equivalent corpus. There's no centralized repository of millions of real-world manipulation tasks, filmed from a first-person perspective, annotated for intent and action, formatted for consumption by neural networks.
Open X-Embodiment, launched in late 2023 and expanded through 2025, aggregated more than 60 datasets spanning 22 robot types and over 500 skills. It's a serious effort. It's also nowhere near sufficient.
The robotics community has begun to coalesce around an answer that feels almost too obvious: film humans.
The Logic of Egocentric Video
Humans already perform the tasks that robots need to learn, in the environments where robots will eventually operate, under conditions no simulation can fully capture. Strap a camera to their heads. Record them working. Train the models on that.
NVIDIA's EgoScale project, unveiled in February 2026, made the case at scale. The researchers compiled more than 20,000 hours of action-labeled egocentric video and used it to pretrain a vision-language-action model. The key finding was a log-linear scaling law: more human video correlated directly with better robot performance on dexterous manipulation. It was the first time anyone had demonstrated a quantifiable relationship between video volume and robot capability.
Other research pushed the thesis further. HumanEgo, published mid-2026, reported success rates above 90% on real-world tasks using just 15 to 30 minutes of human video per task for zero-shot transfer. EgoLive focused specifically on home service and retail routines—exactly the domains where commercial humanoid deployments are expected to land first.
The practical implications are difficult to overstate. Traditional robot training involves building or acquiring hardware, scripting tasks, collecting teleoperation data, iterating on control policies, and hoping the system generalizes to new settings. Training from human video collapses that workflow into: point a camera at someone doing the job, then let the model extract the patterns. The efficiency gain isn't incremental. It's a step change.
NVIDIA's GR00T foundation model, released in spring 2026, drew on more than 20,000 hours of human egocentric footage from EgoScale, supplemented by robot demonstration data. Cross-embodiment learning—where policies trained on one robot type transfer to another—showed improvements of 24% to 45% in success rates when datasets spanned diverse tasks and environments. Variety, it turned out, mattered as much as sheer volume.
Real Deployments, Real Data Loops

The deployment numbers remain modest, but they're no longer hypothetical.
Figure AI partnered with BMW to place Figure 02 humanoids on the manufacturing line at the automaker's Spartanburg plant in 2025. BMW confirmed these as the industry's first humanoid production deployments, with the robots supporting assembly work on more than 30,000 vehicles. Figure 03 units arrived at the same site months later, building on over 1,250 operational hours accumulated by their predecessors.
What BMW extracted from the experience, according to Figure, was operational and failure-mode data that fed directly into hardware revisions and policy updates. That closed loop—deploy, observe edge cases, retrain, redeploy—is the flywheel every robotics company is chasing.
Agility Robotics opened a new software hub in Fremont in mid-2026, explicitly focused on training and testing AI for its Digit humanoid. The company published a case study detailing its use of AWS infrastructure to scale model training and indicated plans to hire roughly 200 people for AI-related roles.
Meanwhile, the KION Group, NavVis, and NVIDIA announced a collaboration centered on scanning large industrial sites to create digital twins, then using NVIDIA's Isaac Sim to generate synthetic sensor data—lidar, RGB, depth, segmentation masks—for training autonomous forklifts and mobile robots. The convergence of physical scanning and simulation suggests a hybrid training regime: real-world captures seed virtual environments that produce endless permutations.
These aren't science projects. They're production systems with procurement cycles, compliance reviews, and quarterly performance targets. And they share a common appetite: more data, from more environments, under more conditions.
A Cottage Industry Takes Shape
Praxis is hardly the only player trying to monetize that hunger.
Cortex AI, which launched out of Y Combinator several months earlier, positions itself as a marketplace that compensates workplaces for hosting egocentric capture sessions and teleoperation trials. Defined.ai announced multimodal robotics data solutions in early 2026, spanning images, video, 3D captures, and teleoperation datasets. Egocentric Network advertises a catalog of nearly half a million videos totaling close to 100,000 hours across 63 countries, all available for licensing. Toloka released HomER v2 in mid-2026, claiming more than 25,000 hours of off-the-shelf egocentric footage. Scale AI, the incumbent data labeling giant, quietly opened a "Physical AI" practice, though details remain sparse.
The competitive landscape is still forming. Some providers emphasize volume and geographic breadth. Others tout annotation quality or vertical specialization. A handful are building two-sided marketplaces that connect data buyers directly with facilities willing to install cameras and sensors.
Praxis claims coverage across tens of thousands of workers inside large, publicly listed companies and unicorns, spanning more than 150 environment types—a differentiator that hints at an enterprise sales model rather than a long-tail marketplace play. The founders describe their goal as building "permanent industrial infrastructure" for data production, framing the offering not as a one-time procurement but as ongoing operational instrumentation.
Whether that resonates with procurement teams, and whether the self-reported scale holds up under independent scrutiny, is less clear. The company has not disclosed customer names publicly.
The Technical Stack Matures

On the infrastructure side, progress is accelerating. Dataset formats are beginning to converge around standards that support multimodal, time-series data at scale. Hugging Face's LeRobot project, updated in 2026, defines a Parquet-based structure designed specifically for large-scale robotics training. NVIDIA's NuRec pipeline, documented over the summer, converts real-world camera and lidar captures into simulatable scenes compatible with Isaac Sim.
The tooling to ingest, process, and serve petabyte-scale video is maturing fast. NVIDIA demonstrated its NeMo Curator processing roughly one million hours of 720p video per day on a cluster of 2,000 H100 GPUs. The compute requirements are staggering, but increasingly feasible for well-funded labs and companies.
What's less settled is the regulatory picture.
The EU AI Act entered into force in mid-2024, with phased applicability stretching into 2028. High-risk AI systems—which may include humanoids deployed in certain industrial or public contexts—face obligations around dataset quality, representativeness, and bias management. The European Data Protection Board issued updated guidance in spring 2026 emphasizing lawfulness, transparency, and necessity for workplace video capture. Large-scale monitoring of employees likely triggers requirements for a data protection officer and a formal impact assessment.
In the United States, the regulatory environment is more fragmented. California's CPRA has covered employee data since early 2023, granting workers the right to know what information is collected, request corrections, and demand deletion. Illinois BIPA restricts biometric data collection and includes a private right of action that has spawned significant litigation. California legislators introduced bills during the 2025–2026 session proposing additional guardrails around "workplace surveillance tools," including restrictions in sensitive areas and obligations on vendors not to resell derivative datasets. Whether those bills become law, and in what form, will shape what's permissible in one of the country's largest industrial markets.
For data infrastructure providers, this means building consent workflows, retention policies, and cross-border transfer mechanisms from the start. It also means educating enterprise customers on their own compliance obligations, since outsourcing data collection doesn't eliminate legal liability.
Who's Training the Robots of 2028?

The opportunity is large enough to justify the complexity. If humanoid deployments scale anywhere near the projected trajectories—and if training those robots demands millions of hours of diverse, real-world human activity filmed from first-person perspectives—then the market for workplace data infrastructure could rival the market for the robots themselves. Perhaps even exceed it, at least in the early years before synthetic data generation reaches parity.
The data gold rush is on. The question is whether providers can move fast enough to meet demand, stay compliant enough to avoid regulatory entanglements, and deliver data rigorous enough to generalize across the chaotic environments where robots will actually need to perform.
Praxis and its cohort are placing that bet. The rest of us will find out soon enough.
Probably by watching whether the humanoids folding laundry in a few years learned their moves from a factory worker in Shenzhen, a warehouse employee in Ohio, or a home care assistant in Stuttgart—all captured on camera, going about their day, unaware they were training their replacements.
Or their coworkers. Depending on how optimistic you're feeling.
