The first humanoid robots to leave pilot programs behind didn't stumble over hardware failures. At BMW's sprawling Spartanburg facility, Figure's machines helped support more than 30,000 vehicles over approximately ten months in 2025. Inside GXO's cavernous warehouses, Agility Robotics' Digit—a bipedal unit with a vaguely unsettling gait—passed the 100,000-tote milestone last November. The machines work, which is perhaps more than the skeptics expected.
But they learn slowly. Too slowly. And making them smarter requires something no laboratory can manufacture at scale: data from the messy, unpredictable real world.
That gap has spawned a new category of infrastructure company—one that builds neither robots nor foundation models, but instead captures the raw fuel both desperately need. YC-backed Praxis AI, which emerged from stealth in July, says it has embedded data-capture infrastructure inside publicly listed companies and unicorn-scale conglomerates, reaching 60,000 workers across five continents and more than 150 environment types. Their pitch borders on the audacious: transform every company into a data vendor for the physical AI economy.
Whether that's visionary or wildly premature remains an open question.
A Data Problem Unlike Any Other
Robotics faces a challenge that looks nothing like the one that powered large language models. While LLMs could scrape the internet—vast, chaotic, but at least digitized—embodied AI systems need action-labeled video from factory floors, warehouses, kitchens, hospital corridors. Environments where cameras don't naturally exist. Where every setting introduces new variables, new lighting, new clutter.
The market opportunity is substantial, if speculative. Estimates for service robotics revenue by 2031 range from $86 billion to $210 billion, depending on which analyst report you trust. Humanoid installations hit 16,000 units in 2025, according to industry trackers, with over 80 percent deployed in China. Counterpoint Research projects cumulative installations will surpass 100,000 by 2027—a figure that may seem conservative or wildly optimistic depending on how quickly manufacturing costs drop and how tolerant employers prove to be.
That growth creates a corresponding hunger for training data. McKinsey observed in June that hardware, AI, and data are finally converging to push machines "out of controlled settings and into the real world." Real-world data, though, doesn't arrive neatly packaged. It requires embedded capture systems, consent frameworks, multimodal sensors. And the operational choreography to gather terabytes from environments that were never designed to produce datasets.
The robotics research community has known for years that cross-embodiment learning could unlock generalization—training on one robot type to improve another's performance. Google DeepMind's RT-X work in 2023 demonstrated precisely that. But by early 2026, the focus shifted decisively toward something more immediately scalable: human egocentric video. Footage captured from a first-person perspective, mirroring what an onboard robot camera might see.
NVIDIA's EgoScale project, released in April, trained a vision-language-action model on 20,854 hours of action-labeled human video and reported something rare in machine learning—a log-linear scaling law. More human video translated predictably into lower validation loss and better real-robot task performance. In the words of NVIDIA's Jim Fan, humans became "the most scalable embodiment." Which is a remarkable claim, when you think about it.
Why First-Person Video Matters

Egocentric video offers advantages that third-person footage simply can't replicate. The viewpoint matches robot sensors more closely. Hand-object interactions, occlusions, motion parallax—all captured naturally rather than reconstructed after the fact. Earlier datasets like Ego4D (released in 2021 with 3,670 hours) and Ego-Exo4D (December 2023, 1,286 hours of dual-view video) laid groundwork, but 2026 saw an explosion of new corpora and capture toolchains.
EgoVerse appeared in April. EgoLive followed on April 26. Open-AoE launched a smartphone-based capture pipeline on July 15. MobileEgo Anywhere, released May 7, enabled low-cost collection by hundreds of contributors—a democratization play that raised as many questions as it answered about data quality and consent.
The research community's enthusiasm is backed by measurable results, at least in controlled experiments. EgoScale's scaling law suggests that adding more diverse, action-labeled human video yields predictable gains. Combined with world models and diffusion policies—methods that simulate possible futures or interpolate between demonstrated behaviors—human video is becoming a cornerstone of foundation models for robotics.
Yet the challenge isn't scientific so much as logistical and commercial. Academic datasets top out in the thousands of hours. Training competitive foundation models may require orders of magnitude more. That's where the vendor ecosystem comes in, and where things get interesting.
The Data Vendors
Praxis AI's model—embedding capture infrastructure inside operating companies—represents one end of a spectrum. The YC profile describes a three-person team led by Rohan Seelamsetty (Computation and Cognition at Penn, Apollo Fellow at Oxford) alongside Dev Karpe (Wharton Huntsman, Fortune 500 supply-chain operations) and Tommy Li (a16z industrial tech VC, IEEE researcher at NASA JPL). The stated reach to 60,000 workers suggests partnerships with large employers, though specifics on funding, revenue, or named partners remain undisclosed beyond the YC listing.
They aren't alone in this space. Far from it.
NLABEL markets itself as holding the "largest egocentric video dataset" and promises "explicit, transferable rights" for training and deployment—a pitch aimed squarely at buyers wary of licensing ambiguity. 1to1.bot offers fleet-based capture and annotation, delivering labeled video to cloud storage. Praxo Labs emphasizes teleoperation infrastructure and real robotic data. EgoData advertises RGB capture "you can trust—at scale," though trust in this context remains somewhat undefined. Tactum Labs, which surfaced in March, provides custom in-the-wild human demonstrations with 3D scanning. MoVo focuses on egocentric kitchen scenes for humanoid foundation models, operating out of Miami and Bogotá. HumanoidLayer positions itself as a marketplace connecting robotics teams to data providers and annotators.
Scale AI, a public company with deep ties to defense and enterprise AI, launched a "Data Engine for Physical AI" in September 2025, targeting robotics foundation models with sensor-fusion expertise. Lightwheel announced a partnership with Manus in July to combine data infrastructure with high-fidelity glove capture for manipulation tasks—the kind of granular hand movement that makes pick-and-place operations possible.
The diversity of approaches reflects a market still sorting itself out. Some vendors emphasize volume and cost. Others pitch rights clarity or vertical specialization. Roland Berger noted in an April report that "data will be the decisive frontier in humanoid robotics," predicting that operator partnerships—agreements with companies that deploy robots or employ workers in target environments—will be central to data acquisition.
Praxis AI's go-to-market, if the claims hold, is a bet on that thesis taken to its logical extreme. Turn employers into data sources. Instrument the workplace. Harvest what workers already do.
The Legal Minefield

Capturing egocentric video inside workplaces introduces legal complexity that academic datasets could comfortably sidestep. In the EU, the AI Act's transparency provisions took effect in August 2026, with high-risk rules—including those covering certain employment and biometric systems—set to apply from December 2, 2027. GDPR imposes strict requirements on workplace monitoring: a lawful basis (often legitimate interests or contract), purpose limitation, transparency, and data protection impact assessments for high-risk processing. If the capture pipeline extracts biometric identifiers—and many modern systems do—special category protections kick in.
In the U.S., the patchwork is even messier. Roughly a dozen states require two-party consent for audio recording, including California, Illinois, and Washington—states where many robotics companies happen to operate. Illinois' Biometric Information Privacy Act (BIPA) has spawned heavy litigation over face geometry and other biometric scans, raising the stakes considerably for any egocentric system that processes facial templates. A May study from Vanderbilt Law found that "bossware" workplace monitoring tools widely share worker data with third parties, often without clear consent—a cautionary tale for continuous capture programs.
The Guardian reported in June on egocentric data contractors (Scale AI among them) running worker-worn camera programs in India, surfacing concerns about surveillance, compensation, and secondary data use. Privacy researchers at ICML 2025 argued that egocentric video poses unique threats to both camera wearers and bystanders, motivating on-device redaction and robust subject rights workflows.
For vendors marketing "explicit, transferable rights," compliance isn't optional—it's the entire product. Buyers building foundation models need clear commercial training licenses, not just research-use permissions. They need audit trails that satisfy EU regulators and contracts that survive BIPA scrutiny. The data infrastructure race isn't only about volume. It's about provenance, consent frameworks, and the operational sophistication to navigate cross-border labor and privacy law without triggering regulatory action or worker backlash.
What Comes Next

The deployment wave continues, at least for now. Figure's F.03 arrived at BMW in June. Agility Robotics expanded its Fremont facility in July and filed for a SPAC merger two days earlier—a financing move that says something about investor appetite, if not necessarily about near-term profitability. Toyota Motor Manufacturing Canada signed a commercial agreement with Agility in February.
These production-adjacent environments—factories, logistics centers, assembly lines—are precisely where embedded data capture makes the most sense, in theory. Workers are already performing the tasks that robots need to learn, in lighting conditions and clutter levels that no lab can fully replicate. The question is whether those workers, and their employers, are willing to become data sources. And on what terms.
The business model remains speculative in key respects. Praxis AI's stated reach to 60,000 workers across 150-plus environment types is unverified by third-party reporting. The three-person team size raises questions about operational scale versus partnership leverage—are they building infrastructure, or brokering relationships? The vendor landscape is crowded with overlapping offerings and unclear differentiation. And the regulatory environment is tightening, not loosening, which could favor professionalized players with legal and compliance infrastructure over scrappy startups.
What seems certain is that egocentric human video has moved from research curiosity to strategic asset. NVIDIA demonstrated a scaling law. The robotics OEMs are deploying in earnest. The EU is writing the rules. The companies that solve the logistical, legal, and commercial complexities of capturing that data inside real enterprises—turning factories into training grounds and businesses into data vendors—may end up owning a critical layer in the physical AI stack.
Whether Praxis AI and its cohort succeed will depend less on the elegance of their capture hardware than on their ability to navigate consent, deliver rights clarity, and scale operations across continents without scandal or regulatory backlash. For now, the race is on, and the finish line is measured not in robot capabilities but in terabytes per month and gigabytes per environment.
The data bottleneck is real. Someone will solve it. The question is who—and at what cost to the workers whose movements, gestures, and daily routines become the raw material for the next generation of machines.
