Praxis Robotics, founded in 2026 and staffed by three founders with backgrounds ranging from NASA research to Andreessen Horowitz deal-making, has entered what may be the hottest niche in artificial intelligence: capturing footage of people doing ordinary jobs so robots can learn by observation.
The San Francisco company emerged from Y Combinator in August 2026 claiming access to 60,000 workers across 150 industrial and residential site types on five continents, all wearing head-mounted cameras that record egocentric video paired with 3D environmental scans. That data, synchronized with motion sensors and sometimes haptic feedback, feeds into vision-language-action models capable of learning manipulation tasks without the painstaking process of teleoperating expensive robot arms. Praxis says it can deliver "calibrated, synced, pose-tracked, QA'd, PII-scrubbed" datasets and build custom capture hardware out of Shenzhen on short timelines.
The pitch arrives at a moment when academic labs have finally proven that human demonstrations can replace costly robot telemetry. NVIDIA published research in February 2026 showing that 20,854 hours of action-labeled human video follows a predictable scaling law for robot policy transfer. University of Maryland researchers reported 92.5% average task success using roughly 30 minutes of egocentric footage per manipulation task, according to work released in June. Carnegie Mellon unveiled an end-to-end pipeline in early August that converts egocentric video into robot-executable episodes through action retargeting and visual synthesis.
Venture investors are writing checks accordingly. Generalist, a "robot brain" startup, raised $200 million on August 24, just two months after closing a $400 million round, Axios reported. Agility Robotics announced a $2.5 billion SPAC deal with a $200 million private placement in late June.
Why Robot Makers Are Filming Humans Instead of Robots
For years, robot training meant teleoperating mechanical arms through thousands of demonstrations. Agility tested its Digit humanoid in Amazon fulfillment centers using recordings captured on hardware, a process that scaled poorly and cost tens of thousands of dollars per setup, according to TechCrunch and SEC filings. The Stanford AI Index 2026 report, published in April, highlighted persistent difficulty on "act in the real world" benchmarks and noted the cost burden of robot training data.
Egocentric video inverts that model. A worker straps on a camera and performs the task naturally. The footage, layered with IMU data, wrist cameras, and glove kinematics, teaches vision-language-action models how humans grip, rotate, and place objects. The approach sidesteps the embodiment problem by capturing human intent and motion, then letting retargeting algorithms translate those actions into robot-compatible commands.
Praxis claims it operates within "publicly listed and unicorn-scale conglomerates," though it declined to name enterprise partners publicly. Founder and CEO Rohan Seelamsetty studied computation and cognition at Penn and completed an Apollo Fellowship at Oxford. COO Dev Karpe brings a Wharton Huntsman background in investment banking and supply-chain operations. CTO Tommy Li worked in tech private equity at Morgan Stanley, industrial-tech venture capital at Andreessen Horowitz, and did research stints with NASA's Jet Propulsion Laboratory and autonomous-vehicle startup PerceptIn.
Three Forces Driving the Data Race

The research tailwind is perhaps the most striking. NVIDIA's EgoScale study found that scaling human video reduces validation loss for dexterous manipulation policies in a power law, much like language models improve with more text tokens. Maryland's HumanEgo work reported zero-shot transfer from minutes of video per task, a threshold that makes commercial collection viable. Carnegie Mellon's Ego2Robot pipeline, released August 3, demonstrated that the retargeting problem isn't solved but is tractable enough for startups to build businesses around it.
Robot deployments are broadening geographically and sectorally. The Association for Advancing Automation reported 9,055 robot orders worth $543 million in North America during the first quarter, with demand spreading beyond automotive and collaborative robots representing 18.1% of units, according to industry publication Automation.com. The International Federation of Robotics' preliminary data, released June 18, showed U.S. industrial robot installations climbed 11% year-over-year to 38,000 units, Manufacturing Dive reported. Google DeepMind made a robotics foundation model available via API in late July, signaling developer-ecosystem momentum.
Regulatory pressure is mounting faster than many in the industry expected. The EU AI Act's transparency obligations took effect August 2, with high-risk AI system rules phasing in through late 2027 and mid-2028, according to the European Commission and EUR-Lex timelines. High-risk systems require governed training datasets with documented provenance. General-purpose AI providers must publish training-data summaries. That compliance burden favors vendors who deliver structured, auditable corpora over ad hoc internal collection. "Robotics is missing its internet-scale training corpus," Praxis wrote in its launch post, a claim that isn't strictly true given Meta's Ego-Exo4D dataset and others, but captures the commercial gap.
At Least Eight Vendors Are Chasing the Same Opportunity

Scale AI operates a "Physical AI" unit with a "robotless egocentric platform" and sensor-fusion job postings. MOVAS AI claims 100,000 hours of egocentric data. iMerit offers egocentric video collection and annotation for embodied AI and participated in a CVPR auto-annotation challenge this year. Grasp Labs, Datoric, Robgence, EgoVista, and ArcheBase all advertise similar services. EGXO Data publishes a "Robotics Data Release Tracker" and technical buyer guides, positioning itself as an industry analyst as much as a vendor.
The academic scaffolding came earlier. Meta FAIR's Ego-Exo4D v2 dataset, updated in April 2024, totals roughly 1,286 hours of video with multiview and multimodal annotations. EPIC-KITCHENS, a longstanding egocentric cooking dataset, saw GitHub updates as recently as late June, though its core data predates recent breakthroughs. HM3D and ScanNet++ v2 provide 3D indoor scans that complement egocentric footage with spatial grounding.
Capgemini's May report identified an "AI-robot-data flywheel" in which early deployments generate training examples that improve future policies. Andreessen Horowitz published pieces on frontier systems for the physical world in April and the physical AI deployment gap in January, both highlighting vision-language-action models and the need for large-scale robot data. ABI Research released a market analysis on robotics foundation models June 22.
Quality, Not Volume, Will Determine Winners

Synchronizing stereo RGB, IMU, pose tracking, and contact labels into robot-ready episodes requires end-to-end pipelines that segment actions, scrub personally identifiable information, and format outputs for frameworks like RLDS or LeRobot. Legal compliance is non-trivial. The European Data Protection Board's April guidance on video devices and the UK Information Commissioner's Office guidance on worker monitoring both stress lawful basis, data protection impact assessments, and transparency. Consent is often unsuitable in workplace contexts. In the U.S., state-level consent laws and employer policies create a fragmented landscape, according to legal analyses from Nolo and law firm Venable published earlier this year.
The next bottleneck may be action labels. While EgoScale, HumanEgo, and Ego2Robot demonstrate that human video scales, they also expose the "embodiment gap" between a human hand and a robot gripper. Retargeting pipelines and robot-visual synthesis are improving, but they remain research problems rather than commodity tools. Startups that can deliver not just raw footage but robot-executable action sequences, with verified success rates in target environments, will command pricing power.
Some projections are extravagant. Research & Markets forecast the physical AI market will grow from roughly $383 billion this year to $3.26 trillion by 2040, though that estimate comes from a vendor forecast and warrants skepticism. The near-term signal is clearer: industrial robot orders are steady, humanoid makers are raising nine-figure rounds or going public, foundation-model labs are shipping robotics APIs, and the EU is enforcing data-governance rules.
The startups that can turn factory floors and warehouses into training grounds at scale, legally and with robot-ready annotations, are solving the problem investors believe stands between today's 38,000 annual U.S. robot installations and tomorrow's embodied-AI economy. Praxis joined that race in August with three people and a claim to 150 environments on five continents. Whether its network delivers on that scale, or whether the founders can out-execute better-funded competitors, remains an open question.
