Three people in a Y Combinator batch think they've figured out how to feed the machines.
The problem, it turns out, isn't that robots are getting dumber. It's that they're hungry in ways the tech industry didn't quite anticipate. While ChatGPT and its cousins gorged themselves on the internet's bottomless text buffet, physical AI systems—the kind that need to fold laundry, pack boxes, and navigate real warehouses—require something messier: egocentric video, 3D scans, sensor data captured in actual kitchens and factories, not labs.
The bottleneck is real enough that Goldman Sachs projected a $38 billion humanoid robot market by 2035. Robotics startups pulled in a record $40.7 billion in 2025. But talk to anyone building these systems and they'll tell you the constraint has quietly shifted. It's not the hardware anymore. It's the training data—the right kind, from the right places, at the right scale.
Which brings us to Praxis AI.
The startup—founded by Rohan Seelamsetty, Dev Karpe, and Tommy Li—is three people deep and barely out of Y Combinator's Summer 2026 batch. Their pitch, according to the accelerator's public profile: turn every company into a data vendor. Capture egocentric video and multimodal sensor feeds from inside industrial and residential environments, then pipe it to the humanoid companies and foundation model labs that need it. The team's YC profile claims access to 60,000 workers across five continents and 150-plus environment types. That's an audacious number for a startup with no disclosed funding beyond YC's standard check. But perhaps it speaks to the urgency of the gap they're trying to fill.
Or perhaps they're very good at pitching.
The Data Deficit the Industry Didn't See Coming
Google DeepMind laid it out plainly enough in January 2024 with AutoRT: robots can't learn from Wikipedia scrapes. They need what the researchers called "experiential data"—sensor logs, video, trajectory recordings collected at scale across different robot bodies and real-world settings. The problem was apparent to anyone paying attention. DeepMind's RT-1, RT-2, and RT-X work emphasized multi-robot data aggregation. The Open X-Embodiment consortium pooled over a million robot trajectories across 22 robot types by 2024, a foundational collection even now, in 2026. DROID added 76,000 manipulation trajectories in March 2024—350 hours across 564 scenes.
Progress, yes. Enough? Not remotely.
Industrial deployment demands robots that can handle long-horizon tasks, generalize across object categories, operate in cluttered human spaces. That requires orders of magnitude more data. And collecting it via robot teleoperation is expensive, slow, doesn't scale to the messy diversity of real conditions.
The industry needed a different approach.
Why First-Person Video Became the New Gold Rush
The robotics community found a workaround in an unlikely place: human egocentric video. Meta's Ego4D program, which has evolved through v2.1 with updates stretching into 2025 and 2026, provides massive volumes of first-person footage. Ego-Exo4D v2, presented at CVPR 2024, pushed further—1,300 hours of video, 221 hours from egocentric viewpoints, synchronized with external camera angles.
The logic is almost obvious in hindsight. Humans already perform the manipulation tasks robots need to learn. Egocentric video captures the world from a robot-like perspective. Train on that, maybe you bootstrap your way to competence.
Papers throughout 2025 and into 2026 refined the idea. EgoLive, published in April, focused on extracting robot lessons from first-person footage. EgoTraj tackled trajectory prediction in May. EgoPhys targeted physics understanding of deformable objects in June, all from egocentric video. Meanwhile, 3D Gaussian Splatting variants emerged to reconstruct dynamic scenes from monocular footage—a technical leap documented across multiple April 2026 arXiv papers that enabled faster world-model generation.
The hardware to capture this data is proliferating, too. EssilorLuxottica reported selling over seven million Ray-Ban Meta smart glasses in 2025, up from roughly two million in the prior two years combined. That's millions of potential egocentric data streams from everyday human activity, just walking around. Compare that to Apple's Vision Pro: an estimated 390,000 units shipped in 2024 according to IDC, with weak 2025 momentum. The consumer appetite for lightweight, first-person capture devices is quietly reshaping the availability of training data, whether anyone planned for it or not.
The Infrastructure Layer Consolidates

Scale AI moved early. The company launched its "Data Engine for Physical AI" in September 2025, then announced a partnership with Universal Robots in March 2026 to produce production-grade robotics datasets. Scale's documentation reflects mature sensor-fusion pipelines—LiDAR, RGB-D cameras, IMUs—all orchestrated into labeled, model-ready formats. The company's pivot toward physical AI coincided with a strategic investment from Meta in June 2025. A signal, perhaps, of where the data infrastructure money is flowing.
Other players are staking claims across different parts of the stack. Labelbox markets a robotics vertical that has processed over one petabyte of data as of January 2026. Encord launched LiDAR and sensor-fusion labeling tools in June 2025. Kognic and Segments.ai offer competing multi-sensor annotation platforms. Appen now advertises "Physical AI training data" services, though with less public detail than Scale's offerings.
Praxis AI's entry point—egocentric video and 3D scans "from inside real industrial and residential environments"—suggests a different wedge. Rather than labeling infrastructure, the company appears focused on primary capture: getting cameras and sensors into the environments where robots will eventually operate, then piping that raw feed to customers.
The specifics remain thin. Praxis's landing page is minimal. As of early July 2026, no independent case studies or customer logos have surfaced in public reporting. The YC profile is the primary source. Which means we're left with the pitch, not the proof.
Real-World Collection Goes Big—and Expensive

The serious money is chasing scaled real-world data collection, not synthetic substitutes.
Figure AI announced "Project Go-Big" in September 2025—a partnership with Brookfield to access over 100,000 residential units for large-scale humanoid pretraining datasets. The ambition is clear: capture heterogeneous in-home environments at a distribution no synthetic or public dataset can match. Figure's bet is that pretraining on the "real distribution" of residential tasks will unlock generalization that lab-collected data cannot. Whether that bet pays off is another question.
NVIDIA's GR00T foundation model series, announced in phases from March 2025 through June 2026, takes a hybrid approach. GR00T N1 was framed as an open, customizable humanoid model. By January 2026, GR00T N1.6 arrived with broader "physical AI" stack updates. The company's GR00T-Dreams blueprint, revealed this year, uses synthetic motion generation to augment limited real-world data—claiming a 36-hour iteration cycle versus three months for hand-collected skills. Unitree's reference humanoid, set to ship late 2026, will carry GR00T's reference design, signaling how quickly data-to-deployment pipelines are compressing.
Google DeepMind's AutoRT, outlined back in January 2024, demonstrated another path: deploy robot fleets equipped with VLM and LLM reasoning to autonomously gather long-tail, in-situ data while maintaining safety gating. The approach trades human teleoperation for machine-orchestrated exploration. A strategy designed to scale data collection without scaling human hours linearly—though how well it works in practice remains to be seen.
The Licensing and Compliance Tangle Nobody Wants to Talk About
Aggregating training data at scale introduces messy licensing constraints. Open X-Embodiment's million-plus trajectories span 22 embodiments and multiple datasets with varied licenses—Creative Commons BY, BY-SA, non-commercial subsets. Mixing these for commercial model training requires careful tracking, the kind of due diligence that slows things down. DROID's 76,000 trajectories carry a CC-BY 4.0 license, relatively permissive. But Ego4D and EPIC-KITCHENS datasets include non-commercial restrictions on subsets, complicating their use in production foundation models.
Regulatory friction is rising too, though it's arriving more slowly than the technology. The EU AI Act entered phased application starting February 2025, with high-risk AI provisions originally slated for August 2026. Political agreements through 2025 and 2026 revised the timeline; current guidance points to December 2027 for standalone Annex III systems and August 2028 for embedded high-risk AI. Companies deploying humanoids in European workplaces or education settings need to track whether their use cases land in regulated categories. Most are hoping they don't.
In the US, audio and biometric capture rules vary by state—a patchwork that makes compliance a headache. Federal baseline is one-party consent under 18 U.S.C. § 2511, but multiple states require all-party consent for recordings. Workplace audio often triggers notice requirements. Illinois's Biometric Information Privacy Act (BIPA) is particularly sharp: it governs face geometry, voiceprints, hand scans with a private right of action, and 2026 litigation has kept compliance salient. Any data collection pipeline involving egocentric video in sensitive environments must navigate this thicket.
Not exactly the kind of problem that gets venture capitalists excited, but the kind that can kill a business model quietly.
The Market Timing Question Hangs Over Everything

Goldman Sachs's February 2024 projection of a $38 billion humanoid market by 2035 assumed 250,000 unit shipments by 2030, concentrated in industrial settings. UBS and Morgan Stanley have floated larger, longer-dated forecasts—millions by 2035, hundreds of millions by 2050—but all acknowledge hardware costs, safety validation, and data availability as gating factors. CB Insights counted over 70 companies in its January 2026 "physical AI models" market map, drawn from the $40.7 billion raised across robotics in 2025.
The timeline matters for data infrastructure startups.
If humanoids scale slowly—constrained by economics or regulation—demand for training data grows incrementally, favoring entrenched players like Scale with established customer relationships and capital reserves. If the market accelerates—driven by breakthroughs in foundation models or step-changes in hardware cost curves—then early-stage specialists like Praxis could ride a wave of demand from labs hungry for proprietary, in-domain datasets. It's a timing bet as much as a technology bet.
McKinsey interviews published in April and June 2026 with robotics executives emphasized that multi-sensor data and "physical AI" are prerequisites for real deployment, but also noted long adoption ramps and uncertain value capture. The industry expects generalist models trained on hybrid corpora—real robot data, egocentric human video, synthetic scenes—to emerge between 2026 and 2028. NVIDIA's roadmap and partner announcements point to meaningful improvements in long-horizon task performance and policy transfer within that window. Which means the window for infrastructure plays is narrow.
What Comes Next—If Anyone Really Knows
The race is on to solve physical AI's data problem before the hardware curve catches up. Whether egocentric video from millions of smart glasses becomes a viable training substrate, or whether companies must revert to expensive in-situ robot teleoperation at scale, remains genuinely uncertain. Praxis AI's positioning—turning workforces into distributed data-capture networks—is one answer. Figure's residential-unit partnerships represent another. NVIDIA's synthetic augmentation via GR00T-Dreams offers a third path, though it's unclear how much synthetics can substitute for real-world messiness.
The bottleneck is real. The market projections are massive. And the question of who supplies the training data that powers the next generation of humanoid robots is very much up for grabs.
For now, we're left with three people in a Y Combinator batch, a bold claim about 60,000 workers, and no public customers. Whether that's the beginning of something big or just another pitch deck that won't survive contact with reality—well, that's the bet.
