The four of them—none older than 22—spent the better part of two months shuttling between India and China, lugging custom camera rigs and force sensors into residential kitchens and factory floors. Their mission was unglamorous, even tedious: capture hours upon hours of human hands doing ordinary things.
While much of the robotics world remains fixated on the sexier questions—transformer architectures, parameter counts, the next clever training trick—Rushil Agarwal, Samay Maini, Raj Patel, and Shloke Patel are betting on a different thesis entirely. The real bottleneck in embodied AI, they argue, isn't compute or cleverness. It's data. Specifically, the kind of high-quality, multimodal data captured in actual homes and workplaces where robots eventually need to function.
The timing feels deliberate, even if the execution remains scrappy. Human Archive, the startup these Berkeley and Stanford dropouts founded, is part of Y Combinator's Winter 2026 batch—a cohort heading toward its Demo Day showcase on March 24. They're entering an industry that has finally, grudgingly acknowledged what researchers at Google DeepMind, Covariant, and NVIDIA have been saying for years: you simply cannot train capable generalist robots without massive amounts of real-world data, aligned across multiple modalities.
Whether that insight translates into a sustainable business is another question.
A Fragmented Ecosystem
The robotics data landscape today resembles what computer vision looked like around 2010—lots of well-intentioned academic efforts, nothing close to ImageNet scale or coherence.
Academic consortia have made notable progress. DROID assembled 76,000 demonstration trajectories across three continents. OXE-AugE expanded to 4.4 million trajectories late last year, reportedly improving cross-embodiment generalization by roughly 24% to 45% for models like OpenVLA. Just last month, RoboMIND 2.0 released 310,000 dual-arm manipulation trajectories, including some 12,000 episodes enhanced with tactile sensing.
These datasets matter. But they're patchwork solutions, collected opportunistically across lab settings with limited embodiment diversity. The data modalities often aren't synchronized—egocentric video might exist without corresponding force feedback, or proprioception logs might miss tactile sensing entirely. For vision-language-action models trying to learn dexterous manipulation, these gaps aren't minor annoyances. They're fundamental limitations.
Commercial players have begun addressing the problem, though mostly to serve their own needs. Covariant built RFM-1 on what it describes as "the largest real-world robot dataset," derived from its warehouse deployment fleets. Figure's humanoid robots accumulated 1,250 operating hours during a pilot at BMW's Spartanburg facility—loading over 90,000 parts across nearly a year of weekday shifts, walking roughly 200 miles in the process.
But this data remains proprietary. And most companies building humanoid robots or foundation models lack comparable deployment scale.
Enter the data collection startups.
Human Archive positions itself as a "multimodal data provider for robotics learning," deploying custom hardware to capture egocentric footage, instance segmentation, hand tracking, depth maps, force feedback, and teleoperation data. According to founder posts from recent months, the team shifted operations to Asia specifically to build what they're calling "the world's largest annotated multimodal dataset" through a distributed labor force.
They're hardly alone. Embodi Labs is assembling what it terms a "Global Archive of Human Skill" through an operator guild network spanning more than 50 countries. TARS Robotics announced its WIYH dataset last October, billing it as the "world's first large-scale real-world Vision-Language-Tactile-Action" collection, backed by funding rounds totaling $242 million. RealMan launched RealSource in December from the Beijing Humanoid Robot Data Training Center, advertising "100% modality completeness" with sub-0.5% frame loss.
Whether these ambitious claims survive contact with reality remains to be seen.
Three Converging Forces

The urgency around real-world data collection stems from three trends that have converged with surprising speed.
First, the architecture debate has largely settled. Vision-language-action models appear to be the path forward. RT-2 demonstrated back in 2023 that web-scale VLM pretraining combined with robot demonstration data could enable semantic understanding for manipulation. OpenVLA, trained on roughly 970,000 robot trajectories from Open-X, became a widely-used open baseline throughout 2025. NVIDIA open-sourced its GR00T N1 humanoid foundation model at GTC 2025, framing it as the dawn of an "age of generalist robotics." Dual-system architectures—high-level reasoning paired with low-level continuous control—are now standard across Figure's Helix, 1X's world models, and academic workshops.
That architectural consensus shifts the bottleneck squarely to data. As Covariant's team noted when launching RFM-1, the only viable path to capable robotics foundation models is "robots deployed in the world collecting a ton of data." NVIDIA CEO Jensen Huang echoed the sentiment at GTC 2025, emphasizing that solving robotics' data problem matters as much as model architecture or scaling laws.
Second, enterprise deployments are accelerating. BMW ran Figure 02 humanoids for 10-hour weekday shifts over 11 months, contributing to production of more than 30,000 X3 vehicles. Agility Robotics overhauled Digit's navigation stack for warehouse deployments with GXO and Amazon; some reports pegged operating costs at around $30 per hour. Apptronik has raised substantial capital across multiple funding rounds through early this year, with pilots underway at Mercedes and a manufacturing partnership with Jabil.
These deployments generate valuable proprietary data. They also reveal capability gaps—particularly for dexterous bimanual tasks, long-horizon sequences, and generalization across environments. Companies training vision-language-action models need complementary datasets that capture residential contexts, manufacturing variations across geographies, and diverse object-human interactions that proprietary fleet data simply cannot cover.
Third, the modalities themselves are evolving beyond simple RGB video. Late last year saw a wave of datasets emphasizing tactile sensing and force feedback—OpenTouch's 5.1 hours of synchronized video-touch-pose data, Hoi!'s 3,048 sequences across four embodiments including custom grippers with tactile sensing. Low-cost sensors like the DIGIT (around $350) and GelSight Mini (roughly $510) are mainstreaming tactile data collection. Vision-tactile fusion challenges dominated Embodied AI workshops at CVPR 2025.
The consensus is increasingly clear: manipulation learning requires touch, not just vision.
Human Archive's pitch—egocentric capture with force feedback and teleoperation data across distributed operators—maps directly to these requirements. By establishing collection infrastructure in Asia, the founders are making a bet on cost arbitrage, environment diversity, and access to manufacturing contexts that U.S.-based academic labs struggle to reach at scale.
Echoes From Adjacent Domains
The strategy mirrors patterns from adjacent fields. Ego4D, the computer vision community's egocentric benchmark, assembled over 3,670 hours of first-person video from 900+ participants across nine countries. That geographic and demographic diversity proved critical for model generalization.
Human Archive appears to be applying the same logic to robotics—capture across regions, settings, and tasks that no single lab or company can access alone. Whether they can execute on that vision with a team of four and limited capital is, well, an open question.
There's also the regulatory dimension, which may be shaping the landscape more than founders initially anticipated. The EU AI Act's high-risk system obligations take effect this August, requiring data governance documentation, transparency notices, and audit trails for robotics deployments in workplaces. California's workplace surveillance regulations, in effect since January 2025, mandate 14-day advance notice and access rights for employee monitoring. Illinois BIPA requires written consent for biometric capture, including hands and gait patterns.
For data collection startups operating globally, compliance becomes both operational overhead and competitive moat. Human Archive's emphasis on "anonymization + QA + annotation pipelines" suggests at least some awareness of these requirements. Companies that can furnish data processing agreements, evidence of pseudonymization, and AI Act-ready lineage documentation in the coming months will have an advantage selling into Europe-exposed customers.
China's embodied AI push adds another dimension to consider. The country accounted for 54% of global robot installations in 2024 and roughly two-thirds of robotics patents, according to analysis published last year. The World Robotics Conference showcased over 100 humanoid models. RealSource, released in December and available on Hugging Face, signals a strategy of open dataset releases that could accelerate global availability while potentially undercutting Western data providers on cost.
What Comes Next

The near-term market will likely stratify into distinct tiers.
Proprietary fleet data from companies like Covariant, Figure, and Tesla will remain highest-value for their specific embodiments and tasks, but won't transfer well to other platforms. Open academic datasets will continue serving as research benchmarks but lack the scale and modality completeness that production models require.
Commercial data providers occupy the middle ground—offering environment diversity and multimodal alignment that single-company fleets cannot capture, but needing to compete on quality, compliance, and cost.
Human Archive's bet on Asia as a collection base makes sense if labor arbitrage holds and regulatory complexity doesn't multiply faster than dataset value can grow. The team has apparently moved quickly since leaving Berkeley and Stanford, hiring operations managers and videographers in India according to job postings visible on their Y Combinator profile. Industry databases show a convertible note of around $500,000 from roughly two months ago, though that figure hasn't been independently verified ahead of YC's upcoming Demo Day.
For founders building embodied AI systems, one trend seems clear: modality completeness will matter more than raw trajectory counts. Models that pair egocentric video with synchronized force, tactile, and language annotations significantly outperform vision-only approaches, as recent results have demonstrated. Dataset providers that can deliver aligned multimodal streams, documented consent and governance protocols, and cross-embodiment coverage will likely command premium pricing.
For investors, the central question is whether data collection becomes a sustainable business or simply a feature that foundation model companies eventually subsume. Covariant built its data moat through deployments; Figure is accumulating operational data at BMW. But neither can easily capture residential contexts or long-tail manufacturing variations across dozens of countries. If vision-language-action generalization continues to scale with embodiment diversity—as recent scaling law research suggests—then independent data providers filling geographic and task gaps might maintain defensibility.
At least in theory.
The Long Game

The robotics data problem is getting solved, one egocentric frame at a time. Whether Human Archive becomes the Ego4D of embodied AI or gets eclipsed by better-capitalized competitors will depend on execution over the next eighteen months—and perhaps more than the founders expected, on their ability to navigate regulatory complexity across multiple jurisdictions.
For now, four dropouts are betting that living overseas, deploying custom hardware, and capturing the messy details of human manipulation at scale is the less obvious—but possibly faster—path to training the robots we keep promising are almost here.
The thing about less obvious paths, though. Sometimes they're less obvious for a reason.
