Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Fintech iconFintechFebruary 23, 2026

Inside the AI Agent That Can Launch a Venture Capital Fund

Inside the AI Agent That Can Launch a Venture Capital Fund
Ai AgentsVenture Capital+2
Climate / Social Tech iconClimate / Social TechFebruary 23, 2026

ewigbyte Raises €1.6M for Zero-Energy Glass Storage Archives

ewigbyte Raises €1.6M for Zero-Energy Glass Storage Archives
Data Center EfficiencyEnergy Storage+3

Founders Mentioned

Samay Maini

Human Archive

saas icon
SaaS

Jensen Huang

Nvidia

saas icon
SaaS

Samay Maini

Human Archive

saas icon
SaaS

Jensen Huang

Nvidia

saas icon
SaaS
SaaS iconSaaS
February 23, 2026
RoboticsTraining DataArtificial IntelligenceAutonomous SystemsAi Hardware

The Race to Build Training Data Infrastructure for Robot Brains

As humanoid robots approach a $5T market, startups and tech giants are scrambling to solve AI's biggest bottleneck: collecting real-world multimodal data to teach robots how to move.

The Race to Build Training Data Infrastructure for Robot Brains

Four students—two from Stanford, two from Berkeley—made what looked like either a brilliant or catastrophically stupid decision sometime in early 2024. They quit school, bought one-way tickets to Asia, and started building what they're now calling "the world's largest annotated multimodal dataset" for robots.

Their company, Human Archive, secured Y Combinator backing. The pitch? That training data, not flashier algorithms or cheaper hardware, will determine which robots actually work when they leave the lab and enter the real world. It's a bet that could look prescient or delusional depending on how the next five years unfold.

They're hardly alone in this particular gamble. Morgan Stanley projects the humanoid robotics market hitting $5 trillion by 2050—roughly 1 billion units deployed across factories, warehouses, and eventually homes. But there's a question underneath all those projections that doesn't get asked enough: where does the data come from to teach a billion robots how to move?

When Simulation Stopped Being Enough

The robotics community spent years in love with simulation. The promise was elegant: generate millions of synthetic trajectories, train your model overnight, deploy. Fast. Clean. Infinitely scalable.

Then reality intruded.

Google DeepMind's RT-X experiments told a different story. Cross-embodiment training on actual real-world data—combining 60-plus datasets across 22 robot types, more than 1 million trajectories total—improved task success rates by roughly 50 percent over single-task baselines. The Open X-Embodiment dataset, pulled from labs worldwide, now covers 527 skills captured across wildly different robotic platforms. Models trained on this messy, real-world data generalized to new environments in ways pure simulation never managed.

Or consider Covariant's RFM-1, an 8-billion-parameter multimodal model trained on production warehouse data. It delivers 1,000 pick cycles per hour with 99 percent-plus precision in some operations. That performance came from ingesting real images, videos, encoder readings, pressure sensors, operational metrics—data from actual robot fleets, doing actual work. No simulation could replicate that fidelity. Not yet, anyway.

"We're seeing a shift across robotics training—from simulation to real-world data," wrote Samay Maini, one of Human Archive's co-founders, in a November post. His company deploys custom hardware to capture egocentric perspectives, force feedback, and teleoperation sequences across residential and manufacturing environments. Whether that's defensible as a business model is another question entirely.

RGB Video Is the Beginning, Not the End

RGB video alone doesn't cut it anymore.

The latest datasets reflect an explosion in sensor modalities: depth maps, tactile feedback, proprioceptive signals, audio streams. Some researchers are even capturing WiFi channel state information and physiological data for human-robot collaboration scenarios, which sounds excessive until you realize collaborative robots need to predict when a human worker is about to reach for something.

The DROID dataset, released in March 2024 by Toyota Research Institute and collaborators, includes 76,000 demonstrations spanning 350 hours, 564 scenes, 84 distinct tasks. Raw stereo HD footage sits alongside force-torque readings and precise joint encoder data. The MiDAS surgical robotics dataset, published in February 2026, captures electromagnetic hand tracking, foot pedal inputs, and surgical video—all synchronized without requiring access to proprietary telemetry systems. (That last detail matters in medical robotics, where equipment manufacturers guard data access jealously.)

For human-robot collaboration in factories, researchers now collect EEG, ECG, electrodermal activity, respiration, and EMG signals to model worker state and intention. Assembly datasets incorporate event cameras, force-torque sensors, microphones to capture contact-rich manipulation. The technical challenge isn't just recording multiple streams. It's synchronizing them with millisecond precision across hardware that was never designed to talk to each other.

Human Archive's approach centers on custom-built capture rigs that handle this complexity. Their annotation pipelines process egocentric clips with instance segmentation, hand tracking, depth reconstruction. Other startups are taking different angles: Embodi develops proprietary wearables to capture motion, force, and egocentric perspectives, marketing itself as an "Apple Store for robotic skills"—a phrase that probably tested well with investors. Nexdata opened what it calls an "Embodied AI Data Factory" in January 2026: 4,000 square meters of configurable environments housing over 100 humanoid robots for large-scale multimodal capture.

Whether any of these approaches actually scale profitably remains an open question.

The Format Wars (and the Surprising Détente)

Digital illustration for article section "The Format Wars (and the Surprising Détente)" in "The Race to Build Training Data Infrastructure for Robot Brains" - A conceptual illustration depicting the transition from digital fragmentation to standardization, vi...

As datasets proliferated, format fragmentation threatened to balkanize the field. Then RLDS—Reinforcement Learning Dataset Standard—emerged as something close to a de facto standard, adopted across the Open X-Embodiment project and Toyota's DROID dataset. Hugging Face extended the ecosystem with LeRobotDataset v3, adding streaming capabilities and multi-episode chunking to handle terabyte-scale collections.

The BEHAVIOR 2025 challenge released 10,000 teleoperation demonstrations—1,200-plus hours—directly in LeRobot format on Hugging Face. Stanford's massive datasets now port cleanly into the same infrastructure.

This standardization matters. It allows researchers to combine datasets from different sources, training models on hundreds of thousands of trajectories without rebuilding preprocessing pipelines for each new data source. But standardization also creates competitive pressure. If the tooling becomes commoditized, data providers compete on capture scale, modality richness, and environment diversity.

That's where Human Archive's move to Asia makes strategic sense, at least on paper. Lower costs for distributed data collection. Access to manufacturing environments at scale. The ability to capture daily-life scenarios across different cultural contexts. Whether it actually works depends on execution quality, which is where most startups die.

NVIDIA's Synthetic Counterargument

NVIDIA's Isaac GR00T platform represents the counterargument to pure real-world capture: what if you could bootstrap massive synthetic datasets from minimal real seeds?

The GR00T N1 open humanoid foundation model, announced at CES 2025, includes blueprints that generate hundreds of thousands of synthetic trajectories in hours. Subsequent releases (N1.5, N1.6) integrated world models that improve physical realism. Jensen Huang, NVIDIA's CEO, declared the "age of generalist robotics" at CES 2025, though his company's approach explicitly combines open models with partners who provide domain-specific real-world data. That last detail is easy to miss in the keynote, but it's critical.

The research suggests synthetic and real data will coexist, not compete. A February 2026 study called EgoScale demonstrated log-linear performance improvements using 20,854 hours of egocentric human video, with strong transfer to a 22-degree-of-freedom dexterous hand using minimal robot supervision. Internet-scale video pretraining appears powerful for high-level planning. But low-level motor control still demands real sensor streams—the texture of a surface, the give in a material, the friction coefficient of an unexpected object.

Companies are finding the sweet spot in hybrid approaches. Use simulation to generate diverse scenarios rapidly, then validate and fine-tune with targeted real-world captures that address specific failure modes. NVIDIA's partnerships with Figure, 1X, Boston Dynamics, Agility, and others all involve this real-plus-synthetic recipe.

The question becomes: who provides the real data layer? And can they charge enough to build a viable business?

The Compliance Minefield Nobody Talks About Enough

Here's the problem nobody wants to discuss in funding announcements: collecting real-world robot training data means recording people. In their homes. At work. On factory floors.

The legal and ethical complexity is substantial, particularly in Europe.

The EU's General Data Protection Regulation requires explicit consent, purpose limitation, and data minimization for any video that might capture identifiable individuals. The European Data Protection Board's Guidelines 3/2019 specifically address video devices, mandating multilayered notice and strict handling of biometric data. In California, two-party consent applies to audio recordings. Workplace surveillance must be disclosed. CCPA and CPRA create additional obligations.

Starting January 2027, the EU Machinery Regulation expands scope to AI-enabled machinery, meaning software updates to robot controllers could trigger reassessment under safety rules that now explicitly account for AI and cybersecurity. ISO standards for industrial robots (ISO 10218-1, updated February 2025) and collaborative robots (ISO/TS 15066) set force and pressure thresholds for human contact.

Data providers building compliant pipelines need consent frameworks, real-time anonymization—blurring faces, muting identifying audio, removing badge numbers—data protection impact assessments, and region-specific retention policies. Companies like LeData explicitly position around "ethical and compliant" datasets with GDPR and EU AI Act framing. This compliance burden creates moats for established players who can afford legal infrastructure. It also raises barriers for smaller operations, which may explain why Human Archive went to Y Combinator first.

The Market (If It Actually Materializes)

Digital illustration for article section "The Market (If It Actually Materializes)" in "The Race to Build Training Data Infrastructure for Robot Brains" - A conceptual illustration depicting the explosive financial growth and industrial integration of hum...

Figure AI raised $675 million in February 2024 from Microsoft, OpenAI, NVIDIA, and Jeff Bezos at a $2.6 billion valuation. The company signed a partnership with BMW to deploy humanoids in factories. 1X, backed by OpenAI, raised $100 million in January 2024 and announced its consumer NEO bot with U.S. preorders opening in late 2025.

These hardware companies all need data. Either they're building internal data engines or they're buying from third parties. Or, more likely, some combination of both.

Goldman Sachs projects the humanoid robot market at $38 billion by 2035 with 1.4 million units shipped annually. Morgan Stanley's more aggressive forecast sees $5 trillion by 2050 with roughly 1 billion units deployed, 90 percent industrial and commercial. UBS estimates 300 million humanoid robots by 2050 with a total addressable market between $1.4 and $1.7 trillion. McKinsey's base case for general-purpose robotics reaches approximately $370 billion by 2040, with China capturing half that market.

Those projections vary wildly, which should tell you something about how certain anyone really is. But even the conservative estimates imply massive data infrastructure demand. Every deployed robot needs to learn dozens or hundreds of tasks. Every new environment requires fine-tuning. Every software update benefits from additional training data. The International Federation of Robotics reported roughly 200,000 professional service robots sold in 2024, up 9 percent year-over-year, with robotics-as-a-service growing 31 to 42 percent annually.

If the market scales as projected—and that's a meaningful "if"—training data infrastructure becomes critical path.

What Separates Winners from the Also-Rans

The competition will be fierce. Startups like Human Archive, LeData, Nexdata, and Embodi compete against tech giants building proprietary data engines: DeepMind's RT-X hub, Toyota Research Institute's DROID operation, Covariant's production fleet data flywheel. Platforms like Viam and Roboto.ai offer full-stack infrastructure that includes data capture, labeling, and training. General AI labeling vendors—Appen, Labelbox, Uber's Scaled Solutions—are extending into multimodal robotics workflows.

Several factors will likely determine who wins.

Capture scale and diversity. Human Archive's distributed collection workforce across Asia offers different scenarios than Nexdata's controlled factory setup or LeData's household manipulation focus. Models trained on cross-embodiment datasets generalize better. Breadth matters, perhaps more than anyone initially expected.

Modality richness. The field is moving beyond RGB-D. Winning datasets will include synchronized force, tactile, audio, proprioception, and potentially physiological signals where relevant. Hardware becomes a differentiator—custom rigs that can capture these streams cleanly without massive post-processing.

Compliance and governance. Companies that build trust through transparent consent frameworks, robust anonymization, and regional legal compliance will win contracts from risk-averse enterprises. The EU AI Act's Code of Practice and incoming machinery regulations make this increasingly important. It's boring infrastructure work, but it's defensible.

Integration with dominant toolchains. Providing data in RLDS and LeRobot formats, with clear documentation and loader code, reduces friction for customers. Compatibility with open policies like Octo and OpenVLA extends market reach. This sounds obvious, but integration work is where many data startups stumble.

The ability to bridge synthetic and real. Data providers who can validate or seed NVIDIA's GR00T blueprints with targeted real captures—closing sim-to-real gaps—offer unique value. Pure synthetic generation will never fully replace real data. But pure real-world capture is too slow and expensive to meet demand. The hybrid model is where sustainable businesses get built, assuming they get built at all.

The Next Eighteen Months Will Tell Us Everything

Digital illustration for article section "The Next Eighteen Months Will Tell Us Everything" in "The Race to Build Training Data Infrastructure for Robot Brains" - A dramatic, high-contrast illustration depicting a sleek humanoid robot engaged in industrial labor,...

The humanoid robotics race is entering its industrial deployment phase, or at least that's the story being told in pitch decks. BMW already has Figure robots in limited production. Agility's Digit is being tested in logistics warehouses. These early deployments will generate a wave of demand for task-specific fine-tuning data—picking up this particular part, navigating that specific layout, handling edge cases discovered only in production.

Simultaneously, the open-source ecosystem is maturing. The X-Embodiment community, which ran a workshop at CoRL 2024, continues pushing for shared datasets and responsible research practices. As more labs contribute to Open X-Embodiment and similar aggregations, the baseline capabilities of generalist policies will rise. So will the bar for what counts as novel or valuable data.

Pieter Abbeel, co-founder of Covariant, told WIRED in March 2024 that "foundation models are the future of robotics," but emphasized they require multimodal, action-grounded data and production data flywheels. That last part—the flywheels—is what separates companies with actual deployment pipelines from those with impressive demos.

The technical bottleneck is becoming clear. Algorithms are advancing rapidly: vision-language-action models, world models for physics reasoning, zero-shot transfer techniques. Hardware costs are dropping faster than Goldman Sachs initially projected. But none of that matters if robots can't learn the long tail of real-world tasks. Folding irregular laundry. Handling transparent or reflective objects. Collaborating with humans who don't follow scripts.

That learning requires data captured in the messy complexity of actual homes, factories, hospitals, and warehouses. It requires force feedback from real objects with unpredictable friction and weight distribution. It requires synchronized multimodal streams that capture how tasks actually get done, not how they ought to get done in simulation.

Four students saw this coming—or thought they did—and dropped out of Stanford and Berkeley to move to Asia and build data infrastructure. Teams at a dozen other startups made similar bets. So did the infrastructure divisions at tech giants who'd prefer not to depend on outside vendors for something this critical.

The race is on. But it's not, ultimately, a race to build better robots. It's a race to build the data infrastructure that makes better robots possible. Whether that infrastructure becomes a winner-take-most market or fragments into regional niches and vertical specialists—that's still anyone's guess.

What's certain is that someone needs to collect millions of hours of multimodal, task-specific, legally compliant robot training data. And they need to do it before the humanoid boom either arrives or collapses under its own weight.

Human Archive and its competitors are betting everything that they'll be the ones to do it. We'll know soon enough if they were right.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Inside the AI Agent That Can Launch a Venture Capital Fund
  • ewigbyte Raises €1.6M for Zero-Energy Glass Storage Archives
  • Mining Millions of Years of Evolution: How Parasites Could Solve Autoimmunity
  • AI Co-Scientists: How Autonomous Research Agents Are Reshaping Discovery
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.