Datoric, a startup so new its two founders are still in the current Y Combinator batch, claims it generated close to a million dollars in revenue over the past 30 days selling custom training datasets for voice synthesis, robotics and what the industry calls world models. The figure, disclosed on the company's Y Combinator profile, has not been independently verified, but if accurate it suggests an unexpectedly fast market for a business founded in 2025 that operates with no employees beyond CEO Nikhil Reddy and CTO Jeffrey Lin.
Their pitch centers on a bottleneck that has grown more acute as companies race to deploy humanoid robots, autonomous vehicles and other forms of what researchers call embodied AI: the scarcity of high-quality, legally defensible training data. Physical AI systems need footage of how humans move through kitchens, how arms manipulate objects, how drivers react in construction zones. That data exists in fragments across research labs and in proprietary corporate vaults, but not in the volumes or formats required to train foundation models that work across different robot bodies or real-world environments.
Industrial robot installations hit 542,000 units worldwide in 2024, more than double the number a decade earlier, according to the International Federation of Robotics. Service robots sold roughly 200,000 units the same year, with transport and logistics models accounting for just over half. Warehouse automation alone became a $31 billion market in 2025, per Fortune Business Insights, and Interact Analysis projects mobile robots will grow at a 19 percent annual clip through 2030. Yet Jensen Huang, NVIDIA's CEO, said in March 2024 that building foundation models for general-purpose humanoid robots remains "one of the most exciting problems to solve in AI today." Goldman Sachs pegs the humanoid market at $38 billion by 2035 in a base scenario; Morgan Stanley Research sees eight million humanoids in the United States by 2040 and a $5 trillion global market within 25 years. UBS wrote in mid-2025 that what it called "brain technology" continues to lag behind the mechanical chassis.
The data gap has both technical and legal dimensions. Open datasets like Ego4D, which contains roughly 3,600 hours of first-person video, and Open-X Embodiment, which had assembled 4.4 million robot trajectories as of December 2025, demonstrate that models can transfer knowledge across different robot platforms. But companies building commercial systems say they need orders of magnitude more footage, and it has to come with clear provenance. Egocentric video of people cooking or folding laundry, teleoperation logs from warehouse robots, audio narration synchronized with physical actions—all captured under known licensing terms that won't expose a manufacturer to class-action liability down the road.
Regulatory walls are closing in. The EU AI Act's transparency and general-purpose model obligations took effect in August 2026; high-risk AI rules phase in over the next two years. Article 10 mandates that training, validation and test datasets for high-risk systems meet specific quality benchmarks. In the United States, the FCC ruled in February 2024 that AI-generated voice calls violate the Telephone Consumer Protection Act. Illinois saw a wave of biometric-privacy lawsuits in May 2026 targeting voice-model vendors accused of training on voiceprints without consent, and Texas extracted a $1.4 billion settlement from Meta in 2024 under the state's biometric statute.
Datoric's argument is that effective data collection requires people who understand the models themselves, not just labor markets for annotation. "The most effective data work happens when the people building the models and the people building the data operate within the same feedback loop," the company wrote in a blog post this past September. Reddy and Lin describe their operation as a "secure training data R&D engine for physical AI," according to their Y Combinator page. They run private, invite-only data-collection apps with what they call per-project isolation and provenance tracking, integrating research, collection, verification and production in a single loop.

Whether that model scales beyond seven figures a month is an open question. Counterpoint Research expects more than 70 percent of AI training datasets will be synthetic within one to three years, especially for autonomous vehicles and robotics where real-world footage is expensive or dangerous to capture. NVIDIA's Cosmos tooling can process 20 million hours of video in 40 days on its Hopper architecture, or 14 days on the newer Blackwell chips, the company has said. But synthetic data struggles with edge cases and the full range of human environments, which is why physical deployments still lean on real-world validation sets.
The market for this kind of data has attracted visible customers. Wayve, the London-based autonomous-vehicle startup, raised $1.2 billion at an $8.6 billion valuation in a Series D announced in February 2026, later adding a $60 million extension from AMD, Arm and Qualcomm in April. The company is training world models on driving footage to support what it calls "embodied AI for any vehicle, anywhere." Figure AI's F.02 humanoid helped produce 30,000 cars at BMW in 2025, the company said last November; the F.03 arrived at BMW's Spartanburg plant at the end of June 2026. Figure announced a partnership in September with Nscale for up to 100,000 GPUs on the NVIDIA Vera Rubin platform. Apptronik, another humanoid maker, opened a 90,000-square-foot "data factory" in Austin on June 30 to generate training footage for its Apollo 2 robot in partnership with Google DeepMind, Forbes reported. Agility Robotics, whose Digit humanoid is being tested by Amazon and Toyota, filed a preliminary SPAC prospectus in July.
The established AI labs have released their own models. Covariant introduced RFM-1, a robotics foundation model trained on text, images, video, sensor data and warehouse deployment logs, in March 2024. Physical Intelligence demonstrated laundry-folding and other generalist policies with its π0 model in 2024 and 2025. DeepMind released Genie, a world model trained on unlabeled internet video, in February 2024, and previewed an update it called "Genie 3" last August. Meta's V-JEPA, a video joint-embedding predictive architecture, arrived in February 2024; Yann LeCun described it as "a step toward a more grounded understanding of the world so machines can achieve more generalized reasoning and planning."

Datoric has posted dataset viewers on Hugging Face for what it describes as 100,000 hours of egocentric residential video, 50,000 hours of industrial egocentric footage, 2,000 episodes of robot teleoperation and 20,000 hours of conversational voice data for text-to-speech training, among others. The company page does not specify licensing or public access. In its September blog post, Datoric wrote that "it is an incredible competitive advantage to not only collect large amounts of data faster than everyone else, but to define what data should exist in the first place."
Other startups have staked out adjacent territory. Scale AI, which built annotation infrastructure for early autonomous-vehicle programs, framed a "Data Engine for Physical AI" in a 2025 blog post. A startup called Physical, at physical.inc, describes itself as "building the data engine for embodied intelligence." Sensei, from Y Combinator's Summer 2024 batch, markets itself as "Scale AI for robotics data" using networks of human teleoperators and low-cost collection hardware. Human Archive, from the Winter 2026 cohort, calls itself "data infrastructure for physical AI." Parallel Domain, which raised a $30 million Series B in 2022, sells synthetic computer-vision data and counts Toyota Research Institute as a customer.
NVIDIA has released open tooling through Isaac Lab, its Cosmos tokenizer and Project GR00T, a foundation-model initiative for humanoid robots announced in March 2024. Open-X Embodiment and its December 2025 extension, OXE-AugE, offer 4.4 million trajectories for cross-embodiment transfer research, though academic datasets typically carry licensing restrictions that block commercial use. Tesla's internal "Data Engine," described by Andrej Karpathy at a 2021 computer-vision conference, deployed models in ghost mode, logged errors, triggered targeted data collection, labeled edge cases with heavy automation and retrained in a closed loop between deployment and dataset expansion.

Datoric's stated approach resembles that iterative design. The company emphasizes that its researchers understand "why a dataset is being collected, how it will be used, what capability it is intended to unlock, and what assumptions need to be tested before a collection is expanded," as it wrote in September. Whether custom datasets or synthetic generation dominates the training mix, or whether open corpora grow fast enough to commoditize the input layer, remains unsettled. The revenue claim from a two-person team less than two years from founding does suggest that some companies would rather pay for tailored, rights-cleared datasets than build collection infrastructure themselves or risk the legal exposure that comes with scraping the internet. Datoric declined to disclose funding details or provide independent verification of the revenue figure, though the number sits on a Y Combinator profile that typically undergoes some internal review.
The bet, in the end, is that the physical-AI wave stalls or scales based on access to the right training data, and that building that data is less a labor-market problem than a research problem. If robots are going to fold your laundry or drive your groceries across town, someone has to teach them what laundry looks like when it's tangled, or what a construction zone looks like at dusk in the rain. Datoric is wagering that it can sell that specificity faster than the rest of the market can synthesize it.
