A San Jose startup that calls itself a "Data Refinery for Physical AI" won't disclose how much money it has raised, but DeepReach said in January that its annual recurring revenue doubled that month. The claim, posted to LinkedIn, offers a glimpse of a budding market that venture capitalists and industrial giants believe could determine whether humanoid robots and embodied AI systems remain lab curiosities or transform factory floors and warehouses worldwide.
The prize is access to the messy, high-fidelity information that robots need to learn: sensor streams from hundreds of environments, synchronized action labels across hardware platforms, diverse task demonstrations that span the mundane and the complex. It is the kind of data that cannot be scraped from the internet or conjured by a language model. It must be captured in the physical world, labeled by hand or sophisticated pipelines, and refined into training sets that cost millions to assemble.
Global industrial robot installations hit 542,000 units in 2024, according to the International Federation of Robotics, with a forecast climb to 575,000 in 2025. U.S. installations rose 11 percent year-over-year to 38,000 units in 2025, the IFR reported in June of this year. Yet those figures mask a fundamental shift. Traditional fixed-arm systems are giving way to AI-powered mobile platforms and humanoid robots that demand orders of magnitude more training data than their predecessors.
BMW announced in late June that its Spartanburg, South Carolina, plant had begun deploying Figure 03 humanoid robots for complex logistics sequencing. The automaker noted that the prior generation, Figure 02, supported production of more than 30,000 X3 vehicles over 10 months last year. "Plant Spartanburg is the birthplace of humanoid robotics in BMW Manufacturing's operational day-to-day activities," Ulrich Wieland, the company's vice president of production control and logistics, said in a press release. Brett Adcock, Figure's founder and chief executive, said the deployment "proved that humanoids are no longer lab experiments."
Amazon unveiled its DeepFleet foundation model and agentic workflow system, Project Eluna, in June, though the company quietly halted its Blue Jay pilot in February after less than six months of testing. KION Group and logistics operator GXO deployed the first AI-supported autonomous industrial truck at a site in Épinoy, France, in March, powered by NVIDIA Omniverse-based digital twins. Agility Robotics opened a Fremont facility in July to accelerate physical AI development and filed a confidential S-4 for a proposed business combination two weeks later.
Goldman Sachs Research estimated in June that global AI infrastructure spending will reach $7.6 trillion from this year through 2031, covering compute, data centers, and power, as AI expands into what the firm calls the "physical economy." The bank's 2024 analysis pegged the humanoid robot market at roughly $6 billion within 10 to 15 years. NVIDIA chief executive Jensen Huang said in early June that "humanoid robots will bring physical AI to the world's largest industries, opening a multitrillion-dollar economic opportunity." NVIDIA announced its Isaac GR00T Reference Humanoid Robot the same day, a research platform built on the Unitree H2 Plus body with Sharpa hands and Jetson Thor compute. The company released GR00T foundation models and the Isaac Lab-Arena evaluation framework in January.
But several industry observers argue the real constraint is not computational power or model architecture. It is data.
"Embodied AI and robotics are bottlenecked not by models, but by access to the physical world," DeepReach wrote on its LinkedIn profile. Hugging Face, the AI collaboration platform, framed the challenge more bluntly in a July blog post: "Physical AI Is Entering the Systems Phase," arguing that the robot-data layer is "one of the most important and least understood parts" of the technology stack. Andreessen Horowitz echoed the theme in January, publishing "The Physical AI Deployment Gap" and emphasizing that real-world data collection, simulation, and deployment infrastructure would become the defensible competitive moat, not larger models.
The scarcity reflects the complexity of the capture process. Robotics foundation models require diverse scenarios, high-fidelity sensor streams, and synchronized action labels across hardware platforms that vary by vendor. The DROID dataset, published in March 2024 and updated through last year, contains 76,000 demonstrations spanning 350 hours, 564 scenes, and 84 tasks. It remains a benchmark for real-world manipulation data. The Open X-Embodiment collaboration, led by Google DeepMind and academic labs from 2023 through last year, consolidated multi-robot datasets that remain a core pretraining resource. AXIS, a browser-based data engine released in July, introduced automatic validation and augmentation pipelines that the authors said improved success rates versus RoboCasa365's simulated kitchens.

Apptronik announced raising over $935 million in February and opened its Robot Park data-collection facility on the last day of June. The company said fleets of Apollo 2 humanoids "continuously collect real-world data" at the site "in partnership with Google DeepMind" to train future foundation models. Apptronik declined to specify dataset sizes or model benchmarks.
Scale AI, which built annotation infrastructure for autonomous vehicles, launched its Data Engine for Physical AI in October of last year, offering collection and labeling pipelines for robotics foundation models. Competitors have emerged across the value chain. Robotic Data operates real-world datasets at what it describes as global scale for autonomous vehicles, humanoids, and drones. Merit Data provides real-world robotics data collection and curation. Operant Data specializes in custom humanoid and warehouse task datasets. Mecka.ai markets "data, evaluation, and deployment infrastructure" alongside commissioned collection services. Egosense.ai sells egocentric point-of-view datasets and has criticized teleoperation as unscalable in vendor materials. Teleop.com offers "certified human operators" for synchronized state-action episodes.
DeepReach lists NVIDIA, Princeton University, UC Riverside, and Unitree among logos on its website under "Trusted by," but the company declined to disclose customer names, dataset sizes, quality metrics, or benchmarked model performance gains. The startup describes itself on LinkedIn as operating "a global network of real-world sites, collection teams, and scenario coverage" with "compliance-verified data." DeepReach is a member of NVIDIA Inception as of January. The company employs between two and 10 people, according to LinkedIn.
A French advertising-technology company called DeepReach, which operates at deepreach.com, is unrelated. Y Combinator's public directories do not list the San Jose startup under the Summer batch as of mid-August, and any claimed affiliation should be treated as unverified.
Cross-border data collection faces tightening regulatory scrutiny. The European Union's AI Act imposes transparency and documentation obligations on general-purpose AI models starting August 2, with staggered high-risk compliance deadlines extending to 2027 and beyond depending on sector and national sandboxes, according to the European Commission. California's Privacy Protection Agency finalized regulations for Automated Decisionmaking Technology, including privacy risk assessments and cybersecurity audits, effective January 1, with ADMT compliance deadlines for certain uses extended to January 1, 2027. China implemented certification measures for cross-border personal data transfers under the Personal Information Protection Law on January 1, requiring security assessments, standard contracts, or certification for data leaving the country, according to the Library of Congress.
OSHA references ANSI/RIA R15.06 and ISO 10218 for robot risk assessments and collaborative-robot safeguarding, though no federal robotics-specific standard exists. A House Homeland Security subcommittee hearing in early June, titled "The AI Security Landscape," included testimony on frontier models, agentic AI, and critical infrastructure risks tied to physical AI systems.
NVIDIA said late-year availability is expected for the GR00T reference robot, which the company positioned as an accelerant for faster data collection and policy iteration on humanoid platforms. Waymo published its World Model for autonomous vehicle simulation in February, offering text-, layout-, and action-controllable generative environments. Wayve released GAIA-3 in December for closed-loop safety evaluation. Both systems point toward hybrid strategies in which real-world data seeds generative simulations that produce augmented training sets. RynnWorld-Teleop, published in July, demonstrated that digital teleoperation via a generative world model could improve success rates when combined with real datasets.
McKinsey, in a late-May report, framed physical AI as the next frontier in industrial automation and highlighted "teach-less robotics" as a tipping point. Deloitte claimed in early March that 58 percent of companies already use physical AI to some extent, with 80 percent adoption projected within two years, though the figures derive from a self-reported enterprise survey.
Deployment gaps remain. Amazon's Blue Jay cancellation in February and mixed pilot outcomes across the industry underscore the long systems-integration phase ahead. Data availability, quality assurance, and safety protocols are rate-limiters, perhaps more constraining than executives anticipated even a year ago. Andreessen Horowitz argued in essays published this year that defensible advantages will accrue to teams that solve compliance-by-design workflows, human-in-the-loop economics, and evaluation infrastructure.

For robotics founders, the calculus is shifting. Coverage depth across domains, tasks, and sites; quality controls for filtering, ranking, and augmentation; and cross-border compliance workflows now compete with model architecture as sources of differentiation. Investors evaluating physical AI infrastructure should look for quantified performance deltas tied to data refinement steps, not marketing claims about dataset size. Enterprise automation leaders face a choice: build proprietary data pipelines in-house or rely on vendors whose benchmarks and compliance practices remain largely unverified in public filings.
The systems phase has begun. But the scaffolding is still under construction.
