Enact emerged from Stanford this past summer through Y Combinator's accelerator with a premise that sounds technical but speaks to robotics' central dilemma: building robots that can learn is one thing, making them work reliably after deployment is another. The startup describes itself as "the post-training layer for physical AI," generating targeted datasets by running robot policies, catching failures, and testing whether new data fixes what broke.
The company has shared almost nothing else. No founder names. No funding figures. No customer list. That silence is notable against the backdrop of an industry grappling with what comes after the demo video.
Robot installations worldwide reached 542,000 units in 2024, more than double the tally from a decade earlier, according to the International Federation of Robotics in a September 2025 press release. Yet the gulf between controlled demonstrations and field reliability remains wide. Figure spent 11 months deploying humanoid robots at a BMW plant before redesigning components for the next generation. Amazon launched its Blue Jay humanoid in October 2025 with optimistic language about safer, smarter work; by February 2026, an update to the original blog post disclosed the program had ended.
When Foundation Models Meet Factory Floors
Robotics companies today confront what might be called a deployment paradox. Foundation models trained on internet-scale data transfer gracefully to new language tasks. Physical environments are less forgiving. Distribution shift, unmodeled dynamics, and edge-case failures proliferate in ways that text and image models rarely encounter.
Figure's F.02 humanoid contributed to production of more than 30,000 BMW X3 vehicles at the Spartanburg plant, loading over 90,000 parts across more than 1,250 runtime hours, the company reported in November 2025. That deployment drove wrist and forearm electronics redesigns for the Figure 03 generation. Real-world reliability learnings, in other words, loop back into hardware and software iteration—sometimes forcing it.
Amazon's Blue Jay trajectory traced a different arc. "The goal is to make technology the most practical, the most powerful tool it can be—so that work becomes safer, smarter, and more rewarding," Tye Brady, Chief Technologist at Amazon Robotics, said in an October 2025 blog post announcing the humanoid. Four months later, an update to that same post carried a single line: "Amazon is no longer utilizing Blue Jay in operations." No elaboration followed.
Industrial robot installations in the United States climbed 11 percent year-over-year to 38,000 units in 2025, the IFR reported in June 2026. Professional service robots in transportation and logistics hit 102,900 units in 2024, up 14 percent, with robot-as-a-service fleets growing 31 percent to more than 24,500 units, according to the IFR's World Robotics 2025 executive summary. Goldman Sachs projected a base case of $38 billion for the humanoid robot market by 2035 in a February 2024 analysis. Gartner predicted in January 2026 that fewer than 20 companies will scale humanoid robots for manufacturing and supply chain to production stage by 2028—a forecast that reflects both the market's potential and its velocity constraints.
The Data Quality Thesis
The reliability gap pivots on data. "The future of robotics AI will be shaped as much by data quality as by model architecture," James Davidson, Chief AI Officer at Teradyne Robotics, wrote in a March 2026 blog post announcing a partnership between Scale AI and Universal Robots. That deal embedded Scale's Physical AI Data Engine into the UR AI Trainer, with plans to release a dataset tapping Universal Robots' installed base of more than 100,000 collaborative robots.
Academic research has begun quantifying the problem. Mobile ALOHA, a January 2024 paper authored by researchers including Chelsea Finn, showed that co-training with 50 demonstrations per task alongside static data increased success rates by up to 90 percent on complex mobile manipulation tasks. A Carnegie Mellon technical report proposed post-training reinforcement-learning adapters for vision-language-action models, aligning them with deployment state distributions. TEMPO, published in August 2026 on arXiv, explored semantic-action decoupled reinforcement learning for post-training. SC3-Eval, a June 2026 paper, proposed using video-generation-based evaluation to grade robot policies without full real-world rollouts—a potentially cheaper path to assessing performance.
Human-in-the-loop recovery methodologies documented in Hugging Face's LeRobot toolkit let operators turn failures into new fine-tuning data. DROID, a March 2024 dataset release with 76,000 demonstrations across 350 hours and 564 scenes, emphasized in-the-wild diversity and encouraged co-training with small amounts of in-domain data. Open X-Embodiment, a multi-institution dataset hub maintained through 2026 by Google DeepMind, unified many robot datasets for training vision-language-action models.
Regulatory pressures amplify the need for auditable workflows. The EU AI Act entered general application in August 2026, with high-risk obligations on AI safety components phasing in through late 2027 and mid-2028, according to the European Commission. The United States revised its industrial robot safety standards, adopting ISO 10218 to replace the 2012 edition. NIST's AI Risk Management Framework, released in January 2023, gained a Generative AI Profile in July 2024.
What Deployment Actually Looks Like

Figure's BMW deployment offers one window into post-training at scale. The company reported cycle time and placement accuracy exceeding 99 percent per-shift targets, with a zero-intervention goal. Reliability analysis from that 11-month run fed directly into Figure 03 hardware redesigns, according to the company's November 2025 report. The updated humanoid arrived at BMW Plant Spartanburg in late June 2026, running the Helix 02 software stack.
Scale AI's Universal Robots integration targets a different inflection point: robots already deployed in automotive, electronics, and logistics factories. "Together, we're creating a foundation where imitation learning can move beyond isolated research projects and become a scalable, industrial capability," Davidson wrote. The partnership embeds data-capture and labeling workflows into existing installations, turning operational environments into training grounds.
Two other Y Combinator Summer 2026 startups tackle adjacent slices of the problem. Instance, founded by CEO Claire Mao and CTO Lucy Cai, provides automated evaluations for robot policies, starting with a success detector benchmarked on more than 10,000 held-out human-labeled episodes across seven robot platforms. The company plans automated reset rigs to reduce the cost of multi-episode evaluation. Parametric, also YC S26, uses reinforcement learning feedback and judge models to tune robots from customer feedback, emphasizing reliability and throughput over form factor.
Enterprise tooling providers occupy a parallel track. Viam offers device, telemetry, and data-management platforms with fleet operations features. PickNik's MoveIt Pro provides deployment and fleet management for robotic arms. Intrinsic Flowstate, Alphabet's web-based environment for building and maintaining robotic applications, went live in July 2026 and announced a FANUC integration later that year. NVIDIA open-sourced Isaac Sim 5.0 in mid-2026, adding teleoperation workflows and OpenUSD-based pipelines, and unveiled the Isaac GR00T N1 open humanoid foundation model at its 2025 GTC conference.
McKinsey analyses published in late 2025 and 2026 cited supply chain bottlenecks, costs, and the need for robust task-specific data, noting limited safety standardization for humanoids compared to collaborative robots and industrial arms. A June 2026 paper, "What Are We Actually Benchmarking in Robot Manipulation?," audited common benchmark suites including LIBERO and CALVIN, cautioning against misinterpretation of benchmark gains. An EMNLP 2025 paper titled "Subtle Risks, Critical Failures" showed that language-model-based policies often miss situational and physical risks—evidence supporting rigorous evaluation design.
The Velocity Constraint

The post-training opportunity opens as foundation models get cheaper and more capable yet remain brittle in real-world deployment. Octo, an open-source generalist robot policy trained on roughly 800,000 trajectories and published at RSS 2024, supports efficient fine-tuning. OpenVLA, presented at CoRL 2025, ships with training notebooks and checkpoints. Diffusion Policy, published in IJRR 2024 with a one-step distillation variant at ICML 2025, demonstrated strong real-world performance before speed optimizations. LiNeS, an October 2024 paper, proposed post-training layer scaling to preserve generalization while improving task performance.
But the constraint is velocity. Agility Robotics opened a Fremont, California facility in July 2026, claiming more than $300 million of multi-year orders for its Digit v5 humanoid subject to milestones and a customer pipeline exceeding 30 companies. 1X opened a NEO factory in Hayward, California for full-scale production with consumer shipments planned later in 2026, citing NVIDIA Jetson Thor and the Isaac platform in an April release. Gartner's prediction of fewer than 20 companies reaching production scale by 2028 reflects the reliability-first pace of industrial adoption—and perhaps an acknowledgment that moving from prototype to production is harder than the initial optimism suggested.
Founders and AI teams building embodied systems should watch three signals, maybe more. Whether data-as-a-service vendors beyond Scale AI emerge with turnkey capture-label-evaluation workflows for specific verticals. Whether automated evaluation infrastructure becomes table stakes for deploying model updates safely—success detectors, sim-to-real bridging methods like the December 2025 PolaRiS framework, video-generation-based grading systems. Whether regulatory documentation burdens under the EU AI Act and refreshed standards tip the economics toward third-party post-training services that bundle compliance artifacts with performance improvements.
Enact's timing sits at that inflection. The tagline frames the gap between a model that works in simulation and one that ships to production. Whether the startup captures that opportunity depends on execution the limited public information does not yet reveal: customer contracts, demonstration hardware, pricing model, team pedigree. What the research and deployment patterns do show is an industry moving from the question of whether robots can learn to the harder question of how they improve after deployment—and who builds the infrastructure to make that improvement systematic rather than heroic.
The startups that figure out the latter may matter more than the ones that solved the former. Learning is table stakes now. Reliability at scale remains the unsolved problem, and the companies addressing it are just starting to emerge from the noise.
