Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSJuly 13, 2026

Robot Learning's Data Quality Problem Gets a YC-Backed Solution

Robot Learning's Data Quality Problem Gets a YC-Backed Solution
YcRobotics+3
Fintech iconFintechJuly 13, 2026

Stoa Raises $2.4M to Turn Idle Cash Into Lifestyle Perks

Stoa Raises $2.4M to Turn Idle Cash Into Lifestyle Perks
FintechDigital Banking+3

Founders Mentioned

Jay Chooi

Robocurve

saas icon
SaaS

Jay Chooi

Robocurve

saas icon
SaaS
SaaS iconSaaS
July 13, 2026
YcRoboticsAi BenchmarkingOpen SourceRegulatory Compliance

Open-Source Robot Benchmarks Tackle Industry's Transparency Gap

YC-backed Robocurve launches Inspect Robots, an open framework for standardized real-world testing, as robotics deployments scale and regulators demand transparency.

Open-Source Robot Benchmarks Tackle Industry's Transparency Gap

There's an uncomfortable conversation happening in robotics labs these days, usually in hushed tones after the cameras stop rolling. Everyone's simulations are perfect. The demo reels are spectacular. But when these machines hit factory floors or distribution centers, something always seems... off. Germany's Fraunhofer Institute called it out this past spring: the gap between what companies present and what their robots can actually do has grown too wide to ignore anymore.

Which makes the timing awkward, to say the least.

The International Federation of Robotics counted 542,076 industrial robot installations in 2024—the sector's held above half a million units annually now—while service robotics revenue is climbing from $26.35 billion in 2025 toward projections that hit $131.9 billion by 2034. BMW put Figure's humanoid robots on the line at its South Carolina plant back in June 2026, following prior use of Figure 02. Agility Robotics took itself public via SPAC on June 24, 2026, waving commercial contracts with Toyota and Amazon. The hardware is shipping, the business models are supposedly scaling, and yet nobody in the industry can agree on a basic question: how do you actually measure whether these things work?

When the Metrics Don't Exist

The structural problem runs deep. Most robotics research still benchmarks in simulation or sanitized lab setups. Getting those results to transfer to messy reality remains notoriously hard—researchers have a term for it, the "reality gap," and it's been exhaustively documented in academic surveys. When companies do report real-world performance, you're typically looking at vendor-supplied numbers with methodologies that are, shall we say, opaque.

"There are no well-run, standardized robotics benchmarks today." That's according to the YC profile of Robocurve, a two-person startup that came through an accelerator batch earlier this year. The founders, Jay Chooi and Aris Zhu, both have evaluation credentials—Chooi from the UK AI Security Institute's Inspect Evals project, Zhu from stints at Amazon Robotics and AGI Labs. Their pitch is pointed: the robotics industry lacks anything resembling what autonomous vehicles have in CARLA or nuScenes leaderboards. No independent, continuous, real-world testing that anyone outside the vendor's PR team can actually trust.

An Open-Source Gambit

Robocurve released version 0.6.0 of something called Inspect Robots back in July 2026—an open-source evaluation framework under an MIT license. The architecture sounds simple on paper: define a benchmark task once, then run any policy (vision-language-action model, classical controller, whatever) on any robot or simulator through a plugin system. Tasks log to structured formats and visualize via Rerun, an open SDK that's become common tooling in robotics labs.

The design borrows heavily from language model evaluation frameworks—standardized interfaces, reproducibility by default, safety-aware error handling for unattended runs. Robocurve's already published some public benchmarks under its GitHub: WorldEvals as a catalog, KitchenBench for bimanual manipulation, adapters for specific robotic arms. According to their YC profile, the team has "ran our first pilot scoring a frontier model on a real robot." (Past tense, singular pilot—not exactly proof at scale yet.)

It's early days. The company structured itself as a public benefit corporation. Its sparse website talks about forecasting "physical automation for societal preparedness," the kind of mission-statement language that could mean everything or nothing. But the timing of the launch suggests Robocurve isn't alone in sensing an opening here.

Everyone Arrives at Once

Digital illustration for article section "Everyone Arrives at Once" in "Open-Source Robot Benchmarks Tackle Industry's Transparency Gap" - A clean, minimalist conceptual image focusing on the standardized analysis and benchmarking of human...

Fraunhofer IPA rolled out standardized third-party analyses for humanoid robots in May, explicitly targeting that credibility gap. MLCommons announced an Edge Agentic Inference benchmark in July—initially focused on on-device LLMs but relevant to robotics controllers down the line. A startup called Positronic Robotics debuted its PhAIL physical-AI benchmark in March. RoboArena, a distributed evaluation platform with roots in the DROID dataset, is running what it calls a live benchmark through the end of the year. Carnegie Mellon published work on scalable real-to-sim approaches for vision-language-action models. The COMPARE ecosystem is attempting to bring coherence to open-source manipulation benchmarking, though "coherence" might be generous.

Even the major platforms are making moves. NVIDIA announced something called Halos for Robotics in June—a full-stack safety system with an ANAB-accredited inspection lab, targeting compliance paths like IEC 61508 and ISO 13849. Agility Robotics was the first adopter. That's primarily a safety play, not pure evaluation, but it reflects the same underlying pressure: as robots move from pilots to production lines, someone needs to certify they're safe and effective. Preferably someone without a sales quota.

The Regulators Are Coming

The EU AI Act's transparency requirements took effect this past August. High-risk AI systems—which can include robotics in industrial settings—will face conformity assessment requirements starting next year for standalone systems, with embedded product requirements following shortly after. In the U.S., ANSI standards for industrial mobile robots were reaffirmed in April. NHTSA continues developing evaluation frameworks for autonomous driving systems, a parallel regulatory track that robotics will almost certainly follow, whether the industry wants it to or not.

"Near-term adoption depends on hardware supply chain, safety and regulatory frameworks, and reliable, repeatable performance in customer workflows," McKinsey noted in a June piece that interviewed MIT's Daniela Rus on the practical limits of current AI approaches. The shift "from tools to teammates," as they framed it, implies that standardized, externally trusted evaluation isn't some nice-to-have feature. It's the bridge to enterprise adoption, full stop.

The AV Parallel, For Better or Worse

The autonomous vehicle analogy cuts both ways, of course. AV development converged around shared datasets and leaderboards partly because regulators demanded proof—and partly because enough money was flowing into the sector to fund the infrastructure. Robotics deployments haven't faced quite the same scrutiny yet. But that's shifting as humanoids enter automotive plants and warehouse floors at meaningful scale.

Goldman Sachs has projected the humanoid market could reach $38 billion by 2035 in a base-case scenario, with an optimistic path to $154 billion. (That report came out in early 2024, though it's still making the rounds in corporate pitch decks.) Whether those numbers hold depends, in no small part, on whether buyers can actually compare performance claims across vendors.

Whether Inspect Robots becomes the standard or just another framework in a fragmented landscape will come down to adoption, not elegant architecture. Open source lowers barriers to entry, sure. But real-world benchmarking is expensive. Someone has to buy the robots, maintain test environments, run evaluations continuously without cherry-picking the good runs. Robocurve hasn't disclosed funding or detailed its business model yet. Other YC companies are tackling adjacent pieces of the puzzle—Physical Turing with real-world humanoid testing, Cortex AI with human-in-the-loop data marketplaces—but the ecosystem remains scattered.

No More Status Quo

Digital illustration for article section "No More Status Quo" in "Open-Source Robot Benchmarks Tackle Industry's Transparency Gap" - A sleek, state-of-the-art humanoid robot meticulously interacting with a polished, minimalist car ch...

What's increasingly clear is that the industry can't afford to muddle through much longer. When BMW puts Figure robots on a production line moving tens of thousands of vehicles, "the demo looked impressive" doesn't cut it as a performance metric. Procurement teams evaluating six-figure robotic systems want apples-to-apples comparisons. Regulators writing safety standards need baseline performance data that isn't marketing copy. And founders building the next generation of embodied AI probably deserve to know whether their sim-to-real transfer actually works before they pitch customers, if only to save everyone some time.

The transparency gap isn't really a technical problem anymore, if it ever was. It's a market structure problem. And markets that scale—particularly those dealing with physical systems that can injure people—eventually demand standardization. There's always an awkward period before that consensus emerges, when the old guard resists and newcomers scramble for position.

Perhaps that's why a two-person startup with an open-source framework thinks it can compete with research institutes and corporate labs. In the absence of a trusted referee, anyone with a decent protocol and plausibly neutral incentives has a shot. Whether Robocurve or someone else fills that role, the role itself seems unavoidable now. Too much capital, too many robots, too many factory floors. Someone's going to build the leaderboard eventually, even if the industry would rather keep showing demo reels.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Robot Learning's Data Quality Problem Gets a YC-Backed Solution
  • Stoa Raises $2.4M to Turn Idle Cash Into Lifestyle Perks
  • Brain Cells as Compute: The Biotech Startups Racing to Solve AI's Energy Crisis
  • Avyn Emerges From Stealth With AI Agent Built to Replace VC Analysts
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.