Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Healthtech & Biotech iconHealthtech & BiotechAugust 23, 2026

Taudia raises $4M to scale blood tests for Alzheimer's

Taudia raises $4M to scale blood tests for Alzheimer's
HealthtechBiotech+3
SaaS iconSaaSAugust 23, 2026

Synthefy raises $6.5M for structured data AI models

Synthefy raises $6.5M for structured data AI models
Seed FundingAi+3

Founders Mentioned

Nalu Concepcion

Idler

saas icon
SaaS

Tony Goss

Idler

saas icon
SaaS

Nalu Concepcion

Idler

saas icon
SaaS

Tony Goss

Idler

saas icon
SaaS
SaaS iconSaaS
August 24, 2026
YcAi BenchmarkingAi TestingB2b SaasStartup Funding

idler raises $9M to build frontier AI evaluation tools

The YC-backed startup sells benchmarks and testing environments to leading AI labs, as model evaluation becomes critical for safety and compliance.

idler raises $9M to build frontier AI evaluation tools

A San Francisco company betting that artificial intelligence labs need better ways to test their systems has pulled in $9 million from investors led by Paradigm, the crypto-focused venture firm that has been expanding into AI infrastructure.

Idler, which emerged from Y Combinator in summer 2025 and describes itself as a "frontier data research lab," builds the specialized benchmarks and evaluation environments that AI developers use to measure whether their models can actually perform tasks like writing code or navigating e-commerce sites. The seed round also drew backing from Long Journey Ventures and several angel investors, including Feross Aboukhadijeh and Dan Posch.

The company's pitch arrives at a moment when model evaluation has become a pressure point for the industry. Regulators are demanding standardized testing protocols—the EU's AI Act now mandates adversarial testing for foundation models—while labs themselves have grown wary of the gap between flashy benchmark scores and real-world performance. OpenAI published guidance on trustworthy third-party evaluations in May 2026, a tacit acknowledgment that the field needs more rigor.

Idler's three founders—Ivan Chub, Nalu Concepcion, and Tony Goss—are building evaluation datasets pulled from actual production environments rather than academic test suites. Their public benchmark suite includes ShelfLife, a 200-task e-commerce challenge built in partnership with what they describe as a live, profitable retailer. When tested between August 13–18, 2026, Anthropic's Claude Opus 5 managed a 69% pass rate. Another benchmark, CorpLaw, draws on anonymized law firm data; Claude scored just under 43% on that battery of 50 legal tasks.

The startup has developed thousands of coding evaluations, though it has not publicly disclosed any named enterprise or frontier lab customers. Concepcion previously worked at Microsoft and Poshmark, while Chub cofounded Dark Forest and worked at Facebook. Goss, the chief technology officer, shares the Dark Forest connection and studied at Tufts. The team has grown to 12 people.

What Idler is selling, essentially, are bespoke datasets and reinforcement learning environments tailored to what foundation model companies actually need to test. Beyond the public benchmarks, the startup offers evaluation collections for long-horizon software engineering tasks across six programming languages, cybersecurity challenges, and what it describes as "recursive self-improvement tests." Those custom offerings are available by request, with no public metrics available, a model that suggests the company is positioning itself as a specialized vendor rather than an open-source contributor.

Digital illustration for article section "Content Section 2" in "idler raises $9M to build frontier AI evaluation tools" - A minimalist, futuristic conceptual representation of a bespoke reinforcement learning environment, ...

The capital will go toward expanding the research team, scaling infrastructure, and accelerating dataset development—standard deployment for a seed round, though perhaps more urgent given how quickly the AI landscape is shifting. NIST released an initial draft of its TEVV-Athlon evaluation framework in late July and plans to run a Generative AI Text Challenge next year with a sequestered testbed. The UK's AI Security Institute is publishing research on pre-release testing. The regulatory scaffolding is being built in real time.

Open-source alternatives exist. EleutherAI maintains the lm-evaluation-harness; OpenAI has its own Evals framework. But neither offers the kind of production-sourced, commercial-grade datasets Idler is betting will command premium pricing. Whether foundation model labs will pay for that remains an open question, though Paradigm's partners Alpin Yukseloglu and Frankie clearly think the market is there. "Great team executing against a timely problem," Paradigm wrote on LinkedIn, in the understated language venture firms use when they're trying not to oversell a bet.

Idler is hiring for multiple engineering roles and a special projects operator, the kind of broad recruiting signal that suggests a company planning to move quickly.

Digital illustration for article section "Content Section 3" in "idler raises $9M to build frontier AI evaluation tools" - A conceptual, highly restrained minimalist composition symbolizing a fast-moving company hiring for ...

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Taudia raises $4M to scale blood tests for Alzheimer's
  • Synthefy raises $6.5M for structured data AI models
  • Sophiie AI raises $5M to automate tradie admin work
  • HexSeed raises $800K to turn CO2 into diamond coatings
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.