Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Healthtech & Biotech iconHealthtech & BiotechAugust 7, 2026

The Race to Build Computers from Human Brain Cells

The Race to Build Computers from Human Brain Cells
YcWetware Computing+3
Climate / Social Tech iconClimate / Social TechAugust 7, 2026

Solinas Integrity Raises $5.5M Series A for AI Water Robots

Solinas Integrity Raises $5.5M Series A for AI Water Robots
Water TechRobotics+3

Founders Mentioned

Karan Brar

hiloop

saas icon
SaaS

Karan Brar

hiloop

saas icon
SaaS
SaaS iconSaaS
August 7, 2026
YcAi AgentsAi InfrastructureAi BenchmarkingEnterprise Ai

YC's hiloop Launches Infrastructure for Self-Improving AI Agents

Summer 2026 batch company debuts platform enabling autonomous research loops, claims new benchmark on Karpathy's autoresearch test amid surging enterprise adoption.

YC's hiloop Launches Infrastructure for Self-Improving AI Agents

The numbers came in overnight—0.9016 bits-per-byte on a benchmark most people outside machine learning research circles have never heard of. But inside those circles, the result, posted by a two-person Y Combinator startup on July 2, mattered enough to ripple through the tight network of engineers building what some are calling the next inflection point in artificial intelligence: systems that improve themselves.

hiloop, a platform barely three months past its YC pitch day, had edged out Recursive Superintelligence, a well-funded research lab whose team exceeds 25 people, by less than a hundredth of a bit. Both teams ran their experiments on Nvidia's B200-class GPUs. Both were chasing the same idea—that the future of AI isn't just smarter models, but smarter processes for making those models smarter.

The benchmark itself, Andrej Karpathy's autoresearch test, is almost comically narrow by design: a coding agent tweaks a small codebase to improve a measurable metric under fixed computational budget. What makes it interesting isn't the task. It's what the task demands from the infrastructure underneath.

hiloop's CEO Karan Brar and CTO Thomas Boser orchestrated 4,188 experiments across 50 B200 GPUs over 48 hours. They used off-the-shelf coding agents from Anthropic and OpenAI with no custom prompting. No hand-tuning. Just parallelized experimentation at industrial scale, managed by the primitives their platform is built on: snapshottable compute environments, modular evaluation harnesses, lineage tracking that survives thousands of experiment forks.

If that sounds like plumbing, it is. But in 2026, plumbing might be the most valuable product category in artificial intelligence.

A Crowded Frontier, Suddenly

Something shifted in the first half of this year. The conversation in AI labs moved from "how do we build better models" to "how do we build systems that build better models." It wasn't a single breakthrough. It was a confluence.

Karpathy's autoresearch concept, articulated in blog posts and code repositories through 2025, caught mainstream attention when Fortune reported in March that Shopify's CEO had used an autonomous research loop to squeeze out a 19% performance gain on an internal task overnight. The result was modest, the method was not. Labs like Recursive Superintelligence and research teams at Anthropic started positioning recursive self-improvement—AI systems that iteratively enhance their own capabilities—as the logical next step beyond foundation model scaling.

By mid-year, the landscape had fractured into two camps. On one side: companies attempting to build full recursive systems in-house. Recursive operates a closed automated research platform and has published state-of-the-art results across multiple benchmarks. Prime Intellect opened a beta "Lab" training platform in May. EverMind launched what it calls a "Raven Agent" harness in July.

On the other side: infrastructure providers rushing to ship the components those systems will need. Modal introduced secure sandboxes for untrusted agent code in April. Weights & Biases expanded its Weave product to support continual improvement loops. Arize, Langfuse, Honeycomb—all announced agent observability tooling between May and June, each betting that whoever solves telemetry and lineage at scale wins the next decade.

The hyperscalers weren't slow to notice. OpenAI's April Agents SDK update added sandbox agents with bring-your-own-sandbox support. Google Cloud launched its Gemini Enterprise Agent Platform on April 22. Anthropic shipped managed agents and updated its usage policies in June and July, clearly watching how customers were starting to chain these systems together.

According to Gartner's May Hype Cycle for Agentic AI, 17% of organizations have already deployed agents in some capacity. More telling: over 60% plan deployment within two years. That adoption curve depends entirely on foundational infrastructure most enterprises don't have yet and may not know they need.

Money, Mandates, Momentum

Digital illustration for article section "Money, Mandates, Momentum" in "YC's hiloop Launches Infrastructure for Self-Improving AI Agents" - A minimalist, futuristic conceptual image representing the converging forces of money, mandates, and...

Three forces are converging, though perhaps not as neatly as anyone would like.

First, spending. Gartner forecasts worldwide AI spending will reach $2.59 trillion in 2026 from $1.76 trillion in 2025, up 47% year-over-year. Infrastructure accounts for more than 45% of that total. IDC projects AI infrastructure spending alone will hit $487 billion this year, tracking toward $1 trillion by 2029. Data center electricity demand is climbing 26% in 2026, per Gartner, largely influenced by AI-optimized servers, with AI-optimized servers consuming roughly 31% of data center power. The capital flowing into this space isn't abstract. It's steel and silicon and electricity contracts.

Second, regulation. The EU AI Act's transparency obligations took effect August 2, requiring disclosure when users interact with AI agents. NIST launched an AI Agent Standards Initiative in February. The White House issued an executive order on AI security in June. The UK's AI Safety Institute has been publishing agentic evaluation toolkits throughout the year. These mandates create immediate demand for infrastructure that bakes in auditability, lineage tracking, provenance—the kind of capabilities hiloop's architecture addresses through its telemetry and query engine.

Third, a subtle but important shift in research priorities. A Springer AI Review paper published in April documented the growing disconnect between benchmark scores and production requirements like cost, safety, workflow integration. McKinsey research from January and July found that organizations redesigning workflows around AI were 5.3 times more likely to deliver value than those simply piloting agents in isolation. Forrester's June "State of Agentic AI in 2026" report noted high adoption intent but uneven operational maturity.

The infrastructure gap, in other words, is cultural as much as technical. Enterprises know they want this. They don't quite know what "this" is yet.

The Decoupling Thesis

Digital illustration for article section "The Decoupling Thesis" in "YC's hiloop Launches Infrastructure for Self-Improving AI Agents" - A conceptual and minimalist architectural composition representing the decoupling of foundational in...

hiloop positions itself in that uncertainty. The platform decouples infrastructure from evaluation harnesses and compute environments—an architectural choice that sounds technical until you consider the implications. Teams can snapshot and fork experiments while maintaining full observability through a SQL query engine built on Apache DataFusion and Arrow. User-defined annotations attach typed metadata to telemetry events, enabling fast filtering over millions of rows. A model gateway keeps provider API keys isolated from sandbox environments.

The company targets ML training teams, post-training groups, and any organization where, as a June 25 blog post put it, "performance is the product." The phrasing echoes Karpathy's autoresearch framing.

Early customer signals remain thin. The homepage lists Erdos Miller, an oil and gas engineering firm, under "Trusted by," though no public case study has emerged. The YC Summer 2026 batch includes several agent infrastructure startups, suggesting the space is crowded but validation from enterprise buyers is still scarce.

Recursive Superintelligence represents the opposite approach. Rather than selling infrastructure, they're building a closed, in-house automated research system. A June 11 article detailed improvements on the NanoGPT benchmark from 79.7 seconds to 77.5 seconds using FP8 attention and other optimizations, alongside the NanoChat result hiloop later surpassed. Recursive's team exceeds 25 people and their focus is recursive self-improvement as a research agenda, not a product anyone outside the company can buy.

CoreWeave, a GPU cloud provider, announced a "Unified Agentic AI" platform in May that positions training, inference, and observability (via Weights & Biases integration) as a closed loop targeting—and here the company used the word without apparent irony—"superintelligence." Crescendo launched an "Optimization Agent" for customer experience in June, framing recursive improvement as a vertical play.

The patterns vary. The underlying thesis does not: the process that improves AI should itself improve. Whether that process belongs inside a platform, a lab, or a tightly integrated cloud service is the question the market hasn't settled.

What Comes Next (Or Doesn't)

Digital illustration for article section "What Comes Next (Or Doesn't)" in "YC's hiloop Launches Infrastructure for Self-Improving AI Agents" - A minimalist, futuristic conceptual image representing parallelized experimentation and infrastructu...

hiloop's July 2 blog post carried a title borrowed from early neural architecture search literature: "Search is enough." The claim is that parallelized experimentation over stock agents can outperform hand-tuned systems, but only if the infrastructure handles lineage, reproducibility, and cost at scale.

Whether enterprises adopt that thesis depends on factors well beyond technical performance. Anthropic's Institute warned in a June 5 essay that full recursive self-improvement could increase control-loss risks, emphasizing the need for oversight tooling and evaluations. arXiv saw a flurry of recursive SI papers in July—surveys, new methods for harness-level improvement, proposals for governed multi-agent architectures. The academic conversation is moving faster than regulatory clarity, which is perhaps not unusual, but still unsettling.

For founders and infrastructure architects, the opportunity is visible even if the shape isn't. The industry knows it needs sandboxes, observability, evaluation harnesses. It does not know whether those primitives consolidate into platforms like hiloop or remain loosely coupled services stitched together by orchestration layers. It does not know how much of the improvement loop should be automated versus human-supervised. And it certainly does not know at what point recursive improvement crosses from engineering optimization into something that requires fundamentally different governance.

hiloop's benchmark achievement is a data point, not a conclusion. The infrastructure for self-improving AI is being built in real time. The teams that succeed will be the ones that solve for reproducibility, safety, and cost before the loops close too tightly to intervene. That's the bet, anyway. The market will decide if anyone placed it correctly.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • The Race to Build Computers from Human Brain Cells
  • Solinas Integrity Raises $5.5M Series A for AI Water Robots
  • The Race to Hack AI Agents Before Attackers Do
  • Black-Owned Knockturnal Launches AI Healthcare Booking in South Africa
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.