Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

Fintech iconFintechOctober 4, 2026

HIFI raises $37M for tokenized money infrastructure

HIFI raises $37M for tokenized money infrastructure
StablecoinsPayment Processing+3
Fintech iconFintechOctober 3, 2026

Pivot57 launches Africa's first institutional intelligence platform

Pivot57 launches Africa's first institutional intelligence platform
Africa TechInstitutional Finance+3
SaaS iconSaaSAugust 22, 2026

Traceforce launches on-device security for AI agents

Traceforce launches on-device security for AI agents
YcAi Agents+3
Fintech iconFintechAugust 22, 2026

Lyon builds private AI models for bank transaction data

Lyon builds private AI models for bank transaction data
YcFintech+3

Founders Mentioned

Rui Wang

EdotEnv

saas icon
SaaS

Rui Wang

EdotEnv

saas icon
SaaS
Fintech iconFintech
August 22, 2026
YcAi AgentsAutonomous TradingAi BenchmarkingReinforcement Learning

EdotEnv builds self-improving AI agents from trading data

YC S26 startup uses quantitative trading environments as benchmarks to train AI models toward recursive self-improvement, targeting frontier labs and fintechs.

EdotEnv builds self-improving AI agents from trading data

Two ETH Zürich-trained engineers left cushy jobs at a London quant shop and a Bay Area chip startup this summer to build what amounts to a gym for AI models that want to learn how to trade. The gym gets harder every time you show up.

EdotEnv, a Y Combinator outfit founded in July by Rui Wang and Michael Zhang, sells something the AI industry has been groping toward for months: a way to measure whether language models can actually improve at sequential decision-making or just memorize answers from training data. Their solution is blunt. Instead of designing another benchmark that models will saturate in six months, they're using financial markets themselves as the testing ground. Markets punish overfitting in real time, and the game resets every morning.

The premise builds on a problem that's been nagging frontier AI labs. Static question-answer benchmarks keep getting gamed. A model trains on enough examples, pattern-matches its way to a high score, and everyone declares victory until someone checks the hold-out set. Wang, who spent years at G-Research building quantitative models, watched this cycle repeat. "With all the benchmaxxing around, evals saturate," he wrote on Hacker News in early August. "Useful benchmarks should increase in difficulty as models advance."

So they built two. The first, called Alpha Autoresearch, hands AI agents anonymized minute-by-minute cryptocurrency price data from 2022 and 2023, asks them to discover predictive patterns, then evaluates performance on a hidden 2024 data set using metrics like Sharpe ratio, profit-and-loss after transaction costs, and maximum drawdown. The second runs agents through 730 consecutive trading days with a $100,000 simulated portfolio, measuring how well they adapt as market conditions shift beneath them.

Both benchmarks use a modified version of Moonshot AI's Kimi Code, a command-line agent framework that EdotEnv outfitted with custom system prompts and trading-specific tools. The company presents itself as a "Quant Neolab working toward RSI"—recursive self-improvement, the idea that models might eventually train themselves by identifying weaknesses and searching for better strategies without human supervision.

That's the pitch to frontier labs, anyway. EdotEnv says it's selling research harnesses and evaluation environments to OpenAI, Anthropic, Coinbase, and Kalshi, though the company declined to specify which relationships are active customers versus target accounts. Wang and Zhang filed the California corporation on July 23, weeks after YC acceptance, and began publishing benchmark results shortly after.

The results so far are mixed. Fable 5 and Grok 4.5 ranked strongest among the agent families EdotEnv tested, but qualitative diagnostics showed models overfitting to tiny validation windows, running shallow iterative searches, and skipping proper ablation tests. "Frontier agents can occasionally produce economically meaningful predictive models, though they do not yet do so reliably," the company's report concluded. The Long-Horizon Planning benchmark found that higher reasoning capacity improved mean and median Sharpe ratios for some models, but performance varied wildly across different rollouts.

Multiple studies published in 2026 have documented what EdotEnv's benchmarks confirm: agent performance depends as much on harness design, system prompts, and tooling as it does on base model capability. A July paper argued that reproducibility and diagnostic telemetry matter at least as much as raw scores. EdotEnv took that to heart, publishing detailed plots showing validation reuse patterns, iterative search depth, and optimizer parameter choices alongside headline metrics.

The timing matters. The algorithmic trading industry hit $20.23 billion in 2026, according to Mordor Intelligence, and generative AI topped JP Morgan's February e-Trading survey as the most influential technology across asset classes. Coalition Greenwich reported in July that 32 percent of U.S. equity brokers now use AI for real-time algorithm optimization, with another 40 percent planning adoption soon. But institutional investors remain cautious about handing over decision-making to black-box systems, especially with regulators circling.

Digital illustration for article section "Content Section 3" in "EdotEnv builds self-improving AI agents from trading data" - A serene, minimalist illustration of an elegant hourglass resting on a clean wooden desk next to an ...

The EU AI Act's transparency obligations became enforceable August 2, with high-risk system rules phasing in through December 2027. The European Securities and Markets Authority issued a February supervisory briefing linking algorithmic trading governance directly to AI Act compliance. The Financial Stability Board opened a June consultation on responsible AI adoption in finance. In the U.S., the SEC has brought multiple "AI washing" cases since 2024, including a May 28 action alleging misrepresentations about AI trading bots. The CFTC updated investor advisories warning that AI cannot predict sudden market shifts.

EdotEnv's emphasis on transparent evaluation controls and contamination-resistant benchmarks—timestamp masking, transformed targets, hidden hold-out sets—addresses both sides of that regulatory pressure. Labs need to show their models work before deployment. Regulators need audit trails that explain how decisions get made.

The company isn't alone in chasing this opportunity. Numerai has run an encrypted stock-market prediction tournament since 2015 and released a "Quantum" dataset update in July 2026. WorldQuant BRAIN launched its 2026 International Quant Championship in May. QFinZero, demonstrated at ACL 2026, provides a unified toolchain for LLM trading agents with deterministic replay. Microsoft's Qlib and the open-source FinRL framework both support reinforcement learning research in finance. Retail-focused platforms like Engine and Cod3x market "self-improving" trading agents to individual traders, with Engine claiming more than 62,000 trade decisions in a single week as of late August.

EdotEnv differentiates by positioning its environments as training grounds for recursive self-improvement rather than one-shot competitions. The company says it's working with frontier labs and academic groups to build post-training environments that continuously increase difficulty so agents can "hillclimb" toward better strategies. In a Hacker News thread, the founders acknowledged the approach remains unproven at scale. "We have experiments showing that agents at least can learn," they wrote, "but we do not yet have full post train runs, mainly due to time."

That gap between vision and execution is typical for a two-person YC Summer 2026 batch company less than two months old. Wang holds an applied mathematics degree from ETH Zürich; Zhang built LLM inference systems at Etched and ran high-frequency options strategies at TransMarket Group before joining. They co-hosted an "Autoresearch Applications Hack" in San Francisco on July 18 that featured a speaker from OpenAI, according to event listings. Jared Friedman is their primary YC partner. Funding remains undisclosed beyond standard YC terms.

The near-term competitive pressure will come from IBM Research and Hugging Face, which launched the Open Agent Leaderboard in May 2026 to compare full agent systems, and Harbor-Index, which debuted in July as a compact multi-benchmark suite. Live evaluation platforms like AI-Trader claim to offer contamination-controlled testing across U.S. stocks, A-shares, and crypto with real-time path stochasticity rather than offline resampling.

Digital illustration for article section "Content Section 5" in "EdotEnv builds self-improving AI agents from trading data" - A serene, minimalist illustration of a beautifully crafted, tiered wooden ranking display representi...

Institutional adoption will likely favor systems that can demonstrate cost-performance trade-offs and survive compliance review. Coalition Greenwich data from Q2 2026 interviews show brokers planning to expand AI use into surveillance, compliance, and strategy selection, but with human oversight remaining central. EdotEnv's focus on research harnesses for frontier labs positions it upstream of that deployment layer—selling tools to improve model training rather than ready-to-deploy trading systems.

Whether recursive self-improvement proves feasible remains an open research question. The Bank of England's July 2026 Financial Stability Report noted Project Logos, a BIS Innovation Hub experiment observing LLM-based portfolio managers in simulated markets, signaling central bank interest in systemic risk from correlated models. Buy-side adoption remains cautious despite rising electronification, according to a June Coalition Greenwich study that found institutional investors still value human judgment and explainability. Among U.S. brokers, only 12 percent currently use AI for compliance and surveillance, though 44 percent plan near-term deployment.

EdotEnv's bet is that financial markets will generate harder problems faster than any hand-labeled benchmark suite, because markets adjust as soon as patterns get discovered. The startup filed incorporation papers in late July. By early August, it had published two benchmarks showing that frontier models can occasionally find predictive signals but don't do so reliably. The road from "occasionally" to "reliably" is what venture capital gets raised to pave. Whether a two-person team can build the highway before the benchmark ecosystem moves on is the question every YC batch company eventually faces.

More stories

  • HIFI raises $37M for tokenized money infrastructure
  • Pivot57 launches Africa's first institutional intelligence platform
  • Traceforce launches on-device security for AI agents
  • Lyon builds private AI models for bank transaction data
  • Amorphic Labs launches commerce infrastructure for AI agents
  • Synthefy raises $6.5M for structured data AI models
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.