David Silver has a theory about what's wrong with artificial intelligence. And in April 2026, he launched his new venture with enough conviction—and capital—to do something about it.
"Human data is like a kind of fossil fuel," Silver told Wired on the day his new company stepped out of stealth mode. By contrast, he argued, "systems that learn for themselves [are] a renewable fuel… learn forever."
Coming from almost anyone else in the AI world, the statement might scan as provocative but hollow. Coming from Silver, it carries different weight. This is, after all, the researcher who built AlphaGo, the system that in 2016 dismantled Lee Sedol, one of the greatest Go players in history, in a match that redefined what machines could do. Silver isn't merely theorizing about a post-LLM future. He's betting $1.1 billion on it.
That figure represents the seed round for Ineffable Intelligence, his new startup, which emerged from stealth on April 27, 2026 with a $5.1 billion post-money valuation. It's the largest seed round in European tech history—and a direct challenge to the industry's prevailing orthodoxy around large language models.
A Wager Against the Grain
The funding comes from a syndicate that reads like the VIP section of deep tech investing: Sequoia Capital and Lightspeed Venture Partners co-led, with participation from NVIDIA, Google, Index Ventures, DST Global, EQT, BOND, Evantic, Flying Fish, the British Business Bank, and the UK Sovereign AI Fund. That's not a typical seed roster. It's the kind of lineup you see when investors believe they're funding a paradigm shift, not a product.
What they're backing is Silver's conviction that the LLM-centric approach dominant in AI today "will fail" because it fundamentally copies human intelligence rather than discovering new forms of it. Ineffable's stated mission—"making first contact with superintelligence"—frames the ambition in almost metaphysical terms. Build a "superlearner," the pitch goes, that "discovers all knowledge from its own experience" through reinforcement learning, untethered from human-generated data.
It's a bold thesis. Perhaps bolder than Silver initially expected it would be when he articulated it in a founder's note dated January 15, 2026, months before going public. In that note, he positioned Ineffable as a place where the "reinforcement learning paradigm can flourish," prioritizing long-term science over near-term commercialization. The company's website suggests that if successful, the impact would be "of comparable magnitude to Darwin."
Lightspeed described the investment as "the next logical step in a decade of research that progressively stripped human priors from AI." Sequoia called it a bet on the "Era of Experience." The consistency in messaging across investor statements is striking—this doesn't look like a diversified hedge. It looks like a conviction play.
The Pedigree
Silver's credibility isn't abstract. It's built on more than a decade of work at Google DeepMind, where he led the reinforcement learning team from 2013 until leaving to start Ineffable in early 2026. AlphaGo made headlines, sure. But it's the trajectory that followed which seems to be what investors are really funding.
AlphaZero, released in 2017, learned chess, Go, and shogi without studying human games—it taught itself purely through self-play. MuZero in 2019 extended the approach to environments where even the rules weren't known in advance. AlphaStar reached Grandmaster level in StarCraft II the same year. AlphaDev, in 2023, discovered faster sorting algorithms than humans had found in decades. AlphaProof, released in 2024, demonstrated advanced mathematical reasoning capabilities using verifier-grounded reinforcement learning.
Each iteration reduced dependence on human knowledge. AlphaGo began with supervised learning on human game records. By the time MuZero arrived, the system was learning world models and decision-making simultaneously, with no human games required. The research arc is clear, even elegant: progressively eliminate human priors and let the system discover structure through raw experience.
Silver's most recent academic work—a December 2025 paper on discovering state-of-the-art reinforcement learning algorithms—hints that the ambition now extends to meta-learning. Not just systems that learn tasks, but systems that discover how to learn.
Reinforcement Learning's Quiet Comeback

Ineffable is launching into an environment where reinforcement learning has, somewhat quietly, moved from research curiosity to production primitive. At least in certain domains.
The evidence is scattered but consistent. DeepSeek-R1, detailed in papers from January and September 2025, used reinforcement learning to incentivize reasoning capabilities, achieving performance comparable to OpenAI's o1 model on several benchmarks. The approach—reinforcement learning from verifiable rewards, or RLVR—has become the de facto method for math and code tasks where unit tests or proof verifiers provide ground truth. Papers throughout 2025 showed systematic improvements in pass rates and benchmark scores when models trained with RL feedback loops rather than pure supervised learning.
World models present another thread in the narrative. Google DeepMind's Genie 2, announced December 4, 2024, generates playable, action-controllable virtual environments from text prompts or images—essentially a platform for training agents in rich simulations. Wayve, the UK autonomous driving company that raised a $1.05 billion Series C in May 2024, unveiled its GAIA-2 world model on March 26, 2025, using controllable video generation to stress-test driving policies in simulation before road deployment.
The infrastructure is starting to catch up, too. NVIDIA announced its Grace Blackwell platform in March 2024 for trillion-parameter training, then followed with the Vera Rubin platform in March 2026, explicitly marketed for "agentic AI" and reinforcement learning workloads. Jensen Huang, NVIDIA's CEO, stated in May: "The next frontier of AI is superlearners—systems that learn continuously from experience… we're thrilled to partner with Ineffable to codesign infrastructure for large-scale RL."
It's worth noting what that signals. NVIDIA doesn't casually throw resources at partnerships. This suggests the company sees reinforcement learning at scale as a genuine compute market, not just a research curiosity.
Infrastructure as Insider Knowledge
The NVIDIA partnership, detailed in a joint blog post on May 13, 2026, reveals something about Ineffable's technical strategy. Reinforcement learning data is generated on-the-fly during training, which creates different pressures than pretraining on static datasets—more demand on interconnect, memory bandwidth, serving infrastructure. The collaboration includes co-design work on Grace Blackwell and exploratory engineering on the upcoming Vera Rubin architecture.
This matters for a practical reason. One persistent criticism of experiential learning is efficiency. An LLM can ingest the compressed wisdom of billions of words of human text. A reinforcement learning agent has to discover patterns through trial and error, which historically has been sample-inefficient and expensive. The infrastructure bet suggests Ineffable believes that scale—and purpose-built hardware—can flip the economics.
The UK government's involvement adds a strategic layer. Both the UK Sovereign AI Fund and the British Business Bank co-invested in the round, announced in a government press release on April 27, 2026. The UK has positioned itself as pursuing a "pro-innovation" regulatory stance while backing domestic AI labs with capital, visa support, and compute access through initiatives like the AI Research Resource. Isambard-AI, launched in Bristol in July 2025, provides 5,448 NVIDIA Grace-Hopper chips; companies in the Sovereign AI program can access up to 1 million GPU hours.
It's a play for tech sovereignty, wrapped in venture capital.
The Capital Climate—and Its Contradictions

Ineffable's $1.1 billion seed lands in the middle of an extraordinary AI investment cycle, but one with curious gaps. The Stanford AI Index 2026, published in April, reports that private AI investment more than doubled in 2025, with generative AI capturing nearly half of all private AI funding. Some 239 generative AI companies reached unicorn status by year-end, according to S&P Global Market Intelligence.
Yet enterprise adoption tells a different story. While 88% of surveyed organizations reported using AI in at least one function in 2025, and 70% had deployed generative AI, actual AI agent deployment—the kind of autonomous decision-making that reinforcement learning enables—remains in the low single digits across most business functions. Consumer surplus from generative AI in the United States reached an estimated $172 billion annually by early 2026, up 54% year-over-year, but economic impact remains concentrated in productivity augmentation rather than autonomous capability.
There are also signs of strain. Employment for U.S. software developers aged 22-25 declined nearly 20% from 2024. A third of surveyed organizations expect workforce reductions related to AI in the next year. The capital is flowing toward frontier capabilities, but the path from capability to commercial deployment isn't exactly straightforward.
Silver's pledge to donate any personal proceeds from Ineffable to high-impact charities, disclosed in the April 27 Wired interview, reads differently in that context. This isn't designed as a near-term monetization play. It's something closer to a research institution with venture backing.
What the Bet Actually Means
The investor base offers clues about what Ineffable really is. Strategic players like NVIDIA and Google care about validating alternative training paradigms. Early-stage firms like Sequoia, Lightspeed, and Index are willing to underwrite decade-long timelines. A sovereign fund is backing a domestic champion. The company had somewhere between 11 and 50 employees as of May 2026, according to its LinkedIn profile, with visible hires from DeepMind including Lasse Espeholt, Wojciech Czarnecki, and Junhyuk Oh.
That's a patient capital structure, built for science rather than rapid scaling.
What makes the bet interesting—genuinely interesting, not just expensive—is the bifurcation it implies. Large language models have demonstrated remarkable capabilities, but they're fundamentally interpolative. They recombine and extend patterns found in human-generated data. Reinforcement learning systems have the potential to be extrapolative, discovering solutions or strategies that don't exist anywhere in the training corpus. AlphaGo's Move 37 in game two against Lee Sedol—a placement that no professional human player would seriously consider—remains the canonical example. It was weird. It was also correct.
The regulatory landscape is shifting, too, in ways that might matter. The EU AI Act, enacted in 2024, phases in obligations between 2025 and 2027, with foundation model transparency requirements starting August 2, 2025 and broader provisions kicking in August 2, 2026. Labs training autonomous agents or world models will need to navigate both general-purpose AI rules and downstream high-risk deployment requirements. The UK, by contrast, has positioned itself as offering a lighter-touch, regulator-led approach. That might explain some of the strategic value of a London headquarters.
The Question That Remains

The question facing CTOs and research leaders isn't really whether reinforcement learning will play a role in the next generation of AI systems. It already does—in reasoning models, robotics, autonomous driving. The question is whether Silver's maximal version of the thesis holds: that systems learning purely from experience, without human data as a bootstrap, represent a genuinely different path to general intelligence.
A $1.1 billion seed round suggests that at least some very well-informed people are willing to find out. The timeline Silver has set—long-term science over near-term product—means we probably won't know the answer for years. But the infrastructure is being built. The capital is committed. And the team that demonstrated experiential learning could defeat the world's best human Go player is now aiming at something considerably more ambitious.
Whether human data is "fossil fuel" or foundational knowledge remains an open question. What's clear is that the alternative hypothesis is now funded at scale, with hardware partnerships in place and a team that's already proven the approach can work—at least in bounded domains. The leap from Go to "all knowledge" is, admittedly, not a small one. But then again, ten years ago, the leap from traditional AI to beating Lee Sedol seemed impossible too.
