In the parlance of artificial intelligence research, there's a problem everyone knows but few want to say out loud: AI agents fail. A lot.
NeoCognition, a Palo Alto startup emerging from Ohio State University's AI labs, thinks it has an answer. On April 21, the company announced a $40 million seed round—among the largest in recent memory for a company barely nine months old—to build what it calls "learning agents" that get smarter on the job rather than remaining static generalists.
The round, which was oversubscribed, drew co-leads Walden Catalyst Ventures and Cambium Capital, with Vista Equity Partners joining alongside a roster of Silicon Valley heavyweights: Intel CEO Lip-Bu Tan, Databricks co-founder Ion Stoica, and AI researchers Dawn Song, Ruslan Salakhutdinov, and Luke Zettlemoyer. Also participating: A&E Investments, Salience Capital Partners, Nepenthe Capital, and Frontiers Capital.
It's the kind of investor lineup that suggests something beyond typical seed-stage enthusiasm. Perhaps it's the market timing—agents are having a moment, after all—or maybe it's the team's research pedigree. More likely, it's the uncomfortable truth the founders are willing to articulate: current AI agents are unreliable in ways that make enterprise adoption a minefield.
The 50% Problem
NeoCognition's pitch begins with what CEO Yu Su describes as the "critical reliability gap" in agent deployments. According to the company's founders, most agents today operate as jacks-of-all-trades, achieving task success rates hovering around 50% in real-world production environments. Flip a coin, in other words.
Recent benchmarks tell a slightly rosier story. When OSWorld was introduced in 2024, the best models achieved roughly 12% success compared to human performance at 72%. By early 2026, according to Stanford HAI's AI Index report, leading agents had improved substantially on similar benchmarks. That's improvement, certainly. But it also means significant gaps remain—a ratio few enterprises would tolerate in mission-critical workflows.
And benchmarks, as academic research published late last year suggests, often miss the point entirely. They don't measure cost. They don't track consistency across multiple runs. They ignore what happens when these systems encounter the messy, unpredictable edges of real operational environments.
NeoCognition's thesis: stop building generalists. Build specialists instead.
Learning on the Job
The company's approach centers on what it calls "world models"—agents that develop deep expertise in the specific micro-environments where they're deployed, learning continuously from experience rather than relying solely on pre-training. Think of it less as a Swiss Army knife, more as a tool honed for one particular task and refined through repetition.
It's an old idea in some ways—specialization has always beaten generalization when stakes are high. But applying it to AI agents at scale is uncharted territory.
The founding team isn't new to this terrain, at least academically. Yu Su, an associate professor in Ohio State's Computer Science and Engineering department, received a Sloan Research Fellowship in February 2025. His co-founders, Xiang Deng and Yu Gu, both earned PhDs from OSU focused on agents and open-world AI systems. Together, they've published work that's become infrastructure for the field: Mind2Web, a web agent benchmark presented at NeurIPS 2023, and follow-on research including SeeAct, Online-Mind2Web, and multimodal reasoning benchmarks MMMU and MMMU-Pro.
At announcement, NeoCognition employed around 15 people, most holding PhDs, according to TechCrunch's coverage.
The Vista Angle

Vista Equity Partners' involvement may be the most telling signal in the cap table. The private equity giant manages north of $100 billion in assets, much of it concentrated in enterprise software companies—exactly the customers NeoCognition is targeting.
Su noted to TechCrunch that Vista's participation could unlock channel access for enterprise deployments, a polite way of saying: here's a built-in testing ground across dozens of portfolio companies.
"The novel learning mechanism enables rapid specialization," said Landon Downs, co-founder and managing partner at Cambium Capital. Ion Stoica, investing as an angel, framed it more bluntly: "Expert-level intelligence" is what makes these systems reliable and cost-effective in high-stakes applications, he argued.
Whether that specialization-through-experience thesis holds up in production remains an open question. But the investor syndicate clearly believes domain expertise will trump general capability for enterprise use cases.
Unusually Large, Unusually Quiet
A $40 million seed is unusual by any standard. In 2024, Carta data pegged the median seed round at $3.5 million; the NVCA's 2026 Yearbook cited median seed pre-money valuations of $16 million as of late last year. For context, Mistral AI's €105 million seed in 2023 made headlines as Europe's largest at the time.
NeoCognition incorporated in Delaware on July 22, 2025, per California Secretary of State records, and spent roughly nine months in stealth before going public. As of the announcement, the company's website—neocognition.io—remained mostly a placeholder, offering little beyond the basics.
That extended quiet period, combined with the heavyweight investor roster, suggests a deliberate strategy: build in private, emerge with credibility and capital to move fast.
What Comes Next

The funding positions NeoCognition to scale its team and deploy learning-based agents into enterprise pilots. Vista's portfolio companies seem like natural early customers, offering controlled environments to prove the technology works under real operational pressure.
But there's a gap between research-grade prototypes and production-hardened systems that handle actual enterprise workflows. Academic benchmarks are one thing; keeping a Fortune 500 company's critical processes running without unexpected failures is quite another.
The bet NeoCognition is making—and that $40 million is backing—is that agents can close the reliability gap not by being better generalists, but by becoming true experts. If they're right, it could reshape how enterprises think about deploying AI in high-stakes environments.
If they're wrong, well. At least the team has runway to find out.
