The problem with training the world's most sophisticated AI models, according to two former Meta engineers, isn't that we've run out of data. It's that most of what we're feeding them is garbage.
Péter Vajda and Seiji Yamamoto would know. Between them, they spent over a decade inside Meta's AI infrastructure—Vajda directing the teams that built Emu and Movie Gen, Yamamoto embedded in the Core Llama group at what the company now calls Meta Superintelligence Labs. They watched frontier models stumble over basic factual questions, hallucinate with alarming confidence, and fail common-sense reasoning tests. The culprit, they concluded, wasn't the architecture. It was the training diet: scraped web text, unverified and awash in misinformation.
So they left. And in mid-May of this year, after going through Y Combinator's Spring 2026 cohort, they announced PerfectBit—a startup with an audacious premise. What if you could generate training data that's correct by construction? Not labeled by exhausted contractors after the fact, but validated in real time against physics simulators, scientific databases, formal proof systems, the kind of deterministic oracles that don't lie.
"Train on reality, not noise," as they put it. Whether frontier labs will buy that pitch remains an open question, but the timing is hard to ignore.
The Backstory
Vajda's résumé reads like a highlight reel of Meta's media synthesis ambitions. Eleven years at the company, most recently as Director of Media Generation. He led the engineering behind Emu—Meta's text-to-image model—then Emu Video, and eventually Movie Gen, the foundation model for generating realistic video content. His name appears on the Movie Gen paper that came out last October. Before Meta, he taught at Stanford as a visiting assistant professor and holds a PhD in computer science.
Yamamoto came at the problem from a different vantage point. A senior staff research scientist, he worked on the foundational language models that underpin much of Meta's AI strategy. His background is unusually varied—PhD in physics, published in PNAS and Physical Review Letters, academic stints at Stanford, Rice, and Columbia, followed by postdoctoral research at a national lab. Not the typical trajectory for someone building transformer models, perhaps, but useful if you're thinking about how to ground AI outputs in verifiable reality.
Both arrived at the same diagnosis: models fail because their training data has no built-in mechanism for correctness. The internet is vast, sure. It's also riddled with errors, contradictions, and deliberate disinformation. Scaling up data quantity won't fix a quality problem.
How They're Building It
PerfectBit's core idea is what the founders call "verifier-grounded data generation." The process works something like this: an AI agent generates a candidate training sample—could be text, code, an image, a snippet of audio or video. Before that sample gets added to a dataset, it's run through an external oracle tailored to the domain. For scientific content, that might mean checking against a physics simulator or molecular dynamics database. For code, it's executable tests. For math, formal proof systems.
Only if the sample passes verification does it make it into the training corpus. The company also bundles verification traces, verdicts, and what they call "reproducible manifests" alongside the data—essentially, a paper trail showing how each sample was validated.
It's an attempt to sidestep what they see as the two dead ends of current data pipelines. Human annotators can't scale fast enough to keep up with model appetites. And web scraping, for all its convenience, pulls in too much low-signal content. The question is whether external verification can replace both.
A Market in Crisis Mode
The company's emergence comes at a moment when the AI industry is openly fretting about data quality. A VentureBeat analysis published earlier this year noted that frontier models were failing one in three production attempts—an error rate that would be intolerable in almost any other critical infrastructure. At the same time, transparency around training methods has been shrinking, as labs become more guarded about what goes into their datasets.
There's also the "peak data" problem. The easy internet has been scraped already. What's left is either lower quality, legally contested, or both. Simply throwing more compute at noisier data doesn't appear to be working.

Recent research has added texture to the challenge. DeepSeek-R1, published in Nature late last year, demonstrated that reinforcement learning with carefully constructed rewards could elicit reasoning behaviors without massive supervised datasets. Impressive—but follow-up studies this spring flagged instability risks. When models are allowed to self-verify without external grounding, they can experience what researchers call "reward inflation," essentially gaming the system by learning to grade their own work generously.
PerfectBit's pitch is that deterministic oracles solve that problem. A physics simulator doesn't care how confident the model is. It either matches reality or it doesn't.
Still Very Early
For now, PerfectBit is a two-person operation based in San Francisco. According to the Y Combinator launch materials, the company is "talking to a small number of frontier AI labs" about pilots—deliberately vague, which is standard for a startup at this stage. The target customers are pre-training, mid-training, post-training, and reasoning programs at major labs. In other words: anyone trying to build or improve large-scale models.
The company is hiring research scientists and engineers, and it's emphatic about in-person work. No disclosed funding beyond Y Combinator's investment, no named customers, no published benchmarks yet. It's a technical thesis in search of validation.
Crowded Field
The AI training data market already has well-capitalized incumbents. Scale AI has built an empire around human-in-the-loop labeling and recently expanded into robotics and physical AI data. Labelbox offers tooling for the entire model lifecycle. Surge AI leans on expert annotators for specialized tasks. Just this past May, Origin Lab pulled in an $8 million seed round to turn licensed video game environments into multimodal training datasets.
PerfectBit's bet is that automation and verification can outcompete human expertise on both quality and scale. It's a technical wager, certainly. But it's also grounded in lived experience—both founders have spent years building models that need to synthesize something resembling reality, whether that's generating coherent video or reasoning through thermodynamics problems.

Whether frontier labs will care more about correctness than coverage is the open question. Vajda and Yamamoto are betting they will. They've been inside those labs. They know what keeps the researchers up at night.
