Peter Vajda and Seiji Yamamoto didn't leave Meta's AI labs to build another chatbot. Vajda had directed Media Generation at Meta. Yamamoto led engineers in the Core Llama group. Both had spent enough time inside frontier model development to know where the cracks were showing—and it wasn't in the architecture.
The problem, as they saw it, was the training data itself. Fundamentally unverifiable, they'd argue. Every example fed into today's models arrives without any external check on whether it's actually correct. Their answer is PerfectBit, a startup backed by Y Combinator's Spring 2026 batch, built around what they call "verifier-grounded" training data. The pitch: every sample gets validated against an external oracle—a formal proof system, a physics simulator, an authoritative database—before it ever reaches a model.
It's a narrow wedge into a sprawling problem. Stanford's AI Index 2026 tallied hallucination rates between 22% and 94% across 26 leading models on factuality benchmarks. Meanwhile, Gartner projected in May 2026 that AI spending will hit $2.59 trillion this year, up 47% from last year, even as enterprises struggle to extract proportional value from those investments. The gap between expenditure and reliability has opened a market, perhaps, for infrastructure that treats correctness as a pre-training problem rather than a post-deployment firefight.
The Hallucination Crisis, Still Unresolved
The industry has known about the hallucination problem for years now. Awareness hasn't solved it. OpenAI's 2025 analysis traced the issue to structural quirks in next-word prediction training. More troubling, a Nature paper published in May 2026 suggested that the very way we evaluate language models might incentivize fabrication—a kind of theoretical pressure toward hallucination even with perfect input data. The authors argued current evaluation frameworks themselves might need rethinking.
Real-world fallout has accumulated. Medical researchers documented LLMs vulnerable to adversarial hallucinations in clinical decision support. Legal scholars at Princeton's Center for Information Technology Policy built cite-checking benchmarks after lawyers started encountering fabricated citations in court filings. A separate Nature study on medical imaging found that just 10% label noise degraded retinal diagnosis accuracy by 3.6 to 4.5 percentage points.
The verification tools market has responded with structured deployments. Google rolled out "Grounding with Google Search" across Gemini API, AI Studio, and Vertex AI throughout 2024 and 2025, embedding live retrieval to curb hallucinations at inference time. Vectara launched an updated Hallucination Leaderboard in late 2025, alongside HHEM 2.x, a factual consistency scoring model for retrieval-augmented generation. Scale AI—which raised $1 billion at a $13.8 billion valuation back in May 2024—has built much of its data engine around evaluation infrastructure.
But most of these efforts target the generation phase. PerfectBit's founders are betting on a different intervention point: the data layer itself, before training even begins.
Why Now? Three Converging Pressures

Regulatory winds are shifting. The EU AI Act entered into force on August 1, 2024, requiring high-risk systems to ensure "high-quality" training, validation, and test data under Article 10. NIST's AI Risk Management Framework and its July 2024 Generative AI Profile explicitly name "confabulation (hallucination)" as a risk demanding data quality controls and provenance tracking. Compliance may eventually mean auditable verification, not just vague assurances.
Then there's the research validation. Microsoft Research published VeruSyn in February 2026, describing a pipeline that generated 6.9 million verified Rust programs—each with formal specifications and proofs attached. Fine-tuning on that corpus improved model performance against Claude Sonnet 4.5 on proof tasks. The VERINA benchmark revealed stark gaps: OpenAI's o3 model hit 72.6% code correctness but only 4.9% proof success in one-shot tests. Writing code that runs is one thing. Writing code you can prove correct is another.
The concept PerfectBit calls "verifier-grounded" or "correct-by-construction" data shows up across recent academic literature under various labels: "verifier-grounded learning," "execution-based supervision," "formal-verification-based data synthesis." The terminology varies, but the core idea doesn't. Establish correctness before use via external oracles—whether that's a theorem prover, a test suite, a physics engine, or a deterministic procedure.
Third factor: market sizing suggests infrastructure plays still have room. IDC projects AI infrastructure spending at $487 billion in 2026, up roughly 53% year-over-year; Q4 2025 alone approached $90 billion. Grand View Research pegged the data labeling market at $1.03 billion in 2023, with related "data collection and labeling" forecast to reach $17.1 billion by 2030. Those figures predate widespread verification mandates. If regulatory or liability pressures mount, the economics could tilt sharply toward correctness guarantees.
Verification in Practice: From Robots to Theorems

Verification approaches are spreading across domains, though not always under that name. In robotics, NVIDIA's Isaac Sim and Cosmos ecosystem functions as a data factory where simulation generates pre-verified trajectories. Isaac Sim 5.0 and Isaac Lab 2.2 were open-sourced over the past year or so, and robotics companies use them to train policies against physics-accurate simulations before real-world deployment. Wayve's GAIA-2 generative world models produce safety-critical scenarios for autonomous vehicles—essentially "pre-verified" scenario generation for edge cases too dangerous to collect in reality.
Mathematical reasoning has its own verified corpora. Theorem-proving datasets like LeanDojo, MiniF2F, and Nemotron-Math-Proofs pair problems with formal proofs that can be mechanically checked. FoVer, presented at ACL 2026 Findings, demonstrated verifier-guided process reward model data synthesis using the Z3 formal checker. Each reasoning step gets verified before becoming training data.
Code generation saw execution-based verification loops as early as CodeT in 2022, with subsequent systems like DOCE and CodeBenchGen using tests and execution feedback to filter or reward correct programs. But there's a meaningful distinction between "runs correctly" and "proven correct." PerfectBit emphasizes the latter—using formal proof systems alongside simulators and databases.
Enterprise efforts have largely focused on inference-time grounding rather than training data quality. Google's Vertex AI Grounding, expanded across the EU in 2025-2026, retrieves from Google Search during generation. Vectara's Tool Validator and Guardian Agent, launched in December 2025, provide step-level audit trails and factual consistency checks for agentic AI systems. Useful, certainly. But these approaches complement rather than replace training data quality improvements.
A Two-Person Team Betting on Oracles
PerfectBit operates as a two-person team out of San Francisco, according to its Y Combinator listing—though team size may have evolved since that posting. The company's website lays out four convictions: data quality is paramount; human annotators don't scale for technical domains; verifiers can serve as oracles; and enhanced natural language can "project" physics, biology, and logic into text. The tagline reads "Train on reality, not noise."
The target customers are specific enough to suggest focus, broad enough to suggest ambition. Frontier labs running pre-training or post-training runs that need correctness-heavy data supplements. Robotics companies building policies from simulation-verified trajectories. AI-for-science teams requiring theorem or proof-checked corpora. Regulated verticals—healthcare, legal, government—with auditability needs baked into procurement.
Market positioning matters here, because the data infrastructure landscape is crowded. Scale AI approached it as a horizontal platform. Cleanlab focuses on label error detection and correction. Snorkel emphasizes programmatic labeling. Synthetic data vendors like Gretel, Mostly AI, and Hazy target privacy-preserving generation. PerfectBit's differentiator is the oracle requirement: every sample verified against an external ground truth before inclusion. Whether that's enough differentiation to command premium pricing is an open question.
The economics hinge partly on how hallucination costs compound over time. Stanford's AI Index 2026 noted that model transparency scores actually declined in 2025 versus 2024, even as capabilities improved. If regulatory scrutiny intensifies or liability risk crystallizes—say, a few high-profile lawsuits over fabricated outputs—the calculus shifts toward correctness guarantees. Gartner's May 2026 forecast observed that while AI spending accelerates, enterprises haven't yet captured proportional value. That's the kind of dynamic that historically drives infrastructure consolidation toward reliability over raw capability.
What Happens Next

Several threads suggest verification infrastructure might outlast the current hallucination panic. Agentic AI systems amplify error propagation in ways that simple question-answering doesn't. When models call tools, chain reasoning steps, and operate semi-autonomously, incorrect intermediate outputs cascade. Gartner and Forrester coverage in 2026 points to rapid agentic AI scaling but high project failure rates absent guardrails. Runtime verification helps, but training on verified data reduces the base error rate—fewer mistakes to catch downstream.
Simulation fidelity keeps improving, too. NVIDIA's investments in Isaac Sim and physics engines make synthetic-but-verified data increasingly viable for robotics and embodied AI. As simulators approach reality, the data they generate becomes training-worthy rather than merely evaluative. This could flip the traditional hierarchy where real-world data is gold standard and synthetic is supplementary.
Formal methods are also scaling beyond academic proofs. VeruSyn's 6.9 million verified programs represent orders of magnitude more formal data than existed a few years ago. Microsoft, Amazon (via Dafny), and other vendors are investing in verification tooling. If formal verification costs drop while hallucination costs rise, the crossover point favors proof-checked data even for general-purpose models. Maybe.
The broader industry bet—shared by PerfectBit and others—is that data curation matters as much as compute and architecture. Andrew Ng's data-centric AI push in 2021-2022 argued this conceptually; the hallucination crisis has made it economic. Companies building verification infrastructure today are wagering that correctness becomes a purchasable property, not just an aspirational goal.
For PerfectBit specifically, the challenge is demonstrating that oracle-verified supplements materially improve model behavior at frontier scale. The research precedent exists: VeruSyn showed gains, VERINA exposed gaps, execution-based loops showed filtering value. But moving from research contribution to enterprise revenue means integrating with existing training pipelines at labs spending tens or hundreds of millions per run.
It's one thing to prove a concept in a controlled academic setting. It's another to convince a frontier lab to route a meaningful portion of its training budget through a two-person startup's verification pipeline. The market will decide whether "correct by construction" becomes essential infrastructure or remains an elegant idea with a San Francisco address and a batch number. The hallucination problem isn't going away. Whether PerfectBit's solution is the one that sticks—well, that's the bet Vajda and Yamamoto are making.
