There's a certain irony in the fact that Peter Vajda and Seiji Yamamoto spent years inside Meta's AI labs helping build systems that generate increasingly plausible text—only to leave and start a company premised on the idea that plausibility isn't good enough.
PerfectBit, the duo's YC-backed venture, launched this year with a thesis that sounds almost quaint in an era of trillion-parameter models: What if the solution to AI hallucinations isn't more compute, but mathematically verifiable training data? Data proven correct against physics simulators. Cross-referenced with authoritative scientific databases. Grounded in formal proof systems. No probabilistic handwaving, as Vajda puts it. Just facts that can be checked.
It's a pitch gaining traction at a peculiar moment. The AI industry—flush with capital and confidence—is simultaneously confronting an uncomfortable reality. Bigger models haven't eliminated hallucinations. In some ways, they've made them more sophisticated, harder to catch, more dangerous in high-stakes domains.
When Scale Stopped Being the Answer
The numbers tell part of the story. Fortune Business Insights pegged the AI training dataset market at $3.59 billion in 2025, with projections pointing toward $4.44 billion in 2026 and $23.18 billion by 2034. That's a compound annual growth rate pushing 23 percent—the kind of trajectory that reflects not just expansion, but a fundamental reassessment of what matters.
Top-tier models have improved, no question. Vendor research from Presenc AI claims summarization hallucination rates dropped to approximately 1.0-2.5 percent in 2026, down from the 3-8 percent range in 2023. Progress, certainly. But those aggregate numbers conceal stubborn failures in specific domains.
A study published in March found that LLaMA-70B-Instruct hallucinated in nearly one out of every five answers on certain medical question-answering tasks—19.7 percent, to be precise. Another benchmark released in February, dubbed HalluHard, demonstrated that even strong models paired with web search still fabricate information roughly 30 percent of the time during multi-turn conversations.
The challenge isn't purely technical. It's structural, maybe cultural. Scale AI—the incumbent giant in data labeling—repositioned its entire brand messaging around "reliability" in May. CEO Jason Droege told Axios the industry needed to close what he termed the "Reliability Race" gap. A month earlier, Gartner predicted that by 2028, explainable AI would drive LLM observability investments to half of all secure GenAI deployments. Gartner also estimated the generative AI model market itself cleared $25 billion last year.
That kind of money attracts attention. And attention reveals cracks.
Three Converging Forces

Vajda and Yamamoto—both veterans of Meta's Superintelligence and Media Generation teams—argue that three trends are converging to make verified training data less of a theoretical luxury and more of a practical necessity.
First: reinforcement learning from verified rewards, or RLVR in the industry shorthand. What began as a research curiosity has become production technique. OpenAI's o1 models, demonstrated in late 2024, showed that chain-of-thought reasoning paired with reinforcement learning could unlock meaningful performance gains. DeepSeek-R1, published in Nature last September, proved that pure RL with verifiable rewards could elicit complex reasoning behaviors without requiring massive supervised datasets. By the start of this year, RLVR had become standard practice at frontier labs working on reasoning models.
PerfectBit frames its approach as pushing RLVR upstream. Rather than verifying outputs after training, the company generates training data that's "correct by construction"—using what it describes as an "agent-verifier loop" grounded in physics simulators, formal proof systems, or curated scientific databases. The claim is that this reduces the compute and iteration cycles needed to achieve reliable behavior. Whether that holds at production scale remains to be seen.
Second: Cloud providers have essentially institutionalized grounding and verification. Google's Vertex AI now offers grounding through Google Search and Vertex AI Search integrations by default. AWS Bedrock shipped contextual grounding checks in 2025 and added cross-account safeguards this April. Azure AI Content Safety, updated at the end of January, includes "groundedness detection" as a core feature. What qualified as experimental middleware two years ago is now table stakes in managed AI services.
Third, and perhaps most consequential: regulatory pressure is mounting. The EU AI Act's Article 10 mandates data quality and governance requirements for high-risk AI systems, with key enforcement milestones hitting August 2. NIST's AI Risk Management Framework added a concept note for trustworthy AI in critical infrastructure on April 7. The International AI Safety Report published in February singled out reliability and data provenance as central safety concerns.
Meanwhile, a March study in the Harvard Data Science Review warned that human evaluator bias can inadvertently reward hallucinations when accuracy is the only optimization metric—a finding that complicates even carefully designed feedback loops.
The regulatory message is unambiguous: "trust me" isn't going to cut it anymore.
The Verification Ecosystem Takes Shape

PerfectBit isn't working in a vacuum, though its focus on generating training data from oracles is relatively novel. Most research and commercial work in this space centers on detection or mitigation after the fact.
Microsoft Research's VeriTrail system, presented at ICLR earlier this year, demonstrates closed-domain hallucination detection with traceability across multi-step workflows. Apple published research in March on hallucination span detection—identifying which specific tokens in a response are factually incorrect. Amazon Science showed how attention maps could detect hallucinations in speech-based language models at inference time.
The academic output has been prolific, if occasionally sprawling. Recent papers explore everything from reasoning-subspace projection (a technique called HARP, published in April) to chance-constrained inference that controls failure rates during decoding. A financial domain pipeline called FinGround, published in late April, achieved a 68 percent reduction in hallucinations versus the strongest baseline under controlled retrieval conditions—rising to 78 percent versus GPT-4o overall.
Real-world deployment patterns vary widely. A clinical study published on medRxiv in February showed that naive retrieval-augmented generation systems could actually increase hallucinations to 43.6 percent in synthetic clinical vignettes. But when researchers introduced "structured patient artifacts"—essentially verified data representations—the rate dropped to 8.4 percent. Braincuber, a RAG vendor, claimed in a March case study that targeted mitigations brought a deployment from baseline to 6.4 percent hallucination rates in two weeks. IrisAgent, which builds customer support AI, reports hallucination rates below 5 percent across enterprise deployments with Dropbox, Zuora, and Teachmint.
The M&A activity tells its own story. Handshake acquired Cleanlab—a startup focused on detecting label and data quality issues—in late January. TechCrunch reported at the time that Cleanlab had attracted interest from "multiple others," underscoring the strategic value of quality-focused data infrastructure.
Scale AI's pivot is perhaps most revealing. The company that built its business on human annotation is now framing its Generative AI Data Engine around evals, safety checks, and RLHF loops. Alexandr Wang, Scale's founder, told Time in a June interview last year that data is "the key bottleneck" for AI progress. By May, Droege was calling reliability the next competitive frontier.
Early Days, Big Questions

PerfectBit's launch arrives at what may be an inflection point—or may just look like one from where we're standing now. The company hasn't disclosed funding beyond its YC backing, and its team of two suggests this is very early days. The YC launch post mentions pilots with "frontier labs," but there are no public benchmarks or customer case studies yet.
That's not unusual for a company founded this year. But it means PerfectBit's claims—however compelling—remain to be proven at scale.
What's less debatable is that the technical trajectory supports their broader thesis. RLVR is becoming standard practice. Grounding and verification are now bundled features in every major cloud AI platform. New benchmarks—FACTS Suite, HalluHard, DO-Bench, FinGround—keep emerging, each exposing residual hallucination problems that prompt engineering alone can't fix.
Regulatory timelines add urgency, whether AI labs like it or not. The EU AI Act's August milestone means high-risk AI systems will soon require auditable data governance. NIST's evolving framework signals similar requirements coming for critical infrastructure applications in the U.S. Content licensing deals—News Corp with Meta in March, Shutterstock's expanded training datasets later that month—suggest that training data provenance is shifting from afterthought to core intellectual property.
Market forecasts support sustained growth. Precedence Research estimates the AI training dataset market at $3.35 billion in 2025, projected to reach approximately $14.94 billion by 2035. IDC argued in April that by 2029, 60 percent of enterprise data platforms will need to unify transactional and analytical workloads to support agentic AI—a shift that hinges on data freshness and verifiability.
PerfectBit's "correct by construction" approach—physics simulators, formal proofs, oracle-grounded loops—remains unproven at production scale. The engineering overhead is real. The compute costs could be prohibitive. But the company is riding three tailwinds: technical momentum around RLVR, institutional adoption of verification layers, and regulatory pressure that makes data provenance a compliance issue rather than a nice-to-have.
Whether verified training data becomes the norm or remains a niche for specialized domains will depend on how well startups like PerfectBit can deliver results that justify the additional complexity. For now, the industry seems to be reaching a consensus: bigger models won't fix this.
Better data might.
