Kara Peterson and Richard DiBona aren't hiding anything. That's the gambit, anyway.
The two founders of Descrybe, working out of Newton, Massachusetts—a suburb better known for its leafy streets than its AI breakthroughs—just published something unusual in the increasingly crowded legal technology space: a 24-page white paper in March 2026 detailing exactly how their legal reasoning engine performed on a standardized bar exam benchmark. The result, announced March 5, 2026, was a perfect score: 200 out of 200 questions answered correctly.
Not 199. Not "within the margin of error." All of them.
ChatGPT, Claude, and Gemini—the trillion-dollar labs' flagship models—scored lower, sometimes significantly so. And Descrybe is daring anyone with the technical chops to prove them wrong.
Whether this represents a genuine breakthrough or an expertly crafted piece of marketing theater depends largely on what happens next. In an industry where "hallucination-free" claims have collapsed under scrutiny and vendor promises routinely outpace reality, methodology matters. Descrybe seems to understand that. The question is whether anyone will bother checking their work.
The Numbers, and What They Might Mean
DescrybeLM—the company's legal reasoning tool, built specifically for this kind of analysis rather than general conversation—ran through questions from the National Conference of Bar Examiners' Multistate Bar Examination Complete Practice Exam between January and February of this year. Closed-book conditions. Each AI system had to pick an answer and explain why.
ChatGPT 5.2 managed 187 correct answers, a 93.5% success rate. Gemini 3 Pro hit 184 (92.0%). Claude Opus 4.5 came in at 177 (88.5%). Respectable scores, certainly. Just not perfect.
Descrybe also evaluated what it calls "reasoning quality"—how well each model explained its thinking, not just whether it picked the right letter. Using a rubric scored by an AI judge (GPT-5.2 in extra high reasoning mode, for whatever that's worth), DescrybeLM notched a 99.70% quality rating. ChatGPT followed at 93.41%, Gemini at 91.45%, Claude at 89.03%.
Perhaps more revealing: Of the 52 questions the three general-purpose models got wrong, 49 were marked "confidently wrong." The errors didn't overlap much, suggesting each system has learned to fail in its own particular way.
That detail—buried in the methodology section—hints at the deeper problem plaguing legal AI. It's not just that these systems make mistakes. It's that they make them with the serene assurance of a partner billing $800 an hour.
The Fine Print Descrybe Actually Published
Here's where it gets interesting. The white paper doesn't try to wave away the limitations.
The MBE question set is commercially available, which means—as the white paper itself acknowledges—Descrybe can't prove their model, or anyone else's, didn't see these questions during training. Each question got a single run, not multiple attempts with cherry-picked results. The evaluation was conducted by the Descrybe team itself, using an AI judge rather than human experts. And crucially, this benchmark tests only multiple-choice legal reasoning. It says nothing about drafting a motion, pulling accurate citations, or navigating jurisdiction-specific research.
Bob Ambrogi, who runs LawSites and has been covering legal tech since before "legal tech" was a venture capital category, flagged these caveats when he wrote about the launch. He hadn't tested the product himself at that point—a refreshingly honest admission in a space where breathless vendor coverage often masquerades as journalism.
This level of disclosure is either naive or shrewd. In a market where competitors tout vague claims about "enterprise-grade accuracy" and "state-of-the-art performance," publishing your scoring logs feels almost provocative.
Purpose-Built Versus the Everything Machines

DescrybeLM doesn't work alone. It's designed to pair with Descrybe's Legal Research Toolkit, which launched last June. The workflow goes like this: use the LRT to find governing law, then feed those results to DescrybeLM to apply the law to your facts.
DiBona told LawSites the system was trained on a curated corpus of more than 100 million structured legal records—over 100 billion tokens of legal-specific material. The company won't disclose its base model architecture or how its retrieval system works, citing competitive reasons. Fair enough. It doesn't scrape the public web by default, a design choice that probably prevents the model from confidently citing cases that don't exist.
"The question we wanted to answer," Peterson said in the announcement, "was whether purpose-built legal AI could meaningfully outperform general-purpose models on core legal reasoning tasks."
It's a pointed question. ChatGPT can write you a poem, plan your vacation, and debug your Python script. DescrybeLM does one thing. The bet is that specialization matters.
Whether that bet pays off depends on what lawyers actually need—and what they're willing to trust.
The Access Play
Descrybe launched publicly in the summer of 2023 with free case law search, offering plain-English summaries before that became table stakes. By last August, the platform had been added to the National Society for Legal Technology curriculum, which reaches more than 350 universities and law schools across 11 countries.
The Legal Research Toolkit pricing, as of August, ran $10 monthly for non-commercial users and $20 for commercial use—a fraction of what traditional legal research platforms charge. The company hasn't published separate pricing for DescrybeLM, though one imagines it won't involve the enterprise sales cycles that define deals with Harvey AI or Thomson Reuters.
Peterson made the ABA's Women of Legal Tech 2024 honors list. The company has collected awards from Anthem and Webby for responsible AI use. LinkedIn pegs the team size between two and ten employees. Available information suggests they remain bootstrapped, though in this corner of the market, funding announcements sometimes lag reality by months.
For a team that small to claim performance advantages over models built by organizations spending hundreds of millions on training runs—well, it invites skepticism. Which may be exactly the point.
The Verification Problem Nobody's Solved

The white paper emphasizes structured reasoning: clear rule statements, explicit application to facts, step-by-step logic chains. The idea is to make outputs easier to verify, catching errors before they become malpractice claims.
That works beautifully on a multiple-choice exam where every question has a clean right answer and a defined scope. Real legal work is messier. Documents run long. Facts conflict. Judges interpret statutes in ways that surprise everyone, including the drafters.
Whether DescrybeLM's approach scales to those conditions remains an open question. Multiple-choice benchmarks make for clean headlines. They're less predictive of how a tool performs when a solo practitioner is trying to draft a response to a motion for summary judgment at 11 p.m. on a Thursday.
Who Else Is in This Fight
The competitive landscape has gotten crowded, fast.
Harvey AI claims over 1,000 clients spanning 60 countries as of January and has inked content partnerships with Wolters Kluwer. Thomson Reuters rolled out guided workflows for CoCounsel last July and made its AI tools available to law schools by September. vLex launched Vincent Studio, pitched as a "maker space for legal AI," in January.
LexisNexis made waves when it launched Lexis+ AI in 2023, touting "hallucination-free" citations. An academic study published the following May found hallucination rates between 17% and 33% on tested tasks for major legal research tools. Turns out "hallucination-free" is harder to deliver than it is to claim.
That's the context Descrybe is stepping into. Every vendor promises accuracy. Every press release uses words like "revolutionary" and "transformative." Publishing full methodology and inviting replication might become a competitive advantage. Or it might just become table stakes if customers and regulators start demanding proof.
As of mid-March, no independent replication of the DescrybeLM benchmark has appeared in academic journals or trade press. That could change quickly. Or it could take months. The legal AI community is small enough that word travels fast when someone's results don't hold up.
What Happens Now

Perfect scores grab attention. Transparency builds trust—or at least, it's supposed to.
Descrybe's bet appears to be that in a market saturated with unverifiable accuracy claims, showing your work matters more than the headline number. The white paper is public. The scoring logs are available. Anyone with access to the MBE practice exam and the technical resources can try to replicate what a two-person team in suburban Boston says they've accomplished.
For a startup competing against billion-dollar labs and entrenched incumbents with multi-decade customer relationships, radical openness is either supreme confidence or smart positioning. Possibly both.
The harder test comes next: whether actual lawyers, working on actual cases with actual consequences, find the tool useful enough to pay for it. Benchmarks make for compelling narratives. Billable hours are a different story.
And if Descrybe's results hold up under independent scrutiny, the industry might have to confront an uncomfortable question: whether the massive scale and general-purpose capabilities that define frontier AI development are necessary advantages—or expensive distractions—when it comes to specialized professional work.
For now, Peterson and DiBona have put their methodology on the record. The rest is up to replication, real-world use, and the slow accumulation of evidence that separates genuine breakthroughs from well-executed demos.
