The white paper dropped on a Thursday morning, unadorned and unexpectedly bold. A husband-and-wife team working out of Newton, Massachusetts claimed their legal AI had just aced a standardized bar exam benchmark—200 questions, zero mistakes, during testing conducted in January and February 2026. Meanwhile, ChatGPT, Claude, and Gemini each stumbled on somewhere between 13 and 23 of the same questions.
That was March 5. The methodology logs went live the same day, scoring rubrics and all, with an open invitation: replicate our results. See if we're right.
The claim, if it survives independent scrutiny, arrives at a useful moment. Law firms and legal tech buyers have been wrestling with a persistent question—whether general-purpose AI systems can truly match the precision of models purpose-built for legal reasoning, or whether the whole vertical-AI thesis is just clever positioning. Descrybe's announcement offers one potential data point. Perhaps more than the founders expected, it also offers a case study in how AI benchmarks are being weaponized in competitive markets where trust is the product.
The Numbers
Descrybe's model—DescrybeLM—went 200 for 200 on the National Conference of Bar Examiners' Multistate Bar Examination Complete Practice Exam during testing in January and February. ChatGPT 5.2 managed 187 correct answers (93.5%). Gemini 3 Pro landed at 184 (92%). Claude Opus 4.5 came in at 177 (88.5%).
Same prompt for each system. No internet access. Single run per question. The company published everything: the rubric, the standardized prompt, per-output scoring logs. They even disclosed the judge model—GPT-5.2 extra high reasoning, the same provider family as one of the systems being evaluated. That detail, tucked into the methodology notes alongside other caveats, is the kind of thing independent researchers will scrutinize.
Descrybe went further, scoring reasoning quality on a separate rubric. DescrybeLM hit 99.70%. ChatGPT scored 93.41%, Gemini 91.45%, Claude 89.03%.
These are vendor-supplied figures. No independent verification yet.
The Team Behind the Claim
Descrybe is Kara Peterson, the CEO, and Richard DiBona, the CTO. They're married. They founded the company in 2023. Before this week's launch, they'd been running a free legal research platform and, as of June 2025, a paid Legal Research Toolkit priced at $10 monthly for non-commercial users and $20 for commercial accounts. (Those are 2025 rates; pricing for DescrybeLM wasn't disclosed in the launch materials.)
"Built differently," DiBona said in coverage of the release. "From the foundation up, specifically for legal reasoning, on more than 100 million structured records."
The company says it processed north of 100 billion tokens preparing its curated legal corpus—primary law, meticulously structured—a figure they disclosed in a LinkedIn post three months before launch. DescrybeLM is designed to work in tandem with the existing toolkit: the Legal Research Toolkit retrieves governing law, then DescrybeLM applies reasoning and drafting against the facts. The new product went live at descrybe.com. The legacy descrybe.ai domain now redirects there, mentioning a free trial.
When Wrong, Confidently So

Across the 52 incorrect outputs from the three general-purpose models, 49 were what Descrybe labeled "confidently wrong." Assertive phrasing. Wrong legal rule, or the right rule misapplied, stated as though certain.
"When these systems were wrong, they were confidently wrong," the white paper notes—a phrasing that captures what many practicing attorneys have started noticing about generative AI in legal contexts. The technology doesn't hedge. It doesn't say "I'm not sure." It just... answers.
The failure analysis also surfaced something else: the errors didn't overlap much between models. If you're cross-checking outputs from ChatGPT and Claude to catch mistakes—a strategy some firms have adopted—Descrybe argues the data suggests that won't reliably work. The models fail on different questions.
Whether that pattern holds outside multiple-choice exams is another matter entirely. The MBE tests reasoning on general legal principles in a controlled format. It doesn't measure citation accuracy in retrieval workflows, jurisdiction-specific edge cases, or drafting quality when facts are ambiguous and the law is unsettled. Real legal work is messier.
The Disclosure List
To their credit, Peterson and DiBona front-loaded the limitations. This was a vendor-conducted benchmark, not independently verified. Single-run evaluation, so variance across attempts isn't captured. The study can't rule out training-data contamination—MBE questions may have appeared in any of the systems' training sets, including DescrybeLM's own.
The team did fine-tune DescrybeLM on a separate NCBE product before the evaluation (not the test set), and no system got access to answer keys during testing. Still, multiple-choice questions are a narrow window into legal AI utility. They don't capture the full range of tasks lawyers actually delegate to these tools—the kind of work done under deadline pressure with incomplete information and high professional liability.
Bob Ambrogi, who covers legal technology at LawSites, walked through the results on March 5 and explicitly noted the product hadn't faced independent testing yet. The company's response: here's the methodology, replicate it yourselves.
The Competitive Landscape

DescrybeLM enters a legal AI field that's gotten noticeably more crowded. LexisNexis rebranded as "Lexis+ with Protégé." Harvey, backed by significant venture funding, runs its own evaluation suite called "BigLaw Bench." Alexi launched a legal reasoning tool in January of last year. Academic groups and industry consortia keep publishing new benchmarks—a signal that the evaluation methodology itself remains contested territory.
Descrybe has some academic traction. The National Society for Legal Technology includes the company's tools in its curriculum, which reaches more than 350 universities and law programs across 11 countries, at least as of last August. The company reported over 50,000 monthly users at that time, though that figure predates the DescrybeLM launch by more than half a year, and user numbers in this space can be... elastic.
The Thesis
The vertical-AI argument is straightforward enough: a model purpose-built on legal data and trained specifically for legal reasoning should outperform a general-purpose system asked to moonlight as a lawyer. The benchmark suggests that might be true for standardized multiple-choice questions.
Whether it holds for the messy, high-stakes work of litigation research and brief writing—that's the test that actually matters. And it's one that law firms will run themselves, vendor benchmarks be damned. The profession has learned, sometimes expensively, that confident-sounding outputs and perfect scores on academic tests don't always translate to reliable performance when a partner's name is on the brief and a client's future hangs in the balance.
For now, Descrybe has published a claim—one that remains without independent verification—and opened the methodology for inspection. The replication attempts, when they come, will tell us whether a two-person startup in Massachusetts has genuinely built something different—or just run a very good benchmark.
