Descrybe, a legal AI company most lawyers have never heard of, just made a claim that sounds almost too clean: perfect score on a 200-question bar exam benchmark, while ChatGPT, Claude, and Gemini each stumbled on double-digit questions.
The usual skepticism applies—Descrybe ran the study itself, after all. But the Newton, Massachusetts outfit did something uncommon for a vendor benchmark: it published a 24-page white paper spelling out its testing methodology, scoring rubrics, and an invitation for anyone to try replicating the results. Whether anyone actually will remains an open question.
The announcement arrived with the launch of DescrybeLM, a legal reasoning engine designed to complement the company's existing Legal Research Toolkit. Co-founder and CEO Kara Peterson framed the release around transparency rather than triumphalism. "We tested ourselves, published our methodology, and invite anyone to replicate it," she said.
For a startup with fewer than ten employees and no disclosed venture backing, that's either confidence or calculated risk. Possibly both.
What the Numbers Show (and Hide)
The white paper—titled "Beyond Confidently Wrong," which tells you something about Descrybe's positioning—details testing conducted between January and February on the NCBE MBE Complete Practice Exam. That's a standardized 200-question multiple-choice test spanning core legal subjects, the kind law students lose sleep over.
Under closed-book conditions (no web access, no external tools), DescrybeLM answered every question correctly. ChatGPT 5.2 scored 93.5%. Claude Opus 4.5 hit 88.5%. Gemini 3 Pro reached 92.0%.
Raw accuracy tells only part of the story, though. Descrybe also scored reasoning quality using a rubric applied by a GPT-5.2 family judge model—a choice that introduces its own complications, which we'll get to. DescrybeLM's explanations scored 99.70% on that rubric. The general-purpose models ranged from 89.03% to 93.41%.
More revealing, perhaps: among the 52 incorrect answers produced by ChatGPT, Claude, and Gemini combined, 49 were flagged as "confidently wrong." That's AI-speak for asserting incorrect legal standards or misapplying the right ones without any hedging language. Claude Opus 4.5 and Gemini 3 Pro also drew overconfidence flags on some correct outputs, according to the paper.
ChatGPT 5.2 and DescrybeLM had zero overconfidence flags in this particular run. Make of that what you will.
A Foundation Built on Structured Records

Richard DiBona, Descrybe's co-founder and CTO, emphasized in coverage by legal tech publication LawSites that DescrybeLM wasn't just another fine-tuned general model. "DescrybeLM was built specifically for legal reasoning, on more than 100 million structured records," he said.
The company claims it processed over 100 billion tokens to prepare its proprietary U.S. primary-law corpus. What it hasn't disclosed: the base model architecture or training details. Competitive reasons, naturally.
The reasoning engine is meant to work alongside Descrybe's Legal Research Toolkit, which launched last June at $10 to $20 per month—a price point that undercuts enterprise legal AI substantially. The toolkit finds relevant authority; DescrybeLM applies legal reasoning to the facts at hand. The new engine supports follow-up questions in an interactive mode, though Descrybe didn't use that feature during the benchmark to keep conditions consistent with the general models.
Descrybe has been operating since 2023, according to its LinkedIn profile, though this is based on indirect confirmation rather than direct company disclosure. It remains bootstrapped, which is unusual in an AI landscape where even modest traction typically attracts venture dollars. The company reported 50,000 monthly users as of August 2025 and was added to the National Society for Legal Technology's curriculum around the same time. That curriculum reached over 350 universities and law programs across 11 countries as of that August announcement.
Not bad for a team you could fit in a minivan.
The Competitors That Weren't There
Here's where the study gets interesting—or conspicuous, depending on your view. Descrybe compared DescrybeLM exclusively to ChatGPT, Claude, and Gemini. Not to Thomson Reuters' CoCounsel Legal. Not to LexisNexis's Lexis+ AI. Not to Harvey, which has raised over $100 million to tackle essentially the same problem space.
The rationale, as outlined in the LawSites coverage, centers on accessibility. The general models are widely available and commonly used by legal professionals who may not have access to expensive enterprise tools. Fair enough, perhaps. But the absence of head-to-head comparisons with other legal-specific systems leaves a fairly obvious gap.
No independent third-party replication has been published yet, though the study only surfaced recently.
Reading the Fine Print

To Descrybe's credit, the white paper includes a disclosure section that anyone evaluating legal tech should read carefully. The limitations are acknowledged plainly: the study was vendor-run, not independently audited. It's possible some questions appeared in the training data of any system tested, including DescrybeLM. Each question received a single pass rather than multiple runs with confidence intervals. And the benchmark covers multiple-choice reasoning only—not drafting, citation accuracy, or the messy synthesis work that defines much of legal practice.
There's also the matter of using GPT-5.2 family models as the judge for rubric scoring, which raises potential bias questions. Descrybe applied the same rubric to all outputs, including its own, which helps. But the choice to use OpenAI's technology to evaluate OpenAI's competitor introduces a variable worth noting.
Bob Ambrogi, who covers legal technology at LawSites and has been tracking this space for years, highlighted these caveats in his launch coverage while noting the unusual degree of methodological transparency for a startup product release. That transparency matters—vendor benchmarks are typically about as trustworthy as a campaign ad.
What This Means (and Doesn't)
If the benchmark results hold up under independent testing, they suggest purpose-built legal models may offer accuracy advantages over general-purpose systems on structured legal reasoning tasks. Not exactly shocking, but useful to quantify.
The gap between 93.5% and 100% on a multiple-choice exam doesn't necessarily translate to superiority in real-world legal workflows, though. Drafting motions, conducting discovery, synthesizing case law across jurisdictions—these tasks involve judgment calls and contextual nuance that multiple-choice tests don't capture.
Pricing and packaging details for DescrybeLM itself remain somewhat opaque. The company's legacy website redirects to descrybe.com with a note about a free trial (no credit card required), but plan details aren't visible without signing in—a pattern established in early 2026. Whether the pricing will track the existing toolkit's $10 to $20 range or climb higher isn't clear yet.
For SaaS founders building vertical AI, Descrybe's approach offers an interesting data point. A small, bootstrapped team claims to have built competitive performance by focusing on a curated, domain-specific corpus rather than competing on model scale. The thesis makes intuitive sense: deep expertise in a narrow domain can beat shallow coverage across everything.
Whether that thesis holds under broader scrutiny—and whether Descrybe can scale its methodology across the full range of legal tasks lawyers actually need—remains to be tested. Publishing the methodology is a start. Getting independent researchers to actually run the tests is another thing entirely.
In the meantime, Descrybe has a story to tell: nine people, no VC money, perfect score. The rest is up to the market.
