In a nondescript laboratory somewhere, right now, a graduate student is closing a notebook on an experiment that didn't work. The catalyst didn't catalyze. The battery electrolyte degraded faster than expected. The alloy came out with the wrong crystal structure. Into the drawer it goes, maybe into a hard drive, almost certainly never to be published. And somewhere else, another researcher—next month, next year—will likely try the same thing, fail the same way, and learn nothing from the failure that came before.
It's a peculiar kind of waste. Not the dramatic kind that makes headlines, but the silent, compounding waste of knowledge that never circulates. By some estimates, approximately 85% of biomedical experimental data never sees publication—a workshop estimate discussed at a National Academies forum, though not a definitive statistic. In electron microscopy, an analysis from 2025 found that more than 90% of recorded data in the specific dataset studied remains trapped in lab computers. Even in chemistry, where precision matters and notation is everything, the Dark Reactions Project at Haverford College has spent years documenting a stubborn truth: failed syntheses—the ones that never make it past peer review—often contain the most critical information for training machine learning models.
Which raises a question that's lately been animating a handful of startups: What if all that "dark data" isn't worthless? What if it's actually valuable?
A Contrarian Wager
Enter 83 Sciences, a three-person team that emerged from Y Combinator with a pitch as straightforward as it is audacious. The name itself comes from their central claim: that more than 83% of experimental data generated in research labs gets discarded. Their plan? Use AI to mine that dark data for breakthroughs in materials needed for clean energy and decarbonization.
The founders—Ian Naccarella, Eric Riesel, and Yankang Yang—bring pedigrees spanning MIT, Stanford, Harvard, Sila Nanotechnologies, and Boston Consulting Group. They're positioning unpublished experimental failures not as embarrassments to be buried, but as raw material for faster papers, patents, and discoveries. On their website, they claim that over $100 billion in R&D value is lost yearly to discarded data—a figure they assert but which lacks third-party verification. They promise to convert a partner's abandoned experiments into a co-authored manuscript in under two months. Their target sectors read like a decarbonization wish list—energy storage, critical minerals and metals, catalysis, life sciences solid forms, semiconductors.
As of mid-2026, the company has disclosed no funding beyond its YC participation. No named partners or case studies appear on its site, either. It's early days, in other words. But the thesis underlying 83 Sciences reflects something broader: a shift in how materials informatics is practiced, accelerated by policy mandates, geopolitical supply chain anxieties, and the slow maturation of AI tools purpose-built for scientific discovery.
Whether 83 Sciences can execute on that thesis—whether anyone can, really—is another question entirely.
The Tools to Extract the Invisible

The materials informatics market itself isn't large, not yet anyway. As per Grand View Research, the market size was $134.6 million in 2023, with forecasts reaching $390.8 million by 2030. Other analysts project faster growth: 360iResearch estimates $211.6 million in 2026, climbing to $583.8 million by 2032. The Business Research Company sees the broader AI-in-materials-discovery market at $740 million this year, potentially hitting $2.77 billion by 2030.
Those numbers sit inside a much larger universe—advanced materials is a trillion-dollar-plus global market—which suggests that while AI-driven discovery is gaining traction, adoption remains concentrated. What's changed recently, though, is the arrival of industrial-grade platforms designed to operationalize dark data mining.
Microsoft's Discovery platform, which went generally available in July 2026, is perhaps the clearest signal that big tech sees scientific dark data as strategic infrastructure. Microsoft positions it as an end-to-end agentic AI system for chemistry and materials, with early use cases spanning energy storage materials and mining reagents. The platform can orchestrate autonomous lab workflows and maintain a "scientific loop" memory across experiments—essentially building institutional knowledge from both successes and failures.
Google DeepMind continues its materials science push, following its November 2023 GNoME release, which predicted 2.2 million candidate stable crystals (though subsequent discourse raised questions about duplicates and validation). At Google I/O 2026, the company previewed "Gemini for Science," signaling continued investment. Meta's Open Catalyst Project released OC25 in 2025, expanding machine learning datasets for electrocatalysis.
Meanwhile, specialized players have carved out niches. Citrine Informatics published a case study showing a global top-five adhesives company identifying a PFAS-free candidate in four months using its platform. Aionics launched its Artificial Molecular Intelligence platform in June 2025, targeting battery electrolytes. Orbital Materials raised a $50 million Series B in May 2026 to discover what it calls "exotic materials." QuesTek continues pushing its integrated computational materials engineering (ICME) model.
The ecosystem is forming, in other words. But infrastructure is only half the story.
Policy as Unlock

Two regulatory shifts are opening the floodgates. The NIH Data Management and Sharing Policy, which took effect in January 2023, requires researchers to submit data-sharing plans alongside grant proposals. More significantly, a 2022 White House Office of Science and Technology Policy memo directed federal agencies to implement zero-embargo public access for publications and research data by December 31, 2026.
That deadline is now. By year-end, a substantial volume of federally funded research data is supposed to become machine-readable and accessible—subject to privacy, security, and intellectual property protections, of course. The move mirrors the European Union's push for FAIR (Findable, Accessible, Interoperable, Reusable) data infrastructures, exemplified by initiatives like NOMAD and FAIRmat.
At the same time, geopolitical supply chain pressures are intensifying demand for materials discovery. The EU's Critical Raw Materials Act entered into force in May 2024, with implementing regulations rolling out through 2025 and 2026. The U.S. Inflation Reduction Act's battery supply chain provisions—including Foreign Entity of Concern restrictions and evolving Treasury guidance—are pushing companies to find substitutes and alternative chemistries, fast.
The result: corporate R&D labs face mounting pressure to accelerate innovation while simultaneously being pushed to share more research data publicly. Dark data mining sits squarely at the intersection of those forces.
And yet.
The Messy Reality of Access
Not all dark data is accessible, even with policy tailwinds. University research data policies reveal a thicket of ownership claims. The University of California system, the University of Minnesota, Stony Brook, the University of South Alabama, and Virginia Tech all assert institutional ownership of research data and lab notebooks. Bayh-Dole provisions govern inventions arising from federally funded research, meaning any unpublished data that contributes to a patentable discovery triggers invention reporting and IP provisions.
For a startup like 83 Sciences—or any platform seeking to train models on unpublished lab data—this creates a complex negotiation landscape. Universities must grant access. Researchers must consent. Sponsors, if involved, may have contractual rights. The two-month timeline 83 Sciences claims for turning discarded data into co-authored manuscripts assumes frictionless access and cooperation. That's... optimistic.
Data quality compounds the challenge. Several papers published in 2025 and 2026 in journals like RSC Journal of Materials Chemistry A, along with reviews on self-driving labs, warn that low-integrity data—undocumented experimental conditions, instrument drift, inconsistent metadata—can degrade AI models and misdirect research. A 2023 NIST analysis found 95% redundancy across some materials datasets, suggesting that indiscriminate data aggregation may introduce noise rather than signal.
The scientific community has responded with calls for trustworthy AI frameworks. The GIFTERS framework, proposed in November 2025, emphasizes governance, integrity, fairness, transparency, explainability, reproducibility, and safety. DeepMind's 2026 policy brief, "Science Needs AI Data Stocktakes," argues for systematic inventories of experimental data to identify what's siloed, what's dark, and what's accessible.
All of which is to say: the technical challenge is daunting, but the institutional and legal challenges might be worse.
Does It Actually Work?

The evidence is mixed but growing. The Dark Reactions Project at Haverford has demonstrated that negative and failed syntheses improve machine learning predictions in chemistry, countering publication bias toward positive results. A December 2024 preprint from a randomized study at a large U.S. firm reported that AI assistance in materials R&D led to 44% more materials discovered, 39% more patent filings, and 17% more downstream product innovation.
Citrine's PFAS-free adhesive case study—though lacking granular performance metrics—illustrates the commercial appeal: a breakthrough candidate in four months for a top-five global player. Uncountable's partnerships with battery materials companies like Mitra Chem and Group1, detailed in case pages from May 2025, suggest that data platforms are embedding themselves in high-stakes R&D workflows.
Autonomous lab demonstrations in 2026—including an agentic self-driving lab for air-sensitive lithium-halide spinels and human-aware robotics platforms—show technical feasibility. Yet reviews and workshop reports note that self-driving lab adoption remains early-stage, limited by costs often exceeding $1 million per customized platform and challenges translating computational findings to manufacturing scale.
In other words, it works, sometimes, for some applications, at considerable expense. Not exactly a ringing endorsement, but not a dismissal either.
What Happens Next
The trajectory for dark data mining in materials science seems clear enough: more data will become accessible, more tools will emerge to structure and analyze it, and more companies will attempt to monetize the insights. The question is whether the value proposition scales beyond niche applications.
For 83 Sciences and its competitors, success hinges on several variables. Can they negotiate data access agreements fast enough to build defensible datasets? Can they maintain data integrity standards rigorous enough to avoid the "garbage in, garbage out" trap that has plagued other AI-for-science efforts? Can they translate computational predictions into experimentally validated materials that companies will actually pay to license or co-develop?
The broader materials informatics market is moving from pilots to portfolio-level deployments, according to case studies from Citrine and Uncountable, but adoption remains uneven. The arrival of platforms like Microsoft Discovery may accelerate adoption by lowering the barrier to agentic AI workflows, particularly for labs that can integrate their existing data estates. Simultaneously, the EU Critical Raw Materials Act and U.S. Inflation Reduction Act supply chain mandates ensure sustained demand for substitutes, new alloys, electrolytes, and recycling catalysts.
If the 83% figure holds—that the vast majority of experimental data never contributes to published knowledge—the opportunity is enormous. But the figure itself is domain-dependent. Biomedical research may discard 85% of data; electron microscopy over 90%; chemistry and materials science likely somewhere in that range. What's consistent is that publication bias systematically excludes failures, near-misses, and negative results. Precisely the data that could prevent the next researcher from repeating the same mistake.
The race now is to see who can capture that knowledge before it disappears into forgotten hard drives. Whether 83 Sciences can compete with better-capitalized players like Orbital Materials, established platforms like Citrine, or tech giants like Microsoft remains an open question. The company is betting that being AI-native and laser-focused on dark data gives it an edge.
By the end of this year, as the OSTP mandate kicks in and more research data floods into public repositories, we'll have a clearer picture of whether that bet pays off. For now, though, it's a wager on waste—on the idea that what we throw away might be worth more than what we keep.
