There's a peculiar inefficiency at the heart of materials science, one that almost everyone in the field knows about but has largely accepted as an immutable fact of research life. You spend months—sometimes years—testing compounds, tweaking temperatures, adjusting pressures. Most experiments fail. Those failures teach you something valuable, perhaps more than the successes do. And then, when you finally write up your work for publication, maybe 5% of what you learned actually makes it into print.
The rest? It vanishes.
Into lab notebooks that gather dust on shelves. Onto hard drives that get wiped when researchers move to new institutions. Into the collective amnesia of a field that publishes only its victories and buries its defeats.
Now a three-person startup called 83 Sciences is making what might sound like an outlandish bet: that this "dark data"—the vast trove of unpublished, mostly negative experimental results—represents one of the most valuable untapped resources in modern science. If they're right, they're sitting on something that could reshape how we discover new materials for everything from batteries to semiconductors to climate technologies.
The timing, at least, seems propitious. Significant venture capital is flowing into materials AI, with multiple companies raising hundreds of millions. Federal mandates are pushing research institutions toward open data. And for the first time, AI models have gotten good enough to actually learn from the messy, contradictory reality of bench science—not just the polished narratives that make it into journals.
Whether a tiny company can capture that opportunity before bigger, better-funded competitors do is another question entirely.
What Gets Lost
The "90% unpublished" figure isn't just startup marketing, though it's important to note that the most rigorous documentation comes from specific fields. A 2025 study in the journal Instruments examined electron microscopy data and found that upwards of 90% of scientifically meaningful EM data never sees publication. For certain subsets, the proportion reached 97%—what the researchers termed "critically low efficiency of data utilization." While this pattern appears across materials science, more field-specific studies are needed to quantify the phenomenon in other domains.
Across the broader landscape of U.S. academic and government labs, a 2023 survey in Cell Reports Methods documented significant volumes and costs associated with unpublished datasets and unused samples. Unfinished projects accounted for roughly half. In the biomedical sciences alone, researchers have estimated that about a quarter of research gets lost to non-publication.
Materials scientists have grumbled about this for years, though perhaps with less urgency than the problem warrants. "Terabytes (sometimes petabytes) from one study result in a few plots in a publication," noted a Nature perspective back in April 2022. A NIST landscape analysis from 2021 flagged fragmented experimental data and the lack of standardized electronic lab notebooks as persistent gaps. The old "file drawer effect"—publication bias favoring positive results—remains stubbornly alive.
What makes this especially maddening is that the unpublished data often contains the most instructive lessons. Failed experiments reveal synthesis pathways that lead nowhere, processing conditions that degrade materials, unexpected side reactions. If you're training an AI model exclusively on published literature, you're learning from a filtered, unrepresentative sample—one that skews heavily toward what worked, not what didn't.
A Confluence of Forces
Three things are happening more or less simultaneously, and their convergence is what makes this moment feel different.
First: the technology has crossed some kind of threshold. DeepMind's GNoME, published in late 2023, predicted roughly 2.2 million candidate crystal structures—expanding the known materials space by an order of magnitude. Berkeley Lab's A-Lab demonstrated autonomous synthesis of dozens of inorganic compounds over 17 days using machine learning and robotics. Microsoft's AI for Science group has introduced tools like MatterGen and MatterSim for generative design and simulations.
By mid-2026, self-driving labs have moved beyond proof-of-concept. Recent preprints describe AI agents handling air-sensitive lithium halide synthesis in gloveboxes, closed-loop thin-film work using laser spike annealing, multi-agent systems orchestrating entire experimental workflows. A July 2026 paper in Communications Materials characterized autonomous labs as entering a new phase—one defined by coordinated, multi-agent frameworks rather than isolated single-agent systems.
Second: venture capital is pouring into the space with an enthusiasm that borders on fervor. Lila Sciences raised $550 million total across two rounds in 2025, with Nvidia participating in the later tranche. Reports indicate that Periodic Labs—founded by former OpenAI researcher Liam Fedus and ex-DeepMind scientist Ekin Cubuk, a co-author on GNoME—closed a $300 million seed round in late 2025. CuspAI reportedly raised $100 million at a roughly $520 million valuation. Orbital Industries pulled in $50 million for its Series B in May 2026.
Market forecasts reflect this surge. Grand View Research pegs the materials informatics market at $134.6 million in 2023, projecting growth to $390.8 million by 2030—a compound annual rate exceeding 16%. Another analysis sees the sector climbing from $183.5 million in 2026 to more than $730 million by 2035.
Third: policy is pushing institutions in the same direction. The White House Office of Science and Technology Policy issued a memo in August 2022—often called the "Nelson memo"—mandating immediate public access to federally funded research and underlying data. Implementation has been rolling out through 2025 and 2026. The Department of Energy launched its Genesis Mission in late 2025, a multi-hundred-million-dollar effort explicitly designed to unify labs, data, compute, and AI infrastructure for scientific discovery.
The Genesis funding opportunity announcement from March 2026 lays out phases aimed at building common platforms and autonomous lab systems. For companies in this space, that represents both opportunity (standardization could expand their addressable market) and competition (if the government builds free infrastructure, what's left to sell?).
Enter 83 Sciences

This is the landscape that 83 Sciences is stepping into—or perhaps stumbling into, depending on how generous you're feeling.
The company describes itself as an "intelligence engine powering the future of research and materials discovery." According to its website, it offers in-person onboarding, tools to automatically capture and sort experimental data, a "lab brain" for querying that data, and AI agents that "understand your failed experiments and propose optimized process conditions."
Target customers span a broad swath of industries: chemicals, energy, batteries, mining, semiconductors, cement, aerospace, pharma, academia, national labs. Basically anyone running materials experiments at scale.
The founding trio brings a mix of industrial strategy, hands-on machine learning, and enterprise AI deployment. Ian Naccarella did stints at MIT (commercializing intellectual property), Sila Nanotechnologies (strategy), and Boston Consulting Group, with degrees from Stanford and Harvard. Eric Riesel holds a PhD in inorganic chemistry from MIT and co-led research on using generative models to solve crystalline structures from powder diffraction data—work that MIT's School of Science highlighted in early 2025. Yankang Yang spent time as a BCG principal leading an internal AI program that reportedly reached more than 30,000 users; he studied computer science and electrical engineering at Harvard.
As of this writing, 83 Sciences has no publicly disclosed customers, case studies, or pricing information on its company profile. That's typical for a company fresh out of Y Combinator—the company is listed as part of YC's Summer 2026 batch, which is consistent with its early-stage status.
They're entering a space where established players have already demonstrated demand. Citrine Informatics, founded in 2013, raised $16 million in a Series C in early 2023 and markets case studies claiming 50–80% reductions in required experiments. Mat3ra (formerly Exabyte) raised a $3 million seed in 2022 and focuses on cloud simulation plus AI workflows. Both target similar industrial verticals but emphasize different layers of the technology stack.
What 83 Sciences seems to be emphasizing—and it's still early enough that "seems" is doing heavy lifting—is the capture layer itself. Getting raw, unpublished, failed data structured at the source, before it disappears. That's adjacent to but distinct from the autonomous lab robotics vendors, the foundation model players like Microsoft, and the computational materials platforms.
Whether that distinction is defensible, or lucrative, remains to be seen.
Real-World Momentum

If this all sounds abstract, there are concrete examples of the technology moving from slideware to functioning systems.
Battery researchers reported in April 2026 on using deep active learning and knowledge transfer to discover lithium metal electrolyte candidates—a collaboration involving teams in China, Brookhaven National Lab, and SES. Earlier work in January 2024 screened more than 32 million candidates, narrowing down roughly 500,000 stable materials for sodium and lithium solid electrolytes.
Self-driving polymer labs at Argonne National Laboratory have demonstrated inverse design across 2024 and 2025. A paper in Digital Discovery this past June detailed the practical engineering challenges—harsh chemistries, multi-step syntheses, air-free environments—that don't make it into the glossy announcements but determine whether these systems can actually generalize.
Catalysis researchers published in-situ atomic-resolution studies of CO₂ reduction catalysts in mid-2024, enabling mechanistic insights that simply weren't possible before. The pattern emerging across these applications: faster iteration, tighter feedback loops, the ability to test hypotheses that would have been impractical with purely human-driven experimentation.
Perhaps the most striking validation comes from researchers pointing out what doesn't work. An ACS Omega perspective published in June 2026 argued that raw experimental data should serve as a verification layer for AI models, noting troublingly high error rates in published datasets and poor resynthesis rates—particularly for metal-organic frameworks—that undermine models trained only on literature. If your training corpus is biased toward what got published, you're learning from a skewed, non-representative sample of reality.
Hurdles and Headwinds

For all the momentum, significant obstacles remain.
NIST issued what it called a "Call for Caution" back in October 2023, emphasizing the need for robustness, reproducibility, and proper validation in AI-accelerated materials research. Those warnings have been echoed in multiple 2025 and 2026 papers on trustworthy AI frameworks and benchmarking challenges. Just because the technology can generate millions of candidate materials doesn't mean those candidates will actually work when someone tries to synthesize them.
Data governance remains messy. The OSTP mandate pushes toward open data, but legal exceptions, intellectual property concerns, and implementation costs create friction. A Government Accountability Office report in June 2026 flagged cost and implementation risks as agencies navigate the transition. Companies like 83 Sciences will need to figure out how to operate in environments where some data must stay proprietary, some must be shared, and the rules keep shifting.
Hardware and operations are harder than software folks often anticipate. Engineering self-driving labs for harsh chemistries, ensuring equipment reliability, generalizing active learning algorithms across different material classes—all of this requires deep domain expertise, not just coding chops.
And then there's the fundamental business question: can anyone actually make money doing this? The Genesis Mission and similar federal programs might accelerate adoption by creating common infrastructure and data standards. But converting national lab pilots into commercial revenue at industrial companies is an entirely different challenge. Chemicals companies and battery manufacturers have their own proprietary data, their own workflows, their own reasons for secrecy. Convincing them to share data—even anonymized, structured data—will require more than just slick demos.
The Long Game
Still, the underlying logic is hard to argue with. Materials discovery is painfully slow. Ten to twenty years from lab to market for some applications. If AI and automation can shave even a few years off that timeline, the economic value is staggering. Batteries, semiconductors, catalysts, climate materials—these sit at the center of multi-trillion-dollar industries.
The question isn't really whether unpublished experimental data holds value. It demonstrably does.
The question is who builds the infrastructure to capture, structure, and learn from it at scale—and whether that infrastructure becomes a platform that others build on or remains fragmented across proprietary systems. For a three-person startup taking on established players and well-capitalized competitors, success will likely hinge on execution speed, early customer wins, and the ability to deliver on a simple but powerful idea: that failure, properly understood, can teach us as much as success. Maybe more.
