Somewhere in a university lab, a graduate student is tossing out another batch of failed synthesis attempts. In an industrial R&D facility, a characterization run that didn't match expectations gets archived under "miscellaneous data, 2025." At a battery startup, engineers shelve test results from an electrolyte formulation that—well, it just didn't work.
Multiply those moments by tens of thousands, and you begin to grasp the scale of what researchers call "dark data": the vast, unmapped territory of experiments that didn't pan out. These aren't measurement errors or lab accidents. They're legitimate scientific inquiries that simply revealed what doesn't happen under certain conditions. And until recently, they've been treated as the cost of doing business—documented in lab notebooks, maybe, but effectively erased from the scientific record.
Here's the twist. A cohort of AI-focused startups now argues that this discarded information represents one of the most underutilized assets in materials science. Their bet? That if you want artificial intelligence to truly understand chemistry—not just simulate it—you need to feed the models what actually happens when researchers mix compounds, heat substrates, and run synthesis protocols. Especially, perhaps, when things go sideways.
The published literature, after all, captures success stories: breakthroughs, validated phase diagrams, performance improvements. What vanishes is everything else. Some estimates suggest the volume of unpublished experimental outcomes dwarfs peer-reviewed results by orders of magnitude. That imbalance is starting to look less like an inevitability and more like a market inefficiency.
When Models Meet Messy Reality
Materials informatics has entered something of a boom phase. Market research firms can't quite agree on the numbers—more recent estimates from Mordor Intelligence place the sector at $160.76 million in 2025, while earlier Grand View Research figures estimated $134.6 million in 2023, climbing toward $390.8 million by 2030. Fact.MR pegged it higher at $208.4 million last year, heading to $1.35 billion by 2036 with an 18.5% compound annual growth rate. The variance is typical for emerging categories, but the directional arrow is unmistakable.
The catalyst arrived in stages. DeepMind's GNoME model, published in Nature in 2023, proposed 381,000 previously unknown stable crystal structures using deep learning—a number that raised eyebrows even in a field accustomed to computational scale. Hundreds have since been experimentally confirmed. Around the same time, Berkeley's A-Lab demonstrated something equally striking: autonomous synthesis systems running closed-loop campaigns that iteratively refine targets based on real-world lab outcomes, not just simulations.
By this year, multi-agent frameworks are orchestrating entire lab operations—managing instruments, planning experiments, learning from each run. The pace has been dizzying. It's also surfaced a technical wrinkle.
Models trained largely on computational predictions or published successes tend to propose materials that look promising on screen but prove stubbornly difficult to synthesize in practice. A technical analysis circulating in materials science circles warns that AI-generated candidates can drift from what researchers call the "structural memory" of experimental chemistry—the tacit, hard-won knowledge of what's actually makeable with real precursors and real lab constraints.
The proposed fix? Ground the models in historical experimental data. Not just the winners. Everything.
The Policy and Infrastructure Convergence
Several forces are aligning to make this once-ignored data suddenly valuable, and not all of them are technical.
Policy matters here more than you might expect. The U.S. Department of Energy's Genesis Mission, unveiled in February 2026, outlined 26 AI-related challenges including digitization of legacy scientific data—an acknowledgment that decades of experimental records sit in formats that modern machine learning can't touch. In June 2026, the National Institute of Standards and Technology awarded $500 million under the CHIPS R&D program to SandboxAQ, explicitly to accelerate AI-driven semiconductor materials discovery. Meanwhile, the European Union's Critical Raw Materials Act, which took effect in May 2024, imposes supply-chain risk assessments that indirectly increase demand for faster materials qualification.
Then there's the technical case, which has hardened considerably. IBM Research demonstrated in mid-2025 that incorporating negative chemical data—reactions that didn't produce the intended product—improves predictive models for reaction pathways. The problem, IBM researchers noted dryly, is the lack of venues to report such outcomes. Nature documented around the same time that recursive training on synthetic data can lead to "model collapse," a technical term for when AI systems trained on their own outputs gradually lose touch with reality. A 2016 Nature study on "dark reactions" in inorganic synthesis showed that folding in negative outcomes directly improves both predictive accuracy and experiment selection.
Infrastructure is catching up, too, though not as quickly as optimists might hope. A community survey last year from NFDI4Chem found that electronic lab notebook adoption reached 34% in organic chemistry and 33% in materials science—higher than many expected, though still leaving two-thirds of labs working with paper or ad hoc digital systems. Hardware-based data sharing via thumb drives dropped from 24% in 2019 to 10% in 2023. Labs are digitizing. Slowly. But the raw material for training AI is becoming accessible in ways it simply wasn't five years ago.
Commercial interest is accelerating, too. McKinsey's ongoing analysis of "scientific AI" suggests substantial potential to lift R&D throughput in sectors close to discovery—chemicals, materials, pharmaceuticals—but cautions that realizing value requires organizational rewiring and careful pipeline governance, not just software deployment. Gartner formalized a Market Guide for materials informatics solutions in mid-2024, giving enterprise procurement teams a defined category to evaluate. IDTechEx observed that executive-led adoption picked up sharply after 2023, with new entrants flooding the market through 2024 and into last year.
From Theory to Deployment

The concept is moving faster than many anticipated. Cabot Corporation, a global specialty chemicals firm, began rolling out Uncountable's R&D platform across its labs starting in late 2020. The goal was straightforward: digitize experiment workflows and create structured lineage tracking, essentially converting what had been dark data into reusable assets. A case study published by Uncountable emphasizes standardized experiment templates and global data access. SCG Chemicals, another client, reported a 20% acceleration in R&D timelines post-implementation—though that figure comes from vendor-provided materials and should be read with the usual grain of salt.
Aionics launched its Artificial Molecular Intelligence platform in mid-2025 and is working with Porsche's Cellforce unit on electrolyte formulation. The AMI system describes real-time autonomous screening: actively learning from every batch, including the ones that fail to meet performance specifications. Orbital Materials raised a $50 million Series B this past May (following a $21 million round in late 2024) and announced a strategic partnership with Amazon Web Services to develop materials aimed at data-center decarbonization. The company positions itself around foundation models trained on diverse materials datasets—presumably including negative results, though the details remain proprietary.
Then there's 83 Sciences, part of Y Combinator's Summer 2026 batch. It's a three-person team split between San Francisco and New York, but the positioning is notable. The pitch: "AI-native materials discovery powered by unpublished experimental data." Their website describes a rapid-turnaround case study in which the team analyzed discarded characterization data, identified a previously unnoticed signature limiting performance, and—using an AI assistant called Dalton alongside human chemists—proposed synthesis modifications and progressed toward manuscript preparation in under two months. The company mentions "dozens of other partnerships in progress," though at this stage that's more aspiration than track record.
What makes 83 Sciences worth watching isn't scale—it's very early—but rather the founding team's backgrounds. Co-founder Eric Riesel holds a PhD in inorganic chemistry from MIT and worked on machine-learning systems for complex materials data, including what MIT described as "one of the first generative approaches for chemical experimental data." Ian Naccarella spent time at Sila Nanotechnologies commercializing battery materials and previously worked strategy at BCG. Yankang Yang, the third co-founder, led BCG's AI program rollout to more than 30,000 users. The combination spans enterprise AI implementation, industrial materials, and generative ML on experimental data—explicitly designed for lab-in-the-loop capture and reuse of messy real-world results.
A job posting for a Founding AI Engineer lists a salary range of $120,000 to $250,000 with 0.5% to 5% equity. They're hiring.
The Hard Questions Ahead
Whether this wave of startups succeeds depends on several open questions, some technical, some organizational, and some just plain hard.
First: Can AI platforms navigate the intellectual property and confidentiality constraints that keep most industrial data locked away? Prior surveys from the Materials Research Society consistently cite proprietary restrictions and IP concerns as top impediments to data sharing. Successful commercial models will need federated learning architectures, data enclaves, or other frameworks that allow companies to contribute experimental outcomes—including failures—without exposing trade secrets. That's not a trivial engineering problem, and it's definitely not a trivial legal one.
Second: Will regulatory frameworks help or hinder? The EU AI Act entered general applicability on August 2, with research exemptions carved out in specific articles, but interpretation is still evolving. As these tools move from "research project" to "commercial product," obligations around transparency, risk assessment, and documentation will increasingly apply. In the United States, EPA's Toxic Substances Control Act imposes a 90-day review before manufacturing or importing new chemical substances, with updated guidance rolling out through last year and into this one. Faster materials discovery won't translate to faster market entry unless regulatory pathways adapt—and there's little indication they will.
Third, and perhaps most important: Will the technical promise actually hold? The critique about "structural memory" suggests that models trained predominantly on computed or AI-generated proposals lose touch with what's synthesizable in a real lab. Mining failed experiments is one potential answer—anchoring predictions in the empirical boundaries of actual bench work. But it's not straightforward. Negative data is messy, context-dependent, often poorly documented. A synthesis that "failed" might have failed for a dozen reasons: wrong temperature, contaminated precursor, instrument drift, simple operator error on a Friday afternoon. Extracting meaningful signal from that noise requires not just machine learning but deep domain expertise and meticulous metadata capture.
Still, the opportunity commands attention. The Dark Reactions Project at Haverford College aggregates failed hydrothermal syntheses as a training corpus. The Open Reaction Database explicitly accepts negative outcomes. NFDI4Chem's survey shows that researchers with better data management practices are already demonstrably more productive. The infrastructure is coming online, even if haltingly.
The Asset Hiding in Plain Sight

For hard-tech founders and R&D leaders, the calculation is shifting. Failed experiments used to be pure sunk cost—time and reagents spent learning what doesn't work, with nothing to show for it but hard-won intuition. Now they might constitute an asset, provided someone can extract the lessons at scale and do it in a way that preserves competitive advantage.
The companies that succeed in this space won't just be building better neural networks or more sophisticated active learning loops. They'll be building trust frameworks that convince industrial labs to share what they've been discarding—or hoarding—for decades. That's a harder problem than the machine learning. It involves IP attorneys, procurement teams, and convincing senior chemists that their failed experiments have value beyond avoiding the same mistake twice.
But if it works? The pace of materials discovery could shift in ways that make even the recent progress look incremental. Every industry from batteries to semiconductors to pharmaceuticals depends on finding better materials faster. Dark data represents a massive, largely untapped training set for the AI systems that might unlock those discoveries.
Which raises one final thought: perhaps the most valuable thing a researcher ever did wasn't the breakthrough that got published. Maybe it was the experiment that didn't work—provided someone else learns from it.
