Mattias Akke spent enough time watching molecular simulations at MIT and drug pipelines at AstraZeneca to develop a particular frustration. The problem wasn't that scientists lacked automation or computing power. It was that nobody could tell when the machine was actually thinking versus when it was just filling in blanks with plausible-sounding nonsense.
So Akke and two colleagues built Halmos Labs, a three-person startup in Lund, Sweden, that announced its software to the public this month, designed to surface every assumption it makes before designing a wet-lab experiment. The timing is deliberate. Biotech companies are pouring money into AI-driven drug discovery while the cost to bring a single drug to market has climbed to $2.67 billion, according to Deloitte's latest industry analysis, and the return on all that R&D investment across the top 20 biopharma firms dropped to 5.9 percent based on 2025 data.
Halmos Labs, part of Y Combinator's Fall 2026 cohort, released a demonstration in October showing an autonomous pipeline that designed a manufacturable mRNA cancer vaccine from a single genetic mutation. The company annotated each step with benchmark performance data and caveats about where the system might be wrong. It's a bet that interpretability matters more than speed when regulators and drug developers start asking hard questions about how an AI reached its conclusions.
The pitch arrives as drug discovery automation is projected to reach $10.0 billion by 2026 and grow to $18.4 billion by 2033, according to Grand View Research. IQVIA noted in May that "early evidence of stronger success rates for AI-enabled programs" is emerging among smaller biotech firms. At the same time, anatomic pathology labs are struggling with 28.5 percent vacancy rates, and the FDA and European Medicines Agency released joint AI principles in January demanding lifecycle validation for any algorithm touching drug development.
Whether a three-person team in Sweden can compete with platform companies and pharmaceutical giants is an open question. But the technical problem they're tackling is real enough that NVIDIA, Benchling, and several well-funded startups have all released competing tools in the past six months.
The Two-Decade Productivity Slide
Pharmaceutical R&D has been stuck in a productivity decline that predates the current AI wave by more than a decade. Deloitte's internal rate of return figure for 2024 represents a steady erosion from double-digit returns ten years earlier. The average cost to develop and approve a single drug asset is projected to be $2.67 billion as of 2026 research. IQVIA's Global R&D Trends report noted that biopharma funding hit $82 billion in 2025, some 44 percent above pre-pandemic averages, yet clinical success rates haven't budged much and timelines remain stubbornly long.
The bottleneck sits upstream in the experimental design phase. A biomedical reproducibility survey reported by Nature in January found nearly three-quarters of researchers perceive a crisis in their field. A July preprint analyzing one life-sciences subfield concluded that 38 of 45 published claims could not be reproduced when other teams tried. McKinsey argued in June that real value is shifting toward end-to-end workflows with agentic AI, though the firm acknowledged that real-world impact remains fragmented.
Lab automation has solved part of the problem. Ginkgo Bioworks now connects more than 100 devices per automation system at its Nebula facility, the company disclosed in its first-quarter earnings deck. Physical throughput has scaled. But deciding what to test next has stayed largely manual or driven by heuristics that don't adapt well to new data. "Maximum information per experiment, minimum experiments per conclusion," reads the tagline on Halmos Labs' Y Combinator profile page.
Built to Show Its Work
Halmos describes its platform as "built to be checked." The system surfaces assumptions and inductive reasoning steps for human review, then samples across competing methodological choices to design experiments that would yield the most information regardless of which scientific hypothesis turns out correct. That's the description on the company's October website, at least.
Akke, who founded the company in 2026, summarized the underlying problem in a September preprint: frontier large language models still fail to decisively outperform statistical baselines when selecting biochemical experiments, and they suffer from what he calls "context-stickiness" and overreaction to fresh data. The paper benchmarked five leading LLMs against classical methods across seven biochemical datasets and introduced diagnostic tools to reveal when models explore poorly or anchor too hard on recent results.
The public demo walks through five phases to design a personalized cancer vaccine from a single mutation. Mutation calling, expression modeling, neoantigen prioritization, RNA sequence design with folding checks, and manufacturability scoring. Each phase comes annotated with benchmark gates and caveats. The company says the pipeline recovered all six known CD8 T-cell responses for a published patient case within its top-10 ranking out of 152 candidates, though Halmos flagged the test as enrichment validation rather than a prospective benchmark.
Akke posted on LinkedIn in early October that the team spent "many months" building the preprint and demo. CTO Linus von Ekensteen Löfgren previously built LLM security tooling at AI Sweden and worked as a quantitative analyst at Danske Bank. The three-person team is backed by Y Combinator's standard deal structure.
The technical claim rests on Bayesian optimal experimental design. Rather than letting a black-box model pick the next assay, Halmos samples multiple plausible analysis pipelines and selects experiments that would best discriminate among them. It's a way of forcing the system to consider what it doesn't know before committing resources to a wet-lab run.
A Rapidly Consolidating Field

Halmos is entering a market that has compressed considerably in recent months. NVIDIA announced its BioNeMo Agent Toolkit on June 23, billing it as software that gives AI agents "the skills of a PhD research assistant and the speed of a supercomputer," CEO Jensen Huang said in the press release. Launch partners included Eli Lilly, Benchling, and hardware vendors Thermo Fisher, Tecan, and Automata.
Benchling followed a month earlier with Benchling Automation, a hardware-agnostic orchestration layer that connects lab instruments to scientific record systems in what the company calls a "lab-in-a-loop." HighRes Biosolutions positioned its Cellario OS as an open execution layer for AI scientists in June, integrating with NVIDIA's toolkit shortly after. Ginkgo Bioworks told investors in August that "autonomous labs are increasingly becoming necessary national infrastructure for American science" and disclosed in September that it would build an autonomous facility for Novo Nordisk's new Waltham R&D site.
LabGenius announced a Sanofi collaboration in December using its EVA closed-loop antibody discovery platform, which the company says achieved 14-fold enrichment over a standard baseline in a bispecific antibody study. Arctoris operates a robotic contract research platform called Ulysses. Recursion has partnered with NVIDIA and maintains more than 23 petabytes of proprietary phenomics data. Insitro launched its TherML platform in January through an acquisition. Carnegie Mellon's AI Science Foundry, selected by the National Science Foundation in July, now connects more than 80 robotically controlled instruments across two cloud labs.
Three other Y Combinator companies joined Halmos in the biotech-automation cohort this fall: Infera, which offers natural-language lab control; Enjamb Labs, building AI workflows spanning preclinical to regulatory stages; and b-12 Labs, focused on code generation for chemistry robots.
What Regulators Want to See
The FDA and European Medicines Agency released joint "Guiding Principles of Good AI Practice in Drug Development" on January 14, covering AI use across the drug lifecycle. The FDA has logged more than 500 AI-component submissions between 2016 and 2023 on its AI in Drug Development portal and qualified its first AI Drug Development Tool, AIM-NASH, in late 2025, according to the agency's spring newsletter. The EU AI Act's transparency obligations began enforcement in August. The UK's MHRA launched an AI sandbox in June to accelerate medicines development.
A March paper in npj Digital Medicine on causal AI in drug development stressed interpretability, overfitting controls, and transparency as essential to meeting regulator expectations. A February article in Chemical Science argued that common attribution methods like SHAP and LIME don't establish causation, which regulators and peer reviewers increasingly expect to see demonstrated. Preprints appearing in June and July on autonomous wet-lab agents introduced runtime guardrails and safety monitors for low-cost robotics platforms, noting that execution validity remains an open research problem.
McKinsey's June analysis of healthcare AI said high performers orchestrate lab-in-loop pipelines rather than deploying point tools, but talent shortages and data platformization remain bottlenecks. BCG wrote in January that AI's potential spans discovery through manufacturing, yet measurable impact stays fragmented. IQVIA's May report suggested watching clinical outcome readouts from AI-enabled programs during the next couple of years. Axios reported in September that optimism about preclinical acceleration has not yet translated to late-stage proof points.
The Execution Test

Halmos Labs is positioning validity and interpretability as competitive advantages in a market where trust determines adoption speed. The public demo annotates data lineage, benchmark performance, and methodological assumptions at every step, which aligns with FDA guidance on model credibility and lifecycle control. The September preprint diagnosing LLM exploration failures suggests the team is building diagnostics into the core system rather than retrofitting explainability onto opaque agents later.
The broader category is moving quickly. Ginkgo CEO Jason Kelly said in May that "autonomous labs will replace the lab bench more quickly than people think." NVIDIA's agent toolkit, Benchling's orchestration launch, and HighRes's AI-scientist positioning all landed within a four-month span earlier this year. University and pharmaceutical adoption is live: Ginkgo is building facilities at MIT, Caltech, Northwestern, and the University of Maryland alongside the Novo Nordisk site.
Technical and regulatory questions remain unresolved. Halmos's own preprint showed that current LLM-based experimental design underperforms classical Bayesian optimization in many domains. Reproducibility and runtime validation for physical automation are active research areas. EU AI Act compliance and FDA lifecycle documentation requirements add overhead that smaller teams may struggle to manage. Workforce vacancies in labs create demand for automation, but commentary accompanying an ASCP survey in December noted AI as a complement to personnel rather than a replacement.
For early-stage investors, the wedge question is whether three people in Sweden can out-execute platform companies and incumbents on credibility infrastructure while regulatory requirements tighten. For drug developers, the question is whether end-to-end pipelines that surface their own uncertainty can compress the billion-dollar, decade-long development cycle enough to justify switching costs.
Halmos Labs has a demo, a preprint, and Y Combinator backing. The next validation happens in wet labs, where assumptions meet actual experiments and the cost of being wrong is measured in months and reagent budgets. That's where interpretability either matters or it doesn't.
