When Steeve Morin talks about his new company, he doesn't start with the technology. He starts with a problem that sounds almost geopolitical: too many AI models, he says, are imprisoned by the chips they run on.
ZML, the Paris-based infrastructure startup Morin founded, announced $20 million in seed funding on July 8, 2026, alongside the launch of LLMD alpha—a server designed to run large language models across NVIDIA, AMD, Google TPU, Intel, and Apple silicon without the usual gymnastics. The pitch is straightforward, maybe deceptively so: build once, deploy anywhere. The execution? That's where things get interesting.
The funding round drew 20VC, LocalGlobe, >commit, AAL VC, Drysdale Ventures, Kima Ventures, Kindred Capital, and Puzzle Ventures. But the cap table's real calling card might be the angels: Turing Award winner Yann LeCun, Docker creator Solomon Hykes, and Hugging Face co-founders Clément Delangue and Julien Chaumond all put money in. When the person who invented convolutional neural networks and the person who containerized software development both show up, you pay attention.
A Compiler Problem Dressed Up as a Business Model
ZML is giving away LLMD alpha for free. Not open source—free. The distinction matters. Morin's team wants usage data before they figure out how to charge, which is either admirably pragmatic or a sign they're still working out the business model. Probably both.
What they've built is a universal LLM server that supports Qwen, Gemma, LLaMA, and Mistral model families across five hardware platforms at launch. The technical packaging is dense: hermetic images per platform (CUDA at 1.7 GB, TPU at 280 MB, Metal via Homebrew at 140 MB), continuous batching, paged attention, tensor-parallel sharding, prefix caching. The stack is written in Zig, uses MLIR and OpenXLA, and explicitly avoids Python in the execution path—a choice that signals Morin's background at Google and his time as VP of Engineering at Zenly, the French social mapping app Snap acquired for nine figures in 2017.
The company's GitHub repository has accumulated 3,500 stars as of mid-2026, respectable for infrastructure tooling but not viral. Performance tables on ZML's blog show throughput comparisons across hardware: Gemma-4-26B and Qwen3.6-27B running on H100, AMD MI300X, Intel B70, Apple M3 Max, and TPU v6e. They also claim "up to 10x" speedup per user with something called DFlash speculative decoding, though benchmark claims in AI infrastructure tend to come with asterisks the size of footnotes.
The underlying compiler and runtime framework got a full rewrite in March 2026. Four major technical updates followed between then and July: the compiler overhaul, a cross-vendor monitoring tool, tokenization performance improvements, and the LLMD server itself. That's a lot of output for a 20-person team, registered in France since November 2023.
Crowded Space, Narrow Wedge

ZML is entering a market that's seen capital pile in at a pace that would make earlier infrastructure cycles look quaint. Inferact, which commercializes vLLM, raised $150 million in January at roughly $800 million valuation. RadixArk launched in May with $100 million to build on SGLang, backed by Accel. Then Baseten closed a $1.5 billion round in June at $13 billion valuation—just months after raising $300 million at $5 billion.
Against that backdrop, $20 million might look modest. But Morin isn't pitching managed inference at scale, at least not yet. He's pitching portability, which is a different wedge entirely.
In an interview with TechCrunch, Morin said ZML is co-designing with chipmakers and plans to support emerging European accelerators from Axelera, Fractile, Kalray, and others. That positioning—hardware-agnostic inference wrapped in European AI sovereignty themes—may resonate more in Paris than Palo Alto. Kima Ventures, backed by French telecom founder Xavier Niel, and LocalGlobe's involvement suggest investors see regulatory and enterprise tailwinds in Europe that favor multi-vendor infrastructure.
>commit, the early-stage open-source fund from Red River West, adds another dimension. ZML isn't open source, but it's building on open standards and positioning itself as the anti-lock-in option. Whether that's enough to compete with well-funded competitors remains to be seen.
The Investor Theory

The cap table reads like a bet on infrastructure pluralism. Harry Stebbings' 20VC has been vocal about backing developer tools that challenge incumbents. LocalGlobe, a London-based firm with a track record in European enterprise software, brings distribution muscle. AAL VC, based in Montreal, focuses on AI infrastructure and developer tools—perhaps seeing ZML as a hedge against NVIDIA-dominated inference economics.
But it's the strategic angels who tell the story investors want to hear. LeCun's public endorsement carries weight beyond capital; he's spent decades thinking about how AI systems scale. Hykes, who created Docker and later Dagger, built his reputation on making deployment less painful. His involvement lends credibility to ZML's claim that inference workloads need the same kind of abstraction that containers brought to application deployment.
Hugging Face's co-founders joining signals something else: interest from a company that's made portability and model distribution central to its $4.5 billion-plus valuation. If Hugging Face sees ZML as complementary rather than competitive, that's a vote of confidence.
What Happens Next

ZML's roadmap includes support for additional Chinese model families—DeepSeek, Kimi, GLM, MiniMax, StepFun are listed as "coming soon"—and Morin has indicated more releases are planned. The team is shipping fast, maybe faster than a 20-person startup usually can. Whether they can maintain that cadence while figuring out go-to-market remains the open question.
The free release strategy buys time, but not forever. Competitors like Baseten and Inferact have demonstrated demand for managed inference at scale, and enterprises pay for reliability and support, not just performance benchmarks. ZML will eventually need to choose: stay infrastructure-thin and charge for tooling, or build up the stack and compete on managed services.
For now, the company seems content to collect data and court chipmakers. With European regulators and enterprises increasingly concerned about concentration in the AI stack—concerns that appear to have intensified through 2025 and into 2026—ZML's timing might be better than its funding round size suggests. Hardware independence isn't just a technical feature; in some markets, it's becoming a regulatory requirement.
Whether that's enough to justify the bet LeCun and company just made, we'll find out when the usage data starts rolling in.
