ZML raises $20M to run LLMs on any chip
What's the deal? French startup ZMLDealroom has a profile for this one. Try Dealroom → has raised $20 million in an early-stage round to build a chip-agnostic inference server. Founder Steeve Morin, former VP of engineering at Snap-acquired Zenly, drew backers including 20VC, Kima Ventures, and LocalGlobeDealroom has a profile for this one. Try Dealroom →.
The cap table: Beyond the lead investors, ZML pulled in Turing Award winner Yann LeCun, Solomon Hykes, and Hugging Face's Clément Delangue and Julien Chaumond. The round ranks among the largest early-stage enterprise software deals in France, sitting in the top 1% of comparable rounds on record.
What's the endgame? ZML this week launched LLMD, a single inference server that reportedly runs open source large language models across Nvidia, AMD, Google TPU, Apple Metal, and Intel Arc — without forcing customers to commit to one chip vendor. Morin's framing is to give people "the power to create their own system" and achieve efficiency gains.
Why now? Capacity is the pull. Nvidia allocation is the practical constraint on most serious inference workloads, and teams are re-pricing everything against whatever accelerators they can secure. If LLMD delivers across AMD, TPU, and Apple Metal without a full port per backend, the math shifts for any lab under GPU rationing.
What could go wrong? LLMD is not open source; it launches free, aimed at learning how people use it — a distribution choice, not yet a business. The market is crowded: Baseten, recently valued at $13 billion, is the incumbent teams benchmark against, while Inferact and RadixArk commercialise the vLLM and SGLang runtimes.
The signal: Chip-agnostic matters only if per-backend performance stays close to a chip-specific runtime — and no head-to-head benchmarks exist yet. But the pitch targets a real audience: teams holding AMD or TPU allocation they cannot fully exploit, hunting for a discount on the Nvidia-scarcity tax.
Read more: AI Weekly