Observed Signal · Jul 8, 2026 · Product Launch · Source: techcrunch · Impact: 3/5 · Sentiment: Positive
ZML Launches Free LLMD Inference Server
ZML, a Paris-based AI startup endorsed by Yann LeCun, has released LLMD, an inference-performance server that aims to run open-source large language models efficiently across many chip types (including Nvidia, AMD, Google TPU, Apple Metal and Intel Arc). The closed-source product is launching free to gather usage data; ZML says the goal is to avoid vendor lock-in, enable mixed-chip deployments, and reduce inference cost and energy for enterprises and clouds. Founder Steeve Morin highlighted co-design work with chipmakers and said the 20-person startup — backed by a $20 million seed round from multiple VCs — plans further releases. The move positions ZML as a competitor in the inference market alongside firms such as Baseten, Inferact and RadixArk, and could help accelerate adoption of non‑Nvidia AI chips.
A startup product enabling multi‑chip inference can reduce vendor lock-in and inference cost, affecting AI infrastructure choices for enterprises and cloud providers; notable but not platform-level policy or major platform technical release.
Track AMD Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- ZML released LLMD, an LLM inference server designed to run models across multiple chip types.
- LLMD supports a variety of chips including Nvidia, AMD, Google TPU, Apple Metal and Intel Arc.
- ZML is a Paris-based startup with about 20 employees and was founded by Steeve Morin.
- ZML has raised $20 million from investors including Harry Stebbings' 20VC, >commit, AALVC, Drysdale Ventures, Kima Ventures, Kindred Capital, LocalGlobe and Puzzle Ventures.
- LLMD is not open source but is launching as a free product to learn about usage.
Connected Companies & Entities
9 Entities mapped“has released inference-performance software that allows a variety of open source large language models to run on a variety of chips — includ...”
“has released inference-performance software that allows a variety of open source large language models to run on a variety of chips — includ...”
“has released inference-performance software that allows a variety of open source large language models to run on a variety of chips — includ...”
“has released inference-performance software that allows a variety of open source large language models to run on a variety of chips — includ...”
“So ZML has competition such as Baseten, recently valued at $13 billion;...”
“the startup’s cap table confirms that other founders are paying attention, including Dagger and Docker founder Solomon Hykes...”
“the startup’s cap table confirms that other founders are paying attention, including ... Clément Delangue and Julien Chaumond from Hugging F...”
“Morin raised $20 million from venture firms including Harry Stebbings’ 20VC, >commit, AALVC, Drysdale Ventures, Xavier Niel’s Kima Ventures,...”
“Zenly, which Snapchat acquired for nine figures in 2017...”
Ontology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
OpenAI and Broadcom unveil Jalapeño inference chip
OpenAI and Broadcom announced Jalapeño, an LLM-optimized accelerator described as OpenAI’s first 'Intelligence Processor' and the initial component of a multi-generation compute platform co-developed with Broadcom and Celestica. OpenAI says the chip was designed from the ground up for modern and future LLM inference, produced from design to tape-out in nine months with assistance from OpenAI models, and that engineering samples are running ML workloads (including GPT‑5.3‑Codex‑Spark). Early internal testing reportedly shows performance per watt substantially better than current state-of-the-art; a detailed technical report will follow. The partners plan gigawatt-scale deployments with data center partners (including Microsoft) beginning in 2026, aiming to lower inference cost, latency, and improve reliability for large-scale interactive LLM products.
Gimlet Labs raises $80M for multi-silicon inference cloud
Gimlet Labs, founded by Stanford adjunct professor and founder Zain Asgar with cofounders Michelle Nguyen, Omid Azizi and Natalie Serrino, raised an $80 million Series A led by Menlo Ventures to commercialize what it calls a "multi-silicon inference cloud." The software orchestrates AI workloads across diverse hardware (CPUs, AI GPUs, high-memory systems), claims 3x–10x inference speedups for the same cost and power, and can slice models to run across different architectures. Gimlet has partnerships with chip makers NVIDIA, AMD, Intel, ARM, Cerebras and d‑Matrix, offers its product as software or via an API/Gimlet Cloud, and targets large model labs and data centers. The company reported eight-figure revenues at launch, has roughly 30 employees, and has now raised $92 million in total including prior seed and angel investments.
Z.ai Says AI Model Built Its Own Inference Infrastructure
Chinese AI company Z.ai (formerly Zhipu AI) published a research paper detailing how its GLM-5.3 model, via an Infra Agent, built and optimized the production inference infrastructure on a cluster of over 100,000 Chinese-made AI accelerators. The process from model adaptation to production readiness took under two weeks, with end-to-end throughput tripling. The company reports performance comparable to Nvidia GPUs and introduced 'Dense Feedback,' where an AI agent uses system metrics to autonomously identify and fix bottlenecks, such as reducing a parallelism bottleneck from 20% to under 1%. While not yet achieving full recursive self-improvement (RSI), Z.ai sees early forms of it. Unconfirmed rumors suggest Google DeepMind may have reached RSI, but Google has not commented. The event occurred in September 2026.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
