ZML/LLMD 20260929.0
ZML released zml/llmd 20260929.0, adding support for deepseek-ai/DeepSeek-V4.1-Flash on NVIDIA and AMD GPUs, the project's first release supporting a frontier model. The release adds 2 model families,…
ZML released zml/llmd 20260929.0, adding support for deepseek-ai/DeepSeek-V4.1-Flash on NVIDIA and AMD GPUs, the project's first release supporting a frontier model. The release adds 2 model families,…
ZML, a Paris startup, released LLMD on July 8, a free, open-source inference server that runs LLaMA, Gemma, Qwen, and Mistral on NVIDIA CUDA, AMD ROCm, Google TPU, Intel oneAPI, and Apple Metal from a…
ZML released LLMD, a free inference server for open large language models that runs across Nvidia CUDA, AMD ROCm, Google TPU, Intel oneAPI and Apple Metal, aiming to decouple AI workloads from proprie…
OpenAI launched GPT-Live, a full-duplex voice model that enables real-time, interruptible conversations, moving AI interaction from transactional prompts to natural dialogue. Meta introduced Muse Imag…
ZML released ZML/LLMD, an inference server written in Zig that runs LLaMa, Gemma, Qwen, and Mistral LLMs on five architectures including NVIDIA CUDA, AMD ROCm, Google TPU, Intel oneAPI, and Apple Meta…
ZML launched LLMD, a free inference server that runs LLaMA, Gemma, Qwen, and Mistral models on NVIDIA, AMD, Google TPU, Intel, and Apple hardware from a single Docker image. Built in Zig and compiled …
Paris startup ZML, backed by AI pioneer Yann LeCun, released free software called ZML/LLMD that runs open-source language models across Nvidia, AMD, Google, Intel, and Apple chips, aiming to break Nvi…
ZML released ZML/LLMD on July 8, 2026 as an alpha inference server for LLaMa, Gemma, Qwen and Mistral models across NVIDIA CUDA, AMD ROCm, Google TPU, Intel oneAPI and Apple Metal targets, aiming to r…
French startup ZML announced a partnership with Scaleway, VSORA, and the Île-de-France region at VivaTech to build a European AI stack independent of Nvidia. ZML's open-source inference software compi…
French startup ZML released ZML/LLMD v2, a free open-source inference server that runs AI models across any hardware platform without code changes, aiming to break NVIDIA's CUDA lock-in. Endorsed by Y…
French AI startup ZML released a free inference server, ZML/LLMD, that enables open-source large language models to run at peak performance across multiple chip types including Nvidia, AMD, Google TPU…