LLMD: Run LLM Inference on Any Chip, One Docker Tag
ZML, a Paris startup, released LLMD on July 8, a free, open-source inference server that runs LLaMA, Gemma, Qwen, and Mistral on NVIDIA CUDA, AMD ROCm, Google TPU, Intel oneAPI, and Apple Metal from aβ¦