{"slug": "windows-ml-gets-llama-cpp-run-gguf-locally-on-any-hardware", "title": "Windows ML Gets llama.cpp: Run GGUF Locally on Any Hardware", "summary": "Microsoft announced llama.cpp support in Windows ML on October 7, letting developers run any GGUF model through a single WinMLServer.exe executable that exposes an OpenAI-compatible endpoint and automatically routes inference across GPU, NPU, or CPU. Windows ML's execution-provider selection covers QNN for Qualcomm NPUs, OpenVINO for Intel NPUs, NvTensorRtRtx for NVIDIA GPUs, and CPU fallback, a capability Microsoft says Ollama and LM Studio lack. Microsoft also contributed CUDA kernel optimizations, speculative decoding, and multi-GPU execution to the upstream llama.cpp project with NVIDIA, and highlighted DeepSeek V4 Flash (284B parameters, 1.6-bit quantized to roughly 60GB) and NVIDIA Nemotron 70B+ (2-bit quantized, under 20GB, launching October 15).", "body_md": "Microsoft just turned Windows into a first-class local AI runtime. On October 7, the company announced llama.cpp support in Windows ML — developers can now point any GGUF model at a single executable and get an OpenAI-compatible endpoint back, with the OS handling hardware routing across GPU, NPU, or CPU automatically. No separate server setup. No manual CUDA configuration. One command.\n\n## What Changed in Windows ML\n\nWindows ML has been Microsoft’s ONNX inference runtime for Windows since 2018 — solid, but limited to ONNX models. That changed today. The new experimental [Text Generation API](https://daily.dev/posts/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml-f2t66cns3) accepts both ONNX and GGUF through the same surface, automatically routing GGUF through llama.cpp and ONNX through the existing runtime. From the developer’s perspective: one API, two model formats.\n\nThe practical entry point is `WinMLServer.exe`, a new command-line server that spins up an OpenAI-compatible endpoint from any GGUF file:\n\n```\nWinMLServer.exe model.gguf --model-id qwen2.5-0.5b --target gpu --port 8080\n```\n\nPoint the standard OpenAI Python SDK at `http://localhost:8080/v1` and your existing inference code runs locally without modification:\n\n``` python\nfrom openai import OpenAI\n\nclient = OpenAI(base_url=\"http://localhost:8080/v1\", api_key=\"not-needed\")\nresponse = client.chat.completions.create(\n    model=\"qwen2.5-0.5b\",\n    messages=[{\"role\": \"user\", \"content\": \"Summarize this code\"}]\n)\n```\n\nThe `--target` flag accepts `gpu`, `npu`, or `cpu`. If the target is unavailable, Windows ML falls back gracefully rather than failing.\n\n## NPU Routing Is the Real Differentiator\n\nThe comparison with existing tools is where this gets interesting. Ollama and LM Studio both run on Windows and both expose OpenAI-compatible endpoints — but neither routes inference to the NPU. Windows ML does.\n\n| Tool | NPU Support | OpenAI Endpoint | OS-Managed | \n|---|---|---|---|\n| Ollama (Windows) | No | Yes | No | \n| LM Studio | No | Yes | No | \n| Windows ML + llama.cpp | Yes | Yes | Yes | \n\nWindows ML auto-selects the right execution provider based on available hardware: QNN for Qualcomm NPUs, OpenVINO for Intel NPUs, NvTensorRtRtx for NVIDIA GPUs, and CPU as fallback. According to the [official announcement](https://blogs.windows.com/windowsexperience/2026/10/07/building-windows-for-hybrid-intelligence/), the platform targets sustained, battery-efficient inference on NPU-equipped devices — which matters for developers running agent workflows that need to stay on all day without saturating the GPU. Copilot+ PC owners sitting on idle NPU silicon now have a practical inference target.\n\nMicrosoft also contributed CUDA kernel optimizations, speculative decoding, and multi-GPU execution back to the upstream llama.cpp project alongside NVIDIA. The integration isn’t a fork; it’s committed code in the main project.\n\n## Models You Can Run Today\n\nAny GGUF file from Hugging Face works with WinMLServer. Microsoft highlighted two flagship models tuned for local Windows inference:\n\n- **DeepSeek V4 Flash** — 284B parameters, 1.6-bit quantized to fit in roughly 60GB. Targets the RTX Spark platform for high-end local inference.\n- **NVIDIA Nemotron 70B+** — 2-bit quantized, under 20GB, launching October 15. The practical choice for most Windows machines with a modern NVIDIA GPU.\n\nFor developers not on cutting-edge hardware, smaller GGUF models — Qwen 2.5, Phi-4, Llama 3.2 — run fine on standard GPUs and the path is identical. Download the GGUF, run WinMLServer, point your SDK at localhost.\n\n## What’s Coming: HydraFusion Goes Local\n\nGitHub’s HydraFusion — the multi-model routing engine that splits tasks by complexity and cost — is being extended to Windows to route between local and cloud models. It’s arriving in VS Code, the GitHub Copilot app, and Copilot CLI in experimental preview later this month.\n\nWhen it lands, the workflow changes: instead of manually deciding which model to call, the router assesses each task and sends routine work to local Windows ML inference while complex tasks go to cloud APIs. For developers paying per-token, this is worth tracking. The [execution providers documentation](https://learn.microsoft.com/en-us/windows/ai/new-windows-ml/supported-execution-providers) already covers the hardware targeting that HydraFusion will use for the local leg.\n\n## What Developers Should Do Now\n\nWindows ML’s llama.cpp support is experimental — available to test but not production-ready. For developers on Windows 11:\n\n1. Check the [Windows ML GitHub](https://github.com/microsoft/WindowsML) for the latest experimental release\n2. Download a small GGUF model (Qwen 2.5 0.5B is ~400MB — reasonable for first tests)\n3. Run `WinMLServer.exe` and point your existing OpenAI SDK code at the local endpoint\n4. On a Copilot+ PC, try `--target npu` and compare throughput against`--target gpu`\n\nThe [Microsoft Execution Containers (MXC)](https://blogs.windows.com/windowsdeveloper/2026/10/07/microsoft-execution-containers-policy-driven-containment-for-ai-agents/) sandbox also shipped today — if you’re building agent workflows that run locally, it’s worth understanding before deploying autonomous code on Windows.", "url": "https://wpnews.pro/news/windows-ml-gets-llama-cpp-run-gguf-locally-on-any-hardware", "canonical_source": "https://byteiota.com/windows-ml-gets-llama-cpp-run-gguf-locally-on-any-hardware/", "published_at": "2026-10-07 20:12:02+00:00", "updated_at": "2026-10-07 20:18:58.866393+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-infrastructure", "developer-tools", "ai-products"], "entities": ["Microsoft", "Windows ML", "llama.cpp", "WinMLServer.exe", "NVIDIA", "DeepSeek V4 Flash", "NVIDIA Nemotron 70B+", "GitHub HydraFusion"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/windows-ml-gets-llama-cpp-run-gguf-locally-on-any-hardware", "markdown": "https://wpnews.pro/news/windows-ml-gets-llama-cpp-run-gguf-locally-on-any-hardware.md", "text": "https://wpnews.pro/news/windows-ml-gets-llama-cpp-run-gguf-locally-on-any-hardware.txt", "jsonld": "https://wpnews.pro/news/windows-ml-gets-llama-cpp-run-gguf-locally-on-any-hardware.jsonld"}}