AI Development on Windows: From PyTorch and Llama.cpp to Windows ML Microsoft added experimental llama.cpp support to Windows ML, letting developers run GGUF models locally on Windows through new task-specific APIs, starting with a Windows ML Text Generation API that accepts both GGUF and ONNX formats. Microsoft said it contributed CUDA kernel optimization, kernel fusion, improved CPU-GPU scheduling, weight repacking and CUDA graphs to llama.cpp with NVIDIA, adding Eagle-3, MTP and D-Flash2 speculative decoding, multi-GPU execution, NVFP4, new model architectures and backend sampling. The updates also cover PyTorch and Triton and target Windows PCs powered by NVIDIA RTX Spark, such as Surface Laptop Ultra. Today we are highlighting improvements across a few key open-source projects many of you already use, as well as updates to Windows ML. To empower developers, we must support the broad range of tools for experimentation and exploration across inference and training, in addition to our production grade native inference stack. Available today, Windows adds experimental llama.cpp support to Windows ML so you can run GGUF models locally through new task-specific APIs. Windows ML is the unified, high-performance local AI inferencing framework for Windows. Our experimental Windows-native Runtime API is now in preview for developers who want more control over how models run and compose. These updates arrive with improvements across the open-source projects many of you already use, including PyTorch, and Triton. On a new generation of powerful Windows PCs powered by NVIDIA RTX Spark, like Surface Laptop Ultra, they give you more ways to run open-source models, build local agentic systems, and develop and optimize your own models on Windows. Building with the llama.cpp open-source community We are very excited about advancements from across the open-source community – and nothing more so than the work happening across the GGUF and llama.cpp ecosystem. On this class of device, great support for the open-source models and frameworks developers already use matters most – and more of it is coming to Windows, natively and across Windows on Arm, than ever before. We’re bringing GGUF support to Windows ML . The GGUF community is moving fast, unlocking new use cases with the latest open-source models and workloads, and llama.cpp is how many developers run them first. With this new, experimental llama.cpp integration, you can pull a brand-new GGUF model from Hugging Face and run it locally through the same Windows ML stack you already use, in just a few lines of code. We’re also contributing directly to llama.cpp. Together with NVIDIA and the broader community, we’ve delivered significant performance improvements – through CUDA kernel optimization, kernel fusion, improved CPU–GPU scheduling, weight repacking, and CUDA graphs. That work also adds Eagle-3, MTP, and D-Flash2 speculative decoding, multi-GPU execution, NVFP4, new model architectures, and backend sampling. Explore: llama.cpp https://github.com/ggml-org/llama.cpp New to Windows ML? For those of you who are new to Windows ML , it is the unified, high-performance local AI inferencing framework for Windows . You can use it to run your own custom local AI workflows across Windows PCs spanning GPUs, NPUs, and CPUs from AMD, Intel, NVIDIA, and Qualcomm. Local inference can help reduce latency, keep workload data on the device, and avoid per-token cloud inference charges. It’s designed to meet you where your models already live – whether you’re bringing something you trained and exported yourself or a popular open-source model from the community. Windows ML also includes the Windows ML CLI https://aka.ms/winmlcli , a command-line tool and set of agent skills for getting your models ready to run with Windows ML. You can use it to convert, optimize, compile, and benchmark your models before you ship. Many app developers are already using the Windows ML stack to enable local AI in their applications: With this release, Windows ML adds an additional, experimental Windows-native inferencing path that runs both ONNX and GGUF models, making it easier to experiment with popular open-source models alongside the ONNX workflows you already use. Experiment with open-source models and GGUF using Windows ML So how do you run a GGUF model on Windows ML? Windows now includes task-specific APIs, starting with the Windows ML Text Generation API . It accepts language models in both GGUF and ONNX formats through one simplified surface. Windows ML automatically selects the right execution engine for your model – including llama.cpp for GGUF – so you can bring your own model, get results back quickly, and focus on building your application instead of the underlying plumbing. For the quickest way to experiment, these APIs also expose an OpenAI-compatible endpoint . You can prototype against a local, on-device model using the same OpenAI SDK you already know – no new API to learn. Start the local Windows ML server with your GGUF model, then point the OpenAI SDK at it: WinMLServer.exe model.gguf --model-id qwen2.5-0.5b --target gpu --port 8080 from openai import OpenAI client = OpenAI base url="http://127.0.0.1:8080/v1", api key="