{"slug": "ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml", "title": "AI Development on Windows: From PyTorch and Llama.cpp to Windows ML", "summary": "Microsoft added experimental llama.cpp support to Windows ML, letting developers run GGUF models locally on Windows through new task-specific APIs, starting with a Windows ML Text Generation API that accepts both GGUF and ONNX formats. Microsoft said it contributed CUDA kernel optimization, kernel fusion, improved CPU-GPU scheduling, weight repacking and CUDA graphs to llama.cpp with NVIDIA, adding Eagle-3, MTP and D-Flash2 speculative decoding, multi-GPU execution, NVFP4, new model architectures and backend sampling. The updates also cover PyTorch and Triton and target Windows PCs powered by NVIDIA RTX Spark, such as Surface Laptop Ultra.", "body_md": "Today we are highlighting improvements across a few key open-source projects many of you already use, as well as updates to Windows ML. To empower developers, we must support the broad range of tools for experimentation and exploration across inference and training, in addition to our production grade native inference stack.\n\nAvailable today, Windows adds experimental llama.cpp support to Windows ML so you can run GGUF models locally through new task-specific APIs. Windows ML is the unified, high-performance local AI inferencing framework for Windows. Our experimental Windows-native Runtime API is now in preview for developers who want more control over how models run and compose.\n\nThese updates arrive with improvements across the open-source projects many of you already use, including PyTorch, and Triton. On a new generation of powerful Windows PCs powered by NVIDIA RTX Spark, like Surface Laptop Ultra, they give you more ways to run open-source models, build local agentic systems, and develop and optimize your own models on Windows.\n\n## Building with the llama.cpp open-source community\n\nWe are very excited about advancements from across the open-source community – and nothing more so than the work happening across the **GGUF** and **llama.cpp** ecosystem. On this class of device, great support for the open-source models and frameworks developers already use matters most – and more of it is coming to Windows, natively and across Windows on Arm, than ever before.\n\nWe’re bringing **GGUF support** to **Windows ML**. The GGUF community is moving fast, unlocking new use cases with the latest open-source models and workloads, and llama.cpp is how many developers run them first. With this new, experimental llama.cpp integration, you can pull a brand-new GGUF model from Hugging Face and run it locally through the same Windows ML stack you already use, in just a few lines of code.\n\nWe’re also contributing directly to llama.cpp. Together with **NVIDIA** and the broader community, we’ve delivered significant performance improvements – through CUDA kernel optimization, kernel fusion, improved CPU–GPU scheduling, weight repacking, and CUDA graphs. That work also adds Eagle-3, MTP, and D-Flash2 speculative decoding, multi-GPU execution, NVFP4, new model architectures, and backend sampling.\n\n**Explore:** [llama.cpp](https://github.com/ggml-org/llama.cpp)\n\n## New to Windows ML?\n\nFor those of you who are new to **Windows ML**, it is the **unified, high-performance local AI inferencing framework for Windows**. You can use it to run your own custom local AI workflows across Windows PCs spanning GPUs, NPUs, and CPUs from AMD, Intel, NVIDIA, and Qualcomm. Local inference can help reduce latency, keep workload data on the device, and avoid per-token cloud inference charges. It’s designed to meet you where your models already live – whether you’re bringing something you trained and exported yourself or a popular open-source model from the community.\n\nWindows ML also includes the [**Windows ML CLI**](https://aka.ms/winmlcli), a command-line tool and set of agent skills for getting your models ready to run with Windows ML. You can use it to convert, optimize, compile, and benchmark your models before you ship.\n\nMany app developers are already using the Windows ML stack to enable local AI in their applications:\n\nWith this release, Windows ML adds an additional, experimental **Windows-native inferencing path** that runs both ONNX and GGUF models, making it easier to experiment with popular open-source models alongside the ONNX workflows you already use.\n\n## Experiment with open-source models and GGUF using Windows ML\n\nSo how do you run a GGUF model on Windows ML? Windows now includes task-specific APIs, starting with the Windows ML **Text Generation API**. It accepts language models in both **GGUF** and **ONNX** formats through one simplified surface. Windows ML automatically selects the right execution engine for your model – including llama.cpp for GGUF – so you can bring your own model, get results back quickly, and focus on building your application instead of the underlying plumbing.\n\nFor the quickest way to experiment, these APIs also expose an **OpenAI-compatible endpoint**. You can prototype against a local, on-device model using the same OpenAI SDK you already know – no new API to learn.\n\n```\n# Start the local Windows ML server with your GGUF model, then point the OpenAI SDK at it:\n#   WinMLServer.exe model.gguf --model-id qwen2.5-0.5b --target gpu --port 8080\nfrom openai import OpenAI\n\nclient = OpenAI(\n    base_url=\"http://127.0.0.1:8080/v1\",\n    api_key=\"<access key printed on startup>\",\n)\n\nstream = client.chat.completions.create(\n    model=\"qwen2.5-0.5b\",\n    messages=[{\"role\": \"user\", \"content\": \"What workloads can I run locally on the powerful NVIDIA RTX GPU in my Surface Laptop Ultra?\"}],\n    stream=True,\n)\n\nfor chunk in stream:\n    print(chunk.choices[0].delta.content or \"\", end=\"\", flush=True)\n```\n\nThis first release ships two task-specific APIs: the **Text Generation API**, which runs your own GGUF or ONNX language model, and the **Speech Recognition API**, which transcribes audio with an ONNX Whisper model. You can also use them together – transcribe voice input with the Speech Recognition API, then pass that text to the Text Generation API running a GGUF model. Just supply the models, and Windows ML handles the rest.\n\n**Explore:** [Text Generation API docs](https://learn.microsoft.com/windows/ai/new-windows-ml/runtime/text-generation) · [Speech Recognition API docs](https://learn.microsoft.com/windows/ai/new-windows-ml/runtime/speech-recognition)\n\n## The Windows ML Runtime API: a new Windows-native inferencing path\n\nUnderneath those higher-level surfaces is the **Windows ML Runtime API** – a new, experimental Windows-native inferencing API for Windows ML, built for performance, deep OS integration, and precise control when you want it. The familiar **ONNX Runtime APIs** stay fully supported for broad compatibility, while the Runtime APIs are where deeper, Windows-native optimizations will land over time – and the two ship side by side, so you can start with what you know and adopt the native path when you’re ready.\n\nWith the Runtime APIs you can:\n\n- **Work with Windows-native data types.** Feed images, video frames, audio buffers, and text directly to models through efficient, zero-copy paths – instead of hand-writing preprocessing and format conversions.\n- **Compose deterministic multi-model pipelines.** Chain multiple models into a single pipeline with explicit, per-stage device placement across CPU, GPU, and NPU. Execution is reproducible and predictable run-to-run – which matters as apps grow richer, like an encoder feeding a decoder.\n- **Load and compile models ahead of time.** Turn a model into a ready-to-run artifact for faster startup, with your device and execution-policy choices applied consistently.\n\n**Drop down to the Runtime API when you need finer control.** The higher-level Text Generation and Speech Recognition APIs are built on it, so you can start high-level and reach for the Runtime primitives whenever you need to.\n\n**Explore:** [Windows ML Runtime API](https://learn.microsoft.com/windows/ai/new-windows-ml/runtime/overview) · [Runtime API samples](https://aka.ms/winml-runtime-samples)\n\n## Open-source AI frameworks for local development on Windows\n\nBeyond llama.cpp, Windows is gaining a wider open-source stack, from model libraries to the tooling that ties them together, with more of it running natively on Windows and Windows on Arm.\n\n### PyTorch and Triton on Windows on Arm\n\n**PyTorch** is the framework most developers reach for across AI development – from research and experimentation to building, training, and fine-tuning models. On Windows, it’s also a natural on-ramp to Windows ML: you can train or fine-tune your own models, including proprietary ones, and then bring them to Windows ML for local inference.\n\nPyTorch now offers official **native Windows Arm64** CPU builds, and NVIDIA publishes CUDA-enabled Windows Arm64 packages for supported hardware. This gives model developers a native foundation for training, fine-tuning, and inference on Arm-based Windows AI systems. The Windows distribution of Triton brings **triton.jit**, **torch.compile**, and custom GPU kernels to supported Windows GPUs. Windows Arm64 compiler and release work extends that path to the optimized kernels used throughout modern AI frameworks.\n\n**Explore:** [PyTorch Arm native builds](https://blogs.windows.com/windowsdeveloper/2025/04/23/pytorch-arm-native-builds-now-available-for-windows/) · [NVIDIA Windows Arm64 PyTorch packages](https://pypi.nvidia.com/nvtorch_oot/) · [Triton for Windows](https://github.com/triton-lang/triton-windows)\n\n### Build with native PyTorch and Triton, then run with Windows ML\n\n**PyTorch** and **Triton** help us complete the full model lifecycle path on Windows. In this example, PyTorch loads a real vision model**, torch.compile** uses Triton-generated GPU kernels on Windows, the original model is exported to the open **ONNX format** and prepared for deployment in an app.\n\n#### Set up the native environment with PyTorch and Triton\n\n```\npymanager install 3.14-arm64\npymanager exec -V:3.14-arm64 -m venv .venv\n.\\.venv\\Scripts\\Activate.ps1\n\npython -m pip install --extra-index-url https://pypi.nvidia.com/nvtorch_oot `\n    \"torch==2.14.0+cu134\" \"torchvision==0.29.0+cu134\"\npython -m pip install \"triton-windows==3.8.0.post29\" `\n    onnx onnxscript numpy\n```\n\nThe benefit of Triton is visible when PyTorch Inductor combines a chain of GPU operations into a generated kernel. This small activation block creates several eager operations but can be fused by torch.compile for significant performance gains you can see comparing eager to triton:\n\n``` python\nimport torch\n\ndef activation_block(x):\n    return torch.nn.functional.silu(x * 1.5 + 0.25).square()\n\ndef time_ms(fn, x, iterations=100):\n    for _ in range(20):\n        fn(x)\n    torch.cuda.synchronize()\n    start = torch.cuda.Event(enable_timing=True)\n    end = torch.cuda.Event(enable_timing=True)\n    start.record()\n    for _ in range(iterations):\n        fn(x)\n    end.record()\n    torch.cuda.synchronize()\n    return start.elapsed_time(end) / iterations\n\nx = torch.randn(4096, 4096, device=\"cuda\")\ntriton_block = torch.compile(activation_block, backend=\"inductor\")\ntorch.testing.assert_close(\n    triton_block(x), activation_block(x), rtol=1e-3, atol=1e-3\n)\n\neager_ms = time_ms(activation_block, x)\ntriton_ms = time_ms(triton_block, x)\nprint(f\"Eager:  {eager_ms:.3f} ms\")\nprint(f\"Triton: {triton_ms:.3f} ms\")\nprint(f\"Speedup: {eager_ms / triton_ms:.2f}x\")\n```\n\nPyTorch remains the programming model while Inductor generates specialized Triton GPU code to be used while working on the model efficiently.\n\n#### Export the portable ONNX model\n\nThe model graph—not the generated kernel—is exported for deployment.\n\n``` python\nimport torch\nfrom torchvision.models import resnet18\n\nmodel = resnet18(weights=None).eval()\nimage = torch.randn(1, 3, 224, 224)\n\ntorch.onnx.export(\n    model,\n    (image,),\n    \"resnet18.onnx\",\n    dynamo=True,\n    external_data=False,\n    input_names=[\"image\"],\n    output_names=[\"scores\"],\n)\n```\n\n#### Build once, benchmark, and ship\n\nThe Windows ML CLI exposes analyze, optimize, quantize, and compile to developers to help them prepare models for their app. This example uses an auto-generated configuration to compose the applicable stages, check the target execution provider, apply compatible graph rewrites and operator fusions, and produce the model and build configuration for the app.\n\n```\nuv python install cpython-3.11-windows-x86_64-none\nuv venv .winml-cli --python cpython-3.11-windows-x86_64-none\nuv pip install --python .winml-cli\\Scripts\\python.exe winml-cli==0.3.1\n.\\.winml-cli\\Scripts\\Activate.ps1\n\nwinml analyze `\n    -m .\\resnet18.onnx `\n    --device gpu `\n    --ep nv_tensorrt_rtx\n\nwinml build `\n    -m .\\resnet18.onnx `\n    -o .\\app-model `\n    --device gpu `\n    --ep nv_tensorrt_rtx `\n    --no-quant\n\nwinml perf `\n    -m .\\app-model\\model.onnx `\n    --device gpu `\n    --ep nv_tensorrt_rtx\n```\n\n#### What’s next: Building momentum with the open-source community\n\nWindows is continuing to work with open-source maintainers, hardware partners, and the Python and AI communities to improve local AI development. Priorities include more native packages, broader kernel coverage, simpler installation, better performance, and clearer support across Windows x64 and Windows on Arm.\n\nYou can follow Windows Arm64 Python package compatibility on [PyEnv-WoA-State](https://github.com/khmyznikov/PyEnv-WoA-State), which links to the upstream issues and pull requests where progress is happening. Use the [WoA Libs tracker](https://github.com/khmyznikov/PyEnv-WoA-State/issues/1) for the latest status and to find ways to contribute.\n\n## Get started\n\nThese updates give you more ways to build and experiment with local AI on Windows using the open-source models and frameworks you already know. The new Windows ML capabilities are experimental, so review the supported scenarios and known limitations before relying on them in production.\n\n- **Train models with PyTorch and Triton** and run them with[Windows ML](https://learn.microsoft.com/windows/ai/new-windows-ml/overview) , as seen in the example above.\n- **Run a GGUF model:** Try the new Windows ML[Text Generation API](https://learn.microsoft.com/windows/ai/new-windows-ml/runtime/text-generation) , powered with[llama.cpp](https://github.com/ggml-org/llama.cpp) .\n- **Explore the native runtime:** Build and experiment with the new[Windows ML Runtime APIs](https://learn.microsoft.com/windows/ai/new-windows-ml/runtime/overview) by installing[Windows ML 2.7.2021 Experimental](https://www.nuget.org/packages/Microsoft.Windows.AI.MachineLearning/2.7.2021-experimental) .\n- **Prepare and compare models:** Use the[Windows ML CLI](https://aka.ms/winmlcli) to convert, optimize, and benchmark your own models for local inferencing with Windows ML.\n- **Share feedback on** [Windows ML GitHub](https://github.com/microsoft/WindowsML)**:** Tell us which models, devices, and workflows you want us to support next.\n\nFrom running GGUF models with llama.cpp in Windows ML to tuning models with PyTorch and Triton, and more work happening across the open-source ecosystem, there are more ways than ever to run open-source models, build local agentic workloads, and develop and optimize your own models on Windows. Bring your own model and workload, try it on your target Windows devices, including new RTX Spark Windows PCs, and tell us what you’d like to see next.", "url": "https://wpnews.pro/news/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml", "canonical_source": "https://devblogs.microsoft.com/foundry-on-windows/build-on-winml-oct-7-26/", "published_at": "2026-10-07 19:40:34+00:00", "updated_at": "2026-10-07 19:49:46.969997+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-infrastructure", "ai-tools", "developer-tools"], "entities": ["Microsoft", "Windows ML", "llama.cpp", "GGUF", "PyTorch", "Triton", "NVIDIA", "Surface Laptop Ultra"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml", "markdown": "https://wpnews.pro/news/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml.md", "text": "https://wpnews.pro/news/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml.txt", "jsonld": "https://wpnews.pro/news/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml.jsonld"}}