cd /news/artificial-intelligence/ai-development-on-windows-from-pytor… · home › topics › artificial-intelligence › article
[ARTICLE · art-147114] src=devblogs.microsoft.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

AI Development on Windows: From PyTorch and Llama.cpp to Windows ML

Microsoft added experimental llama.cpp support to Windows ML, letting developers run GGUF models locally on Windows through new task-specific APIs, starting with a Windows ML Text Generation API that accepts both GGUF and ONNX formats. Microsoft said it contributed CUDA kernel optimization, kernel fusion, improved CPU-GPU scheduling, weight repacking and CUDA graphs to llama.cpp with NVIDIA, adding Eagle-3, MTP and D-Flash2 speculative decoding, multi-GPU execution, NVFP4, new model architectures and backend sampling. The updates also cover PyTorch and Triton and target Windows PCs powered by NVIDIA RTX Spark, such as Surface Laptop Ultra.

by read10 min views2 publishedOct 7, 2026
AI Development on Windows: From PyTorch and Llama.cpp to Windows ML
Image: Devblogs (auto-discovered)

Today we are highlighting improvements across a few key open-source projects many of you already use, as well as updates to Windows ML. To empower developers, we must support the broad range of tools for experimentation and exploration across inference and training, in addition to our production grade native inference stack.

Available today, Windows adds experimental llama.cpp support to Windows ML so you can run GGUF models locally through new task-specific APIs. Windows ML is the unified, high-performance local AI inferencing framework for Windows. Our experimental Windows-native Runtime API is now in preview for developers who want more control over how models run and compose.

These updates arrive with improvements across the open-source projects many of you already use, including PyTorch, and Triton. On a new generation of powerful Windows PCs powered by NVIDIA RTX Spark, like Surface Laptop Ultra, they give you more ways to run open-source models, build local agentic systems, and develop and optimize your own models on Windows.

Building with the llama.cpp open-source community #

We are very excited about advancements from across the open-source community – and nothing more so than the work happening across the GGUF and llama.cpp ecosystem. On this class of device, great support for the open-source models and frameworks developers already use matters most – and more of it is coming to Windows, natively and across Windows on Arm, than ever before.

We’re bringing GGUF support to Windows ML. The GGUF community is moving fast, unlocking new use cases with the latest open-source models and workloads, and llama.cpp is how many developers run them first. With this new, experimental llama.cpp integration, you can pull a brand-new GGUF model from Hugging Face and run it locally through the same Windows ML stack you already use, in just a few lines of code.

We’re also contributing directly to llama.cpp. Together with NVIDIA and the broader community, we’ve delivered significant performance improvements – through CUDA kernel optimization, kernel fusion, improved CPU–GPU scheduling, weight repacking, and CUDA graphs. That work also adds Eagle-3, MTP, and D-Flash2 speculative decoding, multi-GPU execution, NVFP4, new model architectures, and backend sampling.

Explore: llama.cpp

New to Windows ML? #

For those of you who are new to Windows ML, it is the unified, high-performance local AI inferencing framework for Windows. You can use it to run your own custom local AI workflows across Windows PCs spanning GPUs, NPUs, and CPUs from AMD, Intel, NVIDIA, and Qualcomm. Local inference can help reduce latency, keep workload data on the device, and avoid per-token cloud inference charges. It’s designed to meet you where your models already live – whether you’re bringing something you trained and exported yourself or a popular open-source model from the community.

Windows ML also includes the Windows ML CLI, a command-line tool and set of agent skills for getting your models ready to run with Windows ML. You can use it to convert, optimize, compile, and benchmark your models before you ship.

Many app developers are already using the Windows ML stack to enable local AI in their applications:

With this release, Windows ML adds an additional, experimental Windows-native inferencing path that runs both ONNX and GGUF models, making it easier to experiment with popular open-source models alongside the ONNX workflows you already use.

Experiment with open-source models and GGUF using Windows ML #

So how do you run a GGUF model on Windows ML? Windows now includes task-specific APIs, starting with the Windows ML Text Generation API. It accepts language models in both GGUF and ONNX formats through one simplified surface. Windows ML automatically selects the right execution engine for your model – including llama.cpp for GGUF – so you can bring your own model, get results back quickly, and focus on building your application instead of the underlying plumbing.

For the quickest way to experiment, these APIs also expose an OpenAI-compatible endpoint. You can prototype against a local, on-device model using the same OpenAI SDK you already know – no new API to learn.

from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:8080/v1",
    api_key="<access key printed on startup>",
)

stream = client.chat.completions.create(
    model="qwen2.5-0.5b",
    messages=[{"role": "user", "content": "What workloads can I run locally on the powerful NVIDIA RTX GPU in my Surface Laptop Ultra?"}],
    stream=True,
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="", flush=True)

This first release ships two task-specific APIs: the Text Generation API, which runs your own GGUF or ONNX language model, and the Speech Recognition API, which transcribes audio with an ONNX Whisper model. You can also use them together – transcribe voice input with the Speech Recognition API, then pass that text to the Text Generation API running a GGUF model. Just supply the models, and Windows ML handles the rest.

Explore: Text Generation API docs · Speech Recognition API docs

The Windows ML Runtime API: a new Windows-native inferencing path #

Underneath those higher-level surfaces is the Windows ML Runtime API – a new, experimental Windows-native inferencing API for Windows ML, built for performance, deep OS integration, and precise control when you want it. The familiar ONNX Runtime APIs stay fully supported for broad compatibility, while the Runtime APIs are where deeper, Windows-native optimizations will land over time – and the two ship side by side, so you can start with what you know and adopt the native path when you’re ready.

With the Runtime APIs you can:

  • Work with Windows-native data types. Feed images, video frames, audio buffers, and text directly to models through efficient, zero-copy paths – instead of hand-writing preprocessing and format conversions.
  • Compose deterministic multi-model pipelines. Chain multiple models into a single pipeline with explicit, per-stage device placement across CPU, GPU, and NPU. Execution is reproducible and predictable run-to-run – which matters as apps grow richer, like an encoder feeding a decoder.
  • Load and compile models ahead of time. Turn a model into a ready-to-run artifact for faster startup, with your device and execution-policy choices applied consistently.

Drop down to the Runtime API when you need finer control. The higher-level Text Generation and Speech Recognition APIs are built on it, so you can start high-level and reach for the Runtime primitives whenever you need to.

Explore: Windows ML Runtime API · Runtime API samples

Open-source AI frameworks for local development on Windows #

Beyond llama.cpp, Windows is gaining a wider open-source stack, from model libraries to the tooling that ties them together, with more of it running natively on Windows and Windows on Arm.

PyTorch and Triton on Windows on Arm

PyTorch is the framework most developers reach for across AI development – from research and experimentation to building, training, and fine-tuning models. On Windows, it’s also a natural on-ramp to Windows ML: you can train or fine-tune your own models, including proprietary ones, and then bring them to Windows ML for local inference.

PyTorch now offers official native Windows Arm64 CPU builds, and NVIDIA publishes CUDA-enabled Windows Arm64 packages for supported hardware. This gives model developers a native foundation for training, fine-tuning, and inference on Arm-based Windows AI systems. The Windows distribution of Triton brings triton.jit, torch.compile, and custom GPU kernels to supported Windows GPUs. Windows Arm64 compiler and release work extends that path to the optimized kernels used throughout modern AI frameworks.

Explore: PyTorch Arm native builds · NVIDIA Windows Arm64 PyTorch packages · Triton for Windows

Build with native PyTorch and Triton, then run with Windows ML

PyTorch and Triton help us complete the full model lifecycle path on Windows. In this example, PyTorch loads a real vision model**, torch.compile** uses Triton-generated GPU kernels on Windows, the original model is exported to the open ONNX format and prepared for deployment in an app.

Set up the native environment with PyTorch and Triton

pymanager install 3.14-arm64
pymanager exec -V:3.14-arm64 -m venv .venv
.\.venv\Scripts\Activate.ps1

python -m pip install --extra-index-url https://pypi.nvidia.com/nvtorch_oot `
    "torch==2.14.0+cu134" "torchvision==0.29.0+cu134"
python -m pip install "triton-windows==3.8.0.post29" `
    onnx onnxscript numpy

The benefit of Triton is visible when PyTorch Inductor combines a chain of GPU operations into a generated kernel. This small activation block creates several eager operations but can be fused by torch.compile for significant performance gains you can see comparing eager to triton:

import torch

def activation_block(x):
    return torch.nn.functional.silu(x * 1.5 + 0.25).square()

def time_ms(fn, x, iterations=100):
    for _ in range(20):
        fn(x)
    torch.cuda.synchronize()
    start = torch.cuda.Event(enable_timing=True)
    end = torch.cuda.Event(enable_timing=True)
    start.record()
    for _ in range(iterations):
        fn(x)
    end.record()
    torch.cuda.synchronize()
    return start.elapsed_time(end) / iterations

x = torch.randn(4096, 4096, device="cuda")
triton_block = torch.compile(activation_block, backend="inductor")
torch.testing.assert_close(
    triton_block(x), activation_block(x), rtol=1e-3, atol=1e-3
)

eager_ms = time_ms(activation_block, x)
triton_ms = time_ms(triton_block, x)
print(f"Eager:  {eager_ms:.3f} ms")
print(f"Triton: {triton_ms:.3f} ms")
print(f"Speedup: {eager_ms / triton_ms:.2f}x")

PyTorch remains the programming model while Inductor generates specialized Triton GPU code to be used while working on the model efficiently.

Export the portable ONNX model

The model graph—not the generated kernel—is exported for deployment.

import torch
from torchvision.models import resnet18

model = resnet18(weights=None).eval()
image = torch.randn(1, 3, 224, 224)

torch.onnx.export(
    model,
    (image,),
    "resnet18.onnx",
    dynamo=True,
    external_data=False,
    input_names=["image"],
    output_names=["scores"],
)

Build once, benchmark, and ship

The Windows ML CLI exposes analyze, optimize, quantize, and compile to developers to help them prepare models for their app. This example uses an auto-generated configuration to compose the applicable stages, check the target execution provider, apply compatible graph rewrites and operator fusions, and produce the model and build configuration for the app.

uv python install cpython-3.11-windows-x86_64-none
uv venv .winml-cli --python cpython-3.11-windows-x86_64-none
uv pip install --python .winml-cli\Scripts\python.exe winml-cli==0.3.1
.\.winml-cli\Scripts\Activate.ps1

winml analyze `
    -m .\resnet18.onnx `
    --device gpu `
    --ep nv_tensorrt_rtx

winml build `
    -m .\resnet18.onnx `
    -o .\app-model `
    --device gpu `
    --ep nv_tensorrt_rtx `
    --no-quant

winml perf `
    -m .\app-model\model.onnx `
    --device gpu `
    --ep nv_tensorrt_rtx

What’s next: Building momentum with the open-source community

Windows is continuing to work with open-source maintainers, hardware partners, and the Python and AI communities to improve local AI development. Priorities include more native packages, broader kernel coverage, simpler installation, better performance, and clearer support across Windows x64 and Windows on Arm.

You can follow Windows Arm64 Python package compatibility on PyEnv-WoA-State, which links to the upstream issues and pull requests where progress is happening. Use the WoA Libs tracker for the latest status and to find ways to contribute.

Get started #

These updates give you more ways to build and experiment with local AI on Windows using the open-source models and frameworks you already know. The new Windows ML capabilities are experimental, so review the supported scenarios and known limitations before relying on them in production.

From running GGUF models with llama.cpp in Windows ML to tuning models with PyTorch and Triton, and more work happening across the open-source ecosystem, there are more ways than ever to run open-source models, build local agentic workloads, and develop and optimize your own models on Windows. Bring your own model and workload, try it on your target Windows devices, including new RTX Spark Windows PCs, and tell us what you’d like to see next.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @microsoft 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-development-on-wi…] indexed:0 read:10min 2026-10-07 · —