Bringing local models and sandboxed tools to Windows and GitHub Copilot GitHub Copilot will automatically coordinate local and cloud inference by the end of the month, running on-device models on NVIDIA RTX Spark Windows PCs such as Surface Laptop Ultra, Microsoft announced. Microsoft AI built a local version of MAI Code 1.1 Flash, a coding-optimized mixture-of-experts model with 137 billion total and 6.8 billion active parameters, using quantization and speculative decoding to cut memory footprint and latency. Windows also developed Microsoft Execution Containers (MXC) to sandbox interactive and non-interactive agentic coding sessions, with the Surface Laptop Ultra offering up to 128 GB of unified memory and up to 1 petaflop of AI compute. When leveraging agents, developers need both choice and control. They need technologies that offer clear boundaries for their agents and make it easy to choose the model with the right speed, performance, and cost profile for each task. That’s why GitHub offers frontier models from major model providers, as well as options like Project HydraFusion https://github.blog/ai-and-ml/github-copilot/project-hydrafusion-frontier-quality-via-multi-model-orchestration/?utm source=blog-cta-project-hydrafusion-article&utm medium=blog&utm campaign=oct-7-event-oct-2026 , an orchestrator choosing one or multiple models for each task while balancing performance, cost, and latency. It’s also why Windows has developed Microsoft Execution Containers MXC https://aka.ms/WindowsDeveloperMXC to help secure interactive and non-interactive agentic coding sessions. Coming by the end of the month, GitHub Copilot will determine when a task is best handled by on-device intelligence and when it should leverage cloud-scale models. Rather than forcing developers to manage infrastructure decisions themselves, GitHub Copilot automatically coordinates local and cloud inference behind the scenes. For NVIDIA RTX Spark Windows PCs like Surface Laptop Ultra, that means we are enabling local coding in GitHub Copilot with powerful local inference models and hardware capable of delivering a great experience at the edge. The result is poised to be the next step in the HydraFusion vision: intelligent orchestration that spans not just multiple models, but multiple compute environments including the edge. GitHub Copilot can run commands in these environments with controlled access to files, networks, system capabilities, and credentials. Developers can automate with confidence and security in mind. Why memory matters for a local coding agent Local inference starts with a memory budget. Surface Laptop Ultra is built around NVIDIA RTX Spark, with up to 128 GB of unified memory and up to 1 petaflop of AI compute. With a discrete GPU, dedicated video memory is an important constraint: moving model data between system memory and the GPU can add overhead. Unified memory gives the CPU and GPU access to a shared physical pool. That makes more capacity available to the workload, but it doesn’t make all of it available to model weights. The operating system, your applications, and the inference runtime need memory, too. So does the key-value cache, which stores attention state for tokens the model has already processed. As an agent reads files and receives tool results, its context can grow, increasing memory use and the work needed to process the next request. Keeping a model loaded between requests can avoid repeated loading work. However, it doesn’t guarantee constant response time; context length, memory pressure, and the rest of the workload still matter. That’s why the useful question isn’t just whether a model fits, but how it behaves over a complete coding task. Introducing MAI Code 1.1 Flash for local coding To bring this local-development experience to life, Microsoft AI developed a local version of MAI Code 1.1 Flash, a coding-optimized mixture-of-experts model with 137 billion total and 6.8 billion active parameters. The on-device work applies quantization and speculative decoding to reduce the model footprint and improve end-to-end responsiveness while preserving the task completion and tool-use quality that matter in an agent loop. Quantization reduces the precision used to represent model weights and activations, lowering memory requirements. Because code doesn’t degrade gracefully, the release evaluation must measure coding-task success as well as footprint: a single incorrect token can produce a syntax error, wrong identifier, malformed tool call, or broken diff. Speculative decoding trades additional working memory for higher decode throughput and lower end-to-end latency. A drafter proposes candidate token blocks and the target model verifies them. For a coding agent, the important tradeoff is whether the smaller model can still complete the same tasks. A smaller footprint is useful only if changes in code quality and tool use are understood. With our first shipping version of MAI Code 1.1 Flash on Surface Laptop Ultra we achieve the following performance at different context lengths, with peak memory usage of 75.5GB at 256k context. At 64k and 128k context, prompt-processing throughput reaches 923.5 and 769.8 tokens per second, respectively. The quantized version of MAI Code 1.1 Flash we use on device retains capability impressively compared to the Bfloat16 cloud variant, coming in at 53GB, an 80% reduction in size. | Benchmark | Dataset size | MAI Code 1.1 Flash | GPT OSS 120B | MAI Code 1.1 Flash Quantized on Device | |---|---|---|---|---| | SWE-Bench Verified | 500 | 72.6% | 32.0% | 70.80% | | Terminal-Bench 2.1 | 89 | 62.9% | 23.6% | 66.29% |