# GitHub Copilot Goes Local: MAI Code Flash, $2,599 Gate

> Source: <https://byteiota.com/github-copilot-goes-local-mai-code-flash-2599-gate/>
> Published: 2026-10-09 10:07:25+00:00

GitHub Copilot is about to route your coding tasks to a model sitting on your own machine. No cloud round-trip, inference stays on device — that’s the pitch. The catch: you need hardware that starts at $2,599 and ships October 16. And Microsoft still hasn’t answered the one question that makes the whole privacy angle worth anything: how much of your code does Auto mode quietly forward to the cloud?

## What’s Actually Shipping

The feature is called HydraFusion, and it’s been running in Copilot as a cloud-only routing system since May 2026. It evaluates each coding task across four dimensions — reasoning, code generation, debugging, and tool use — then picks a single model, a cascade, or a draft-critique-revise workflow depending on complexity. That 86-millisecond routing decision happens before any model call.

What’s new is the local leg. Starting end of October, HydraFusion can route tasks to MAI Code 1.1 Flash running on your machine instead of always calling out to the cloud. The model is a mixture-of-experts architecture: 137 billion total parameters, 6.8 billion active per forward pass, 256K context window. Microsoft quantized the on-device build down to roughly 53GB — about 80% smaller than the full model — while keeping its 72.6% SWE-bench Verified score intact.

Two modes ship with the rollout. Auto lets HydraFusion decide where each request goes, balancing task complexity, local hardware, and cost. Explicit forces everything through your local model. The rollout covers Copilot CLI, the Copilot app, and VS Code by end of October. [Microsoft’s full announcement](https://commandline.microsoft.com/local-models-sandboxed-tools-github-windows/) has technical details on sandboxing and the MXC execution container.

## The Hardware Wall

Local inference requires an NVIDIA RTX Spark GPU — Blackwell architecture, currently only available in the [Surface Laptop Ultra starting at $2,599.99](https://the-gadgeteer.com/2026/10/07/surface-laptop-ultra-starts-at-2599-as-microsoft-puts-more-ai-work-on-the-pc/). The 24GB entry config ships October 16; the 128GB version you’d need to run MAI Code 1.1 Flash at full context without hitting memory pressure costs $5,899.99.

On that hardware, performance is genuinely good: 923 tokens per second at 64K context, 770 tokens per second at 128K. Peak memory use hits about 75.5GB at maximum context length — meaning most of the entry model’s memory budget is already spoken for.

If you’re on a MacBook, an AMD machine, or any laptop without RTX Spark — which is virtually every developer machine currently in service — local inference doesn’t apply to you. HydraFusion still works on any hardware; it just routes everything to cloud.

## What Microsoft Won’t Say

Here’s where the privacy pitch gets complicated. Microsoft has confirmed the rollout timeline and published performance numbers. What it hasn’t disclosed: how much repository context Auto mode sends to cloud models when it routes a task there, whether developers can inspect routing decisions, and whether Auto can be restricted to local-only. Subscription tiers and pricing for local inference are also unannounced.

This matters because “going local” as a privacy story only holds up if you know when you’re actually local. In Auto mode, you don’t. [The New Stack flagged the same gap](https://thenewstack.io/https-thenewstack-io-copilot-local-inference-routing/): the routing decision is opaque, and you could be sending code to Azure on half your completions without any indication.

One important caveat: selecting a local model keeps inference on device, but does not stop the agent from making network requests through its tools. Full air-gap requires BYOK mode — set `COPILOT_OFFLINE=true` in Copilot CLI — which has been available since April 2026.

## What Developers Can Do Now

If you’re not buying a $2,600 laptop to access local inference, you already have options. The Ollama integration GitHub Copilot added in March 2026 routes inference to locally running models — any hardware with sufficient VRAM, zero cloud, no telemetry when paired with `COPILOT_OFFLINE=true`. It works today, on existing machines, for free.

Microsoft’s implementation adds polished Auto routing and better hardware optimization, but the underlying capability isn’t new. [A detailed HydraFusion technical breakdown](https://nerdleveltech.com/github-hydrafusion-multi-model-copilot-routing) is worth reading if you want to understand the routing mechanics before the rollout hits your editor.

Enterprises in regulated industries — fintech, healthcare, defense — are the clearest early adopters for the premium hardware path. For everyone else, the privacy question remains open until Microsoft publishes what Auto mode actually does with your code.
