The News, Explained

Microsoft announced on October 7, 2026 that GitHub Copilot will choose between an on-device model and cloud models according to the coding task. The company says this orchestration will start arriving by the end of October. Its first named target is Surface Laptop Ultra, a Windows PC with NVIDIA RTX Spark. Microsoft does not present the feature as a simultaneous release for every Windows PC or current Copilot user. Source

Running an AI model on the PC is called local inference. The device’s own memory and processors perform the calculation instead of sending the input to a remote service for that work. Microsoft says Surface Laptop Ultra has up to 128 GB of unified memory and up to one petaflop of AI compute. Its first local coding model is MAI Code 1.1 Flash, a mixture-of-experts system with 137 billion total parameters and 6.8 billion active for a task. Source

The device build uses quantization and speculative decoding. Quantization stores model values with fewer bits to reduce storage and memory demand. Speculative decoding prepares likely next-token candidates to speed up a response. Microsoft lists a 53 GB package for the quantized model and 75.5 GB of peak memory use at a 256,000-token context. Those figures come from the company’s October 5 test with a particular ARM64 and CUDA setup and a synthetic coding workload, so results can vary by device, configuration and project. Source

There are two routes to the feature across Copilot CLI, the Copilot app and VS Code. Copilot can automatically choose local or cloud inference after assessing the task. Developers can also explicitly select MAI Code 1.1 Flash through the Windows ML provider, or choose a model exposed by an OpenAI-compatible local endpoint. The announcement does not specify the task-by-task routing rules or plan-level eligibility. Source

The sandbox introduced alongside local inference addresses tool permissions rather than model location. Shell commands launched by an agent normally inherit the account’s access. Microsoft Execution Containers can apply policies for files, networks, credentials, system capabilities and execution paths. The implementation uses ProcessContainer-based isolation on Windows, Seatbelt on macOS and bubblewrap on Linux. Source

With sandboxing enabled, shell commands and, by default, local Model Context Protocol servers and language servers run inside the restricted process boundary. An MCP server connects an AI agent to external tools or data. Copilot’s built-in file tools are checked by policies inside Copilot, but they do not run as operating-system-isolated child processes. Remote MCP servers also sit outside the local process boundary and receive connection-policy checks. Users therefore need to review both the allowed file and network scope and where each tool runs. The CLI exposes these controls through /sandbox; the Copilot app offers a project-level Sandbox new sessions setting. Source

OYOPICK’s Take

The verified change treats model selection and tool permissions as separate layers. A local model changes where computation happens, while a sandbox reduces what an agent-launched process can reach. Developers still need to evaluate access to project files, credentials and external connections alongside latency and network dependence.

OYOPICK sees a path toward more personal and controllable AI coding environments in this combination. That prospect depends on clear indicators for automatic routing, local execution and cloud handoffs, plus an understandable boundary around built-in tools and remote MCP servers. Supported hardware, plan availability, real power use and production performance remain questions for the rollout. This is a conditional outlook based on the published architecture.