GitHub Copilot Local Sandboxing Reaches General Availability as the CLI Adds Ollama Model Discovery and Microsoft Details a 53GB Local Coder
GitHub made Copilot local sandboxing generally available, added local Ollama model discovery to its CLI, and Microsoft detailed a 53GB on-device MAI Code 1.1 Flash.
Overview
GitHub said on October 7 that local sandboxing for Copilot is now generally available, according to a GitHub Changelog post. The same day, GitHub added a way to pick local Ollama models inside the Copilot CLI, per a separate changelog entry, and Microsoft described an on-device coding model that Copilot is meant to route work to, in a post on the Microsoft Command Line blog.
What We Know
Sandboxing reaches general availability. The feature is generally available in the Copilot CLI, the Copilot app, and VS Code sessions that use Agent Host, GitHub says. It was introduced in the Copilot app in public preview on September 23, as previously reported. GitHub says the sandbox is powered by Microsoft eXecution Container (MXC), which translates one sandbox policy into native operating-system controls on Windows, macOS, and Linux.
According to the changelog, developers and organizations can limit the files and directories that agent-run commands can read or modify, and control access to the internet, local networks, Git credentials, and GitHub CLI credentials. The policies can also cover local MCP and language servers where supported, and enterprise-managed settings can require sandboxing and enforce policies that developers cannot weaken. GitHub adds that sandbox policies apply to tool execution regardless of which model Copilot uses, and that local sandboxing is included with GitHub Copilot at no additional cost.
The Microsoft post fills in the mechanics. It says MXC is an open-source library from the Windows team, and that Copilot uses the BaseContainer tier of the ProcessContainer backend on Windows, Seatbelt on macOS, and bubblewrap on Linux. Microsoft writes that shell commands and, by default, local MCP servers and language servers run inside the process boundary, while built-in file tools run inside Copilot itself, where requests are checked against the effective policy but are not OS-enforced child-process isolation. Remote MCP servers sit outside the local process sandbox, the post says.
Local models in the CLI. Starting in CLI version 1.0.94-0, the /model command can discover supported models from a running local Ollama instance, according to GitHub. Discovery does not add models automatically: a user chooses a discovered model, reviews its provider and endpoint, and then confirms either “Add and use for this session” or “Add without switching.” Ollama and the model must already be installed, since the flow does not install a runtime or download models, and models must support tool calling and streaming.
GitHub also cautions that choosing a local model does not turn on offline mode or disable GitHub telemetry. Offline mode remains an explicit choice through COPILOT_OFFLINE=true, and the changelog notes that a remote provider can still receive prompts and code context over the network, even in offline mode.
An on-device model. The Microsoft post describes a local version of MAI Code 1.1 Flash, which it calls a coding-optimized mixture-of-experts model with 137 billion total and 6.8 billion active parameters. After quantization and speculative decoding, the on-device version comes in at 53GB, which Microsoft says is an 80% reduction in size from the Bfloat16 cloud variant. Microsoft reports the following benchmark results, tested October 5, 2026:
- SWE-Bench Verified: 72.6% for MAI Code 1.1 Flash, 70.80% for the quantized on-device version.
- Terminal-Bench 2.1: 62.9% for MAI Code 1.1 Flash, 66.29% for the quantized on-device version.
On a Surface Laptop Ultra, which Microsoft says is built around NVIDIA RTX Spark with up to 128 GB of unified memory, the post reports peak memory use of 75.5GB at 256k context, and prompt-processing throughput of 923.5 and 769.8 tokens per second at 64k and 128k context.
Microsoft says developers will have two options across the Copilot CLI, the Copilot app, and VS Code: let Copilot’s Auto orchestration choose between local and cloud inference, or select a local model directly, either MAI Code 1.1 Flash through the Windows ML provider or an OpenAI-compatible local endpoint. The post says the Auto behavior is coming by the end of the month. GitHub’s changelog likewise says it is announcing intelligent routing with local models and that availability will follow.
What We Don’t Know
- Neither GitHub nor Microsoft has published an exact date for the Auto routing between local and cloud models; the Microsoft post says only that it is coming by the end of the month.
- The benchmark figures are Microsoft’s own, from a single test run on one date, and the post does not describe independent replication. The quantized model scores lower than the cloud variant on SWE-Bench Verified and higher on Terminal-Bench 2.1.
- The sources do not say whether the on-device model will run on hardware other than the NVIDIA RTX Spark PCs the Microsoft post targets.
- GitHub has not said how many developers have enabled the preview sandbox since September, or how often agents hit its limits.