Microsoft and Nvidia Bring Local AI to Windows: Surface Laptop Ultra From $2,599 and 120B Models On-Device
Microsoft put numbers on the Surface Laptop Ultra on October 7, 2026: its first laptop with Nvidia RTX Spark, from $2,599, up to 128GB of unified memory and models with more than 120 billion parameters running on the machine itself.
The story isn't just the laptop: it's where Microsoft wants AI to run.
What Was Announced on October 7, 2026
At its first big Windows event in two years, held in San Francisco with Satya Nadella, Pavan Davuluri and Jensen Huang on stage, Microsoft and Nvidia closed the launch of the Surface Laptop Ultra. The hardware had been shown months earlier; October 7 brought the prices, the shipping date and, above all, how the device fits into the new Windows strategy.
Surface Laptop Ultra: 15 Inches, RTX Spark and Up to 128GB of Unified Memory
It is a 15-inch laptop with a mini-LED display, an Arm CPU with up to 20 cores, a Blackwell GPU with up to 6,144 cores and up to 128GB of unified memory, with full CUDA support. Microsoft positions it for creators, developers and AI builders, not for someone who just browses.
Read also
The technical key is that memory. On these machines the CPU and the GPU share the same memory, so the model does not have to fit inside a separate graphics card. That is why 128GB of unified memory can load models a 24GB GPU cannot even open.
Prices and Dates: From $2,599, Shipping October 16
- Surface Laptop Ultra: from $2,599, with pre-orders already open.
- Shipping: Friday, October 16, 2026.
- Memory configurations: 32, 64 or 128GB according to the official spec sheet, and it cannot be upgraded later.
- Surface RTX Spark Dev Box: from $5,999, shipping in November.
Surface RTX Spark Dev Box: the $5,999 Mini PC Arriving in November
Alongside the laptop, Microsoft is taking orders for the Surface RTX Spark Dev Box, a desktop machine from $5,999 with 128GB of unified memory and 2TB of storage, aimed at developers. It arrives in November and, at first, is sold in the United States only.
Don't mix the two prices up: $2,599 is the laptop's entry point; $5,999 is the Dev Box's floor.
It Doesn't Come Alone: Dell, HP, Lenovo, ASUS and MSI Ship the Same Day
Nvidia is not betting on a single horse. RTX Spark laptops from Dell, HP, Lenovo, ASUS and MSI launch on the same day, October 16. Microsoft and Nvidia want the platform to belong to the whole ecosystem, not to one brand.
What RTX Spark Is and Why Everyone Talks About Memory
An SoC With CPU, Blackwell GPU and Unified Memory
RTX Spark is an all-in-one Nvidia chip in the GB10 family already seen in AI development machines: an Arm CPU, a Blackwell GPU and memory shared between both inside the same package. That design is what lets a large model fit into a desktop machine.
Unified Memory: Why the Whole Model Has to Fit
Model weights are stored at reduced precision (quantization): fewer bits per number, less memory and a controlled quality loss. The mental math is simple: more parameters and more context demand more memory. If the model does not fit, the system falls back to disk and everything crawls.
Up to 1 Petaflop and Up to 120B Parameters: What Those "Up To"s Mean
Those are configuration-ceiling figures. A petaflop is a measure of AI compute at reduced precision, not a speed promise; models with more than 120 billion parameters only fit in the high-memory configurations. On the base model, those on-stage numbers do not apply.
The Number Nobody Shared On Stage: Tokens Per Second
What users actually notice is how many tokens per second it produces and how fast it answers, and those numbers are not public yet. Any local-versus-cloud performance comparison based on guesses is worth nothing.
The Windows Changes Developers Should Care About
Truly Hybrid: Copilot Tapping Models Running on Your PC
Microsoft calls it hybrid intelligence: the same feature can route a task to the local model (private, no per-token cost, offline) or to the cloud (more capable and up to date). The system, or the developer, decides where each job runs. That path already existed: we covered it in our guide to local AI with Ollama.
Windows ML Adds llama.cpp: GGUF Models Locally
The most technical piece: Windows ML adds experimental support for llama.cpp to run models in GGUF format from Windows, through new task-specific APIs. It is experimental, not stable: good for testing, not for production.
Agents Without Setup and the OpenClaw Gateway
Windows adds a getting-started experience to spin up agents without configuring anything and a native gateway for OpenClaw, the open standard for agents. The idea is that a Windows PC can host local agents the way it hosts applications today. The world of standards connecting AI to your tools is explained in Model Context Protocol.
Copilot+: the Requirements That Leave Most PCs Out
Many of the new features depend on Copilot+, which requires a 40 TOPS NPU, 16GB of RAM and 256GB of storage, meaning compatible hardware, and a rollout that varies by market. They are not available "now" on any machine.
The Awkward Numbers
$2,599 Against the 16-Inch MacBook Pro
The entry price sits below the 16-inch MacBook Pro's $2,999, and the press in the room compared them right away. Microsoft is aiming at the same creator and developer audience as Apple.
The Memory Bill: RAM Shortages and Rising Prices
The reason repeated in every report is awkward: memory is scarce and expensive because of AI infrastructure demand, and that pushes device prices up. A $2,599 laptop is also a story about the price of RAM, the same shortage that led Micron to warn that memory will stay tight through 2028.
What Microsoft Didn't Say: Power Draw, Battery Life and Real Performance
There were no tokens per second, no battery figures and no power draw under load. Those will come with the independent reviews in the week of October 16.
What This Means for You
If You Code
A decent local model next to your editor gives you code privacy, zero per-token cost and offline work. The cloud still wins on frontier models, huge contexts and teamwork. Almost nobody will pick one: they combine.
If You Create 3D, Video or Audio
For AI pipelines (upscaling, asset generation, cleaning up takes) unified memory matters more than a big GPU: it decides which model fits and at what quality you work. The real limit is memory, not core count.
If You Buy Hardware
Learn to read the configurations: 18 or 20 CPU cores, 5,120 or 6,144 GPU cores and, above all, 32, 64 or 128GB of memory. On these machines RAM is soldered and cannot be upgraded, so the decision happens at purchase, not later.
What Comes Next
The independent reviews in the week of October 16 (tokens per second, power draw and battery), Windows ML with llama.cpp leaving the experimental stage and how the rest of the industry responds. If Apple or the manufacturers make a move on memory or price, that is where it will show.
Conclusion
The Surface Laptop Ultra is the showcase for a bigger bet: that AI runs on your machine and not only in the cloud. It is an announced bet, not a result, and we still need to see real performance and whether the price of a PC with this memory holds up. In the meantime, we keep telling you on the blog what actually changes and what is just stagecraft.


