cd /news/ai-infrastructure/lenovo-just-shoved-a-120b-parameter-… · home topics ai-infrastructure article
[ARTICLE · art-122758] src=promptcube3.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Lenovo just shoved a 120B parameter model into a laptop

Lenovo has introduced the Yoga Pro 9n laptop, which can run a 120-billion parameter AI model locally with a 1-million token context window, powered by NVIDIA's RTX Spark chip featuring a 20-core Grace CPU and 6,144 CUDA cores sharing 128GB of unified memory. The company also unveiled the ThinkCentre X Ultra (ThinkCentre X Tiny), a 1.6-liter desktop with AMD Ryzen AI Max+ PRO 495 and 128GB unified memory, which can be clustered in groups of four to form a mini-supercomputer. These releases mark a significant step in making Windows a viable platform for local AI agents, challenging Apple's dominance in the local AI workstation space.

read3 min views1 publishedSep 7, 2026
Lenovo just shoved a 120B parameter model into a laptop
Image: Promptcube3 (auto-discovered)

We're talking about a 1.65kg laptop that manages to pack 128GB of unified memory. According to NVIDIA's specs for the RTX Spark chip, this setup can actually handle a 120-billion parameter model with a 1-million token context window locally. For those who remember the struggle of fitting a decent model into 16GB of VRAM, this is a massive jump.

The death of the memory bottleneck #

The real magic here isn't just "more RAM," but the unified memory architecture. In a standard Windows rig, the CPU and GPU are constantly fighting over who gets to move data across a slow bus. The RTX Spark changes that by letting the 20-core NVIDIA Grace CPU and the 6,144 CUDA cores of the Blackwell RTX GPU share the same physical memory pool.

Since the model weights and input data only need to be stored once, you aren't wasting space. This is the same play Apple has been making with their M-series, but now it's hitting the CUDA ecosystem. Lenovo is claiming they can keep an 80W TDP under control in a 16.7mm chassis using their X Power cooling, which is impressive—or optimistic—depending on how loud the fans actually get.

Windows vs. Mac in the local AI war #

Apple has been dominating the "local AI workstation" vibe with the M5 Max and Ultra, pushing unified memory up to 512GB in the Mac Studio. They've even demonstrated 4-machine clusters via Thunderbolt 5 to run the Kimi K2.6 model. While the MacBook Pro (M5 Max) with 128GB RAM can hit around 79 token/s on a 4-bit quantized 120B model, the long context still eats memory for breakfast.

The entry of the Yoga Pro 9n means Windows is finally fighting back in the Agent host space. Until now, if you wanted a stable environment for multi-agent workflows and long contexts, you either suffered through Windows driver hell or migrated to Linux. By bringing the RTX Spark and FP4 precision to a consumer laptop, Lenovo is making Windows a viable home for local AI Agents.

More than just laptops #

If you don't want a laptop, Lenovo also released the ThinkCentre X Ultra (ThinkCentre X Tiny). It's a 1.6-liter box featuring the AMD Ryzen AI Max+ PRO 495 series and 128GB of unified memory. The wild part is the clustering; you can link four of these together to create a mini-supercomputer on your desk. A few other oddities from their lineup:

  • Project Aeroblade: Uses Frore Systems AirJet solid-state cooling to kill the fan entirely in a chassis under 10mm.
  • Project Swan: A rollable screen that expands from 14 to 17 inches.
  • Yoga Tab Plus Gen 2: A 4K 144Hz tablet with an 8192-level pressure pen.

It feels like we're finally moving past the "AI is just a chatbot in a browser" phase and into the "my hardware can actually process a million tokens without calling a server" phase. Whether the price tag makes it accessible to anyone other than enterprise execs remains to be seen.

Next Huawei's latest hardware dump proves they're obsessed with →

a practical ChatGPT prompt guide, with plenty of directly applicable cases.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @lenovo 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/lenovo-just-shoved-a…] indexed:0 read:3min 2026-09-07 ·