Local AI is moving from "toy" status to actually handling heavy AMD's new Ryzen AI Max 400 (Gorgon Halo) chips with up to 192GB of unified memory can now run 320B-parameter models like GLM-5.3-Flash locally, while the Threadripper Halo Station with 2TB system memory and 576GB HBM3E supports models exceeding 1 trillion parameters. Microsoft's Project Zenith for developer PCs with at least 64GB unified memory includes WSL, VS Code, and GitHub Copilot CLI, and Microsoft Execution Containers sandbox AI agents. In AMD demos, the local Laguna S 2.1 model on Strix Halo scored 70.3 on a software engineering benchmark, beating cloud-based Claude Sonnet 5, which costs about 90 Euros for 10 million output tokens versus near-free local inference. Local AI is moving from "toy" status to actually handling heavy The hardware tiers are getting wild. We've gone from the Ryzen AI 400 Gorgon Point which handles 24B models, up to the Strix Halo with 128GB of unified memory capable of running 200B parameter models. But the real beast is the Gorgon Halo Ryzen AI Max 400 , pushing unified memory up to 192GB. This allows local execution of models like GLM-5.3-Flash with 320B parameters. I find it interesting that AMD is using these open-weight models as the new benchmark for hardware success—it's no longer about gaming FPS, but about how many billions of parameters you can fit in RAM. For those who need even more, the Threadripper Halo Station is basically a liquid-cooled supercomputer for your desk. It packs a 96-core Threadripper PRO and up to four AMD Instinct MI350P cards. With system memory hitting 2TB and HBM3E capacity up to 576GB, it can supposedly run models exceeding 1 trillion parameters. This is a massive leap for a local AI workflow, essentially creating a private shared node for small dev teams so they don't have to rely on metered cloud APIs. On the software side, the "unmetered intelligence" vision depends on how Windows handles this hardware. Microsoft is rolling out Project Zenith for developer-grade gear requiring at least 64GB unified memory , which comes pre-loaded with WSL, VS Code, and GitHub Copilot /en/tags/github%20copilot/ CLI. They're also testing Microsoft Execution Containers to give agents a sandbox so they don't accidentally delete your home directory while trying to "organize your files." The performance gap is closing fast too. We're seeing models like Qwen3.5-9B outperforming much larger predecessors on GPQA benchmarks. In the AMD demo, the local Laguna S 2.1 model on Strix Halo actually beat the cloud-based Claude /en/tags/claude/ Sonnet 5 on a software engineering benchmark with a score of 70.3. When you realize 10 million output tokens on Sonnet 5 costs about 90 Euros while local is essentially free minus electricity , the math for local deployment becomes a no-brainer. The roadmap is clear: move the "personal context"—your calendars, private files, and specific habits—away from the cloud and onto silicon you actually own. If we can run 300B+ models on a laptop like the new HP ZBook codename Sundance , the cloud becomes a choice for scale, not a necessity for intelligence. Next Building a product with AI is a trip—you can go from zero to a → /en/threads/8948/