llama.cpp Llama.cpp, the open-source project that runs AI models locally on consumer hardware, has launched a new installer and command-line tool at llama.app, enabling users to run frontier AI models such as Alibaba's Qwen 3.6, Google's Gemma 4 and Gemma 3, and OpenAI's GPT-OSS entirely on their own machines with no API keys or telemetry. The tool pairs with the local coding agent Pi via the pi-llama plugin, automatically discovering local models and keeping files and requests on-device, and is optimized for any hardware from laptops to clusters. AI that lives on your computer. Open-source, private & always local. Run frontier AI entirely on your machine. No API keys, no telemetry, no limits. Own your models and conversation data. curl -LsSf https://llama.app/install.sh | sh Pair it with a local coding agent. Run llama serve , install the pi-llama plugin and launch Pi https://github.com/earendil-works/pi . It will automatically discover your local model. No config, no API keys. Files stay on your machine, requests never leave it. 1. Serve a model llama serve 2. Install the pi-llama plugin pi install git:github.com/huggingface/pi-llama 3. Run Pi, everything is set pi Optimized for any hardware. From your laptop to a cluster, llama.cpp runs on whatever you have. Same binary, same models, same hand-tuned kernels for every GPU and CPU. Run your first model Qwen 3.6 Alibaba's next-gen natively multimodal reasoning models. Dense and MoE variants that rival models many times their size on coding and vision tasks. Gemma 4 Google's most capable open models, built from Gemini 3 technology. Supports multimodal reasoning, agentic workflows, and 140+ languages. GPT-OSS OpenAI's first open-weight models since GPT-2. Built for reasoning, agentic tasks, and developer use with function calling and tool use capabilities. Gemma 3 Google's multimodal models built from Gemini technology. Supports 140+ languages, vision, and text tasks with up to 128K context for edge to cloud deployment.