Run frontier AI entirely on your machine. No API keys, no telemetry, no limits. Own your models and conversation data.
curl -LsSf https://llama.app/install.sh | sh
Pair it with a local coding agent. #
Run llama serve
, install the pi-llama
plugin and launch Pi. It will automatically discover your local model. No config, no API keys. Files stay on your machine, requests never leave it.
llama serve
pi install git:github.com/huggingface/pi-llama
pi
Optimized for any hardware. #
From your laptop to a cluster, llama.cpp runs on whatever you have. Same binary, same models, same hand-tuned kernels for every GPU and CPU.
Run your first model #
Qwen 3.6
Alibaba's next-gen natively multimodal reasoning models. Dense and MoE variants that rival models many times their size on coding and vision tasks.
Gemma 4
Google's most capable open models, built from Gemini 3 technology. Supports multimodal reasoning, agentic workflows, and 140+ languages.
GPT-OSS
OpenAI's first open-weight models since GPT-2. Built for reasoning, agentic tasks, and developer use with function calling and tool use capabilities.
Gemma 3
Google's multimodal models built from Gemini technology. Supports 140+ languages, vision, and text tasks with up to 128K context for edge to cloud deployment.