Atomic Chat: Run Open-Weight LLMs Locally on Windows, macOS & Linux Atomic Chat, a new application, enables users to run open-weight LLMs from Hugging Face locally on Windows, macOS, and Linux, and exposes them through an OpenAI-compatible API for use with coding agents, CLIs, and IDE plugins. It supports models like Llama, Gemma, Qwen, Mistral, and Phi, and offers features such as speculative decoding, Flash Attention, and TurboQuant. The tool can also connect to cloud providers like OpenAI, Anthropic, Mistral, and Groq when needed. Want to run an AI model locally, but still use it with the tools you already have? Atomic Chat makes that possible. It lets you run open-weight LLMs from Hugging Face on your own computer, then exposes them through an OpenAI-compatible API so coding agents, CLIs, IDE plugins and other apps can use your local models too. You can run models such as Llama, Gemma, Qwen, Mistral and Phi, use Atomic Chat as a regular AI chat app, or connect it to tools such as OpenCode, Goose and Kilo Code. Your local conversations and API keys can stay on your machine, while cloud providers such as OpenAI, Anthropic, Mistral and Groq are available when you need them. Under the hood, it also supports multiple inference engines and performance features such as speculative decoding, Flash Attention and TurboQuant on supported models and hardware. So you're getting more than a local chatbot. Atomic Chat can act as the local AI layer behind the rest of your setup.