LLMario Manny Singh released LLMario v0.2.0, a free Apache-2.0 open-source desktop app that runs open models such as Qwen, Gemma, Llama and gpt-oss locally on macOS 13+ (Apple Silicon and Intel) and Windows 10/11 x64 in preview. LLMario detects the user's chip and memory, estimates model memory requirements and refuses to load a model that will not fit, then drives MLX on Apple Silicon or llama.cpp elsewhere, with no accounts, cloud or telemetry. The Windows installer is not yet code-signed, so Windows may display a "Windows protected your PC" warning. New Now on Windows preview → https://github.com/mannysinghx/LLMario/releases/tag/v0.2.0 Run open AI models privately on your Mac or PC LLMario is a free, open-source app for chatting with open models like Qwen, Gemma, Llama and gpt-oss on your own computer. It picks the right engine for your hardware, checks the model fits in memory before loading it, and never sends your messages anywhere. Free · Apache-2.0 · macOS 13+ Apple Silicon and Intel · Windows 10/11 x64 preview · command line and local API included · all downloads and checksums https://github.com/mannysinghx/LLMario/releases/latest How it works LLMario doesn't reinvent AI engines. It runs proven open-source engines for you, correctly configured for your machine. You ask Chat in the LLMario window, use the llmario command, or connect your own apps through the built-in local API. LLMario plans It detects your chip and memory, estimates what the model needs, and refuses with a clear reason if it won't fit, instead of freezing your computer. The right engine runs it MLX Apple's engine, usually fastest on Apple Silicon or llama.cpp runs everywhere, including Windows . LLMario starts, watches and stops it for you. Answers stay on your computer Replies stream back as the model writes them. No accounts, no cloud, no telemetry. Your messages are never logged. Get started in 3 steps 1. 1 Install an engineEngines do the actual AI work. Install one or both : Recommended: install both. MLX is usually fastest; llama.cpp runs the widest range of models. brew install llama.cpp pip3 install mlx-lm No Homebrew? Get it at brew.sh https://brew.sh .Windows uses llama.cpp. Run this in PowerShell or Command Prompt, then restart LLMario so it finds the engine: winget install ggml.llamacpp No winget? Download a Windows zip from llama.cpp releases https://github.com/ggml-org/llama.cpp/releases and add its folder to your PATH. 2. 2 Install LLMarioDownload LLMario for macOS .dmg, Apple Silicon and Intel , open it, and drag LLMario into Applications. It is signed and notarized by Apple, so it opens without warnings. Download for macOS https://github.com/mannysinghx/LLMario/releases/download/v0.2.0/LLMario-0.2.0-macos-universal.dmg Download LLMario for Windows 10 and 11, 64-bit and run the installer. It installs for your account only, with no admin rights needed. Windows 10 may also install Microsoft's WebView2 component; Windows 11 already has it. Download for Windows https://github.com/mannysinghx/LLMario/releases/download/v0.2.0/LLMario-0.2.0-windows-x64-setup.exe Preview: the installer isn't code-signed yet, so Windows may show "Windows protected your PC". Click More info → Run anyway .Just the command-line tool? Download llmario for Windows .zip https://github.com/mannysinghx/LLMario/releases/download/v0.2.0/llmario-0.2.0-windows-x64.zip , unzip it anywhere and run llmario.exe .Builds and installs LLMario.app and the llmario command. Needs Rust https://rustup.rs and Xcode Command Line Tools xcode-select --install . The script checks both first. git clone https://github.com/mannysinghx/LLMario.git && cd LLMario && ./scripts/install-macos.sh Just the command-line tool: add --cli-only . Check prerequisites without installing: --check .On Windows, build the command-line tool with Rust and the Visual Studio Build Tools: cargo install --path crates/cli --locked in the cloned folder. 3. 3 Pick a model and chat 1. Open LLMario : from Launchpad or Spotlight on a Mac, or the Start menu on Windows. 2. Click Models → Library . Every model shows what it's good at, its size, and whether it fits your computer. 3. Click Download on one marked recommended. A small model like Qwen3.5 4B or Gemma 4 E4B is a good first choice. 4. Select it in the top bar and start typing. The first message loads the model, which takes a few seconds. Already have a model file? Drag a .gguf file or, on a Mac, an MLX model folder onto the window. It's used where it is, never copied. 4. Open What you get 🧭 Model library 37 current open-model families with exact Hugging Face names. Search by task chat, reasoning, coding, agents, multilingual, long context, small and see what fits your computer before downloading. 🧠 Memory check first LLMario estimates weights plus conversation memory before loading and explains the limit, with a smaller setting that would fit. ⚙️ Right engine, automatically MLX on Apple Silicon, llama.cpp on Intel Macs and Windows. LLMario also knows which model types your installed engine version can load. 🔒 Private by default Runs offline once a model is downloaded. Local-only network access, no accounts, no telemetry, and messages are never written to logs. ✅ Verified downloads Every file is checked against Hugging Face checksums and pinned to an exact version. Models already on your computer are reused. 🔌 Works with your apps An OpenAI-compatible API on your own computer, so tools and scripts that speak OpenAI can use local models instead. Models you can run A curated library of current open models, verified against Hugging Face. The app shows the full list with sizes and licenses. Full list with exact file names: docs/MODELS.md https://github.com/mannysinghx/LLMario/blob/main/docs/MODELS.md . Check each model's license before commercial use. Configure In the app: Settings | Setting | What it does | |---|---| | Performance profile | Latency : one chat, fastest replies default . Balanced : up to 4 requests at once, e.g. your chat plus other apps. Throughput : many parallel requests. | | Context length | How much conversation the model remembers. Longer uses more memory. | | Max reply length | Upper limit on tokens per answer. Raise it for "thinking" models. | | Temperature | 0 = predictable, higher = more creative. | | System prompt | Instructions applied to every chat. | | Chat history | Stored only on this computer. Turn it off or clear it anytime. | Advanced: config file Optional. Lives at ~/.llmario/config.toml Windows: %USERPROFILE%\.llmario\config.toml and applies to the app, the command line and the API. server port = 11500 local API port runtime profile = "latency" latency | balanced | throughput context = 16384 tokens per conversation memory limit gb = 24 cap LLMario's memory use backends prefer = "mlx" mlx | llamacpp See the effective settings with llmario config . Models are stored in ~/.llmario/models . For developers Command line llmario doctor your hardware, engines, what fits llmario model catalog coding browse the library llmario model pull qwen3.5-9b best version for your machine llmario run qwen3.5-9b "Hello " chat in the terminal llmario serve start the local API llmario bench -m qwen3.5-9b measure speed and memory OpenAI-compatible API llmario serve listens on http://127.0.0.1:11500/v1 , on your computer only unless you enable remote access with an API key. python from openai import OpenAI client = OpenAI base url="http://127.0.0.1:11500/v1", api key="local" r = client.chat.completions.create model="qwen3.5-9b", messages= {"role": "user", "content": "Hello "} print r.choices 0 .message.content Questions Is it really private? Yes. Models run on your own computer. After a model is downloaded, LLMario works offline. The API listens only on your computer by default, messages are never logged, and there is no telemetry. Which computers work? Mac: macOS 13 or later. Apple Silicon M1 and newer is fastest and can use both engines; Intel Macs use llama.cpp. Windows preview : Windows 10 or 11, 64-bit, with llama.cpp. LLMario uses an NVIDIA graphics card when it finds one; otherwise models run on the processor, which is slower. Memory decides model size: 8 GB suits small models 1–4B , 16 GB mid-size up to ~9B , and 32 GB+ larger ones. The app shows exactly what fits. A model says "too large" or "needs a newer engine" Too large : pick a smaller model, or lower the context length in Settings. Needs a newer engine : the model uses a new architecture your installed engine doesn't support yet. Update it with brew upgrade llama.cpp or pip3 install -U mlx-lm on a Mac, or winget upgrade ggml.llamacpp on Windows, or choose the model's other engine version. "Engine not installed" / nothing happens when I chat Install an engine step 1 , then reopen LLMario. Models → Engines shows what LLMario found. llmario doctor gives full details. Windows says "Windows protected your PC" The Windows preview isn't code-signed yet, so Microsoft SmartScreen shows this for new apps. Click More info → Run anyway . Only do this for the installer downloaded from this site or the GitHub releases page https://github.com/mannysinghx/LLMario/releases . You can check the download against its .sha256 file with Get-FileHash in PowerShell. macOS says the app can't be opened Official releases are notarized by Apple and open normally. If you received a build from someone else, prefer the GitHub release or install from source. Does it use Ollama or send data to a server? No. LLMario runs the open-source engines llama.cpp and MLX directly. It only goes online to download a model you choose from Hugging Face. How do I remove it? Mac: delete LLMario from Applications. Windows: Settings → Apps → LLMario → Uninstall. Models and settings are in ~/.llmario Windows: %USERPROFILE%\.llmario ; delete that folder to free the space. Remove a source-built command-line tool with cargo uninstall llmario .