{"slug": "llmario", "title": "LLMario", "summary": "Manny Singh released LLMario v0.2.0, a free Apache-2.0 open-source desktop app that runs open models such as Qwen, Gemma, Llama and gpt-oss locally on macOS 13+ (Apple Silicon and Intel) and Windows 10/11 x64 in preview. LLMario detects the user's chip and memory, estimates model memory requirements and refuses to load a model that will not fit, then drives MLX on Apple Silicon or llama.cpp elsewhere, with no accounts, cloud or telemetry. The Windows installer is not yet code-signed, so Windows may display a \"Windows protected your PC\" warning.", "body_md": "[New Now on Windows (preview) →](https://github.com/mannysinghx/LLMario/releases/tag/v0.2.0)\n\n# Run open AI models privately on your Mac or PC\n\nLLMario is a free, open-source app for chatting with open models like Qwen, Gemma, Llama and gpt-oss on your own computer. It picks the right engine for your hardware, checks the model fits in memory before loading it, and never sends your messages anywhere.\n\nFree · Apache-2.0 · macOS 13+ (Apple Silicon and Intel) · Windows 10/11 x64 (preview) · command line and local API included · [all downloads and checksums](https://github.com/mannysinghx/LLMario/releases/latest)\n\n## How it works\n\nLLMario doesn't reinvent AI engines. It runs proven open-source engines for you, correctly configured for your machine.\n\n### You ask\n\nChat in the LLMario window, use the `llmario` command, or connect your own apps through the built-in local API.\n\n### LLMario plans\n\nIt detects your chip and memory, estimates what the model needs, and refuses with a clear reason if it won't fit, instead of freezing your computer.\n\n### The right engine runs it\n\n**MLX** (Apple's engine, usually fastest on Apple Silicon) or **llama.cpp** (runs everywhere, including Windows). LLMario starts, watches and stops it for you.\n\n### Answers stay on your computer\n\nReplies stream back as the model writes them. No accounts, no cloud, no telemetry. Your messages are never logged.\n\n## Get started in 3 steps\n\n1. \n1### Install an engineEngines do the actual AI work. Install one (or both): Recommended: install both. MLX is usually fastest; llama.cpp runs the widest range of models. \n\n```\nbrew install llama.cpp\npip3 install mlx-lm\n```\n\n No Homebrew? Get it at [brew.sh](https://brew.sh) .Windows uses llama.cpp. Run this in PowerShell or Command Prompt, then restart LLMario so it finds the engine: \n\n```\nwinget install ggml.llamacpp\n```\n\n No winget? Download a Windows zip from [llama.cpp releases](https://github.com/ggml-org/llama.cpp/releases) and add its folder to your PATH.\n2. \n2### Install LLMarioDownload **LLMario for macOS** (.dmg, Apple Silicon and Intel), open it, and drag LLMario into Applications. It is signed and notarized by Apple, so it opens without warnings.[Download for macOS](https://github.com/mannysinghx/LLMario/releases/download/v0.2.0/LLMario-0.2.0-macos-universal.dmg)Download **LLMario for Windows** (10 and 11, 64-bit) and run the installer. It installs for your account only, with no admin rights needed. Windows 10 may also install Microsoft's WebView2 component; Windows 11 already has it.[Download for Windows](https://github.com/mannysinghx/LLMario/releases/download/v0.2.0/LLMario-0.2.0-windows-x64-setup.exe)**Preview:** the installer isn't code-signed yet, so Windows may show \"Windows protected your PC\". Click**More info** →**Run anyway** .Just the command-line tool? Download [llmario for Windows (.zip)](https://github.com/mannysinghx/LLMario/releases/download/v0.2.0/llmario-0.2.0-windows-x64.zip) , unzip it anywhere and run`llmario.exe` .Builds and installs `LLMario.app` and the`llmario` command. Needs[Rust](https://rustup.rs) and Xcode Command Line Tools (`xcode-select --install` ). The script checks both first.\n\n```\ngit clone https://github.com/mannysinghx/LLMario.git && cd LLMario && ./scripts/install-macos.sh\n```\n\n Just the command-line tool: add `--cli-only` . Check prerequisites without installing:`--check` .On Windows, build the command-line tool with Rust and the Visual Studio Build Tools: `cargo install --path crates/cli --locked` in the cloned folder.\n3. \n3### Pick a model and chat\n  1. Open **LLMario** : from Launchpad or Spotlight on a Mac, or the Start menu on Windows.\n  2. Click **Models** →**Library** . Every model shows what it's good at, its size, and whether it fits your computer.\n  3. Click **Download** on one marked recommended. A small model like*Qwen3.5 4B* or*Gemma 4 E4B* is a good first choice.\n  4. Select it in the top bar and start typing. The first message loads the model, which takes a few seconds.\n Already have a model file? Drag a `.gguf` file (or, on a Mac, an MLX model folder) onto the window. It's used where it is, never copied.\n4. Open \n\n## What you get\n\n### 🧭 Model library\n\n37 current open-model families with exact Hugging Face names. Search by task (chat, reasoning, coding, agents, multilingual, long context, small) and see what fits *your* computer before downloading.\n\n### 🧠 Memory check first\n\nLLMario estimates weights plus conversation memory before loading and explains the limit, with a smaller setting that would fit.\n\n### ⚙️ Right engine, automatically\n\nMLX on Apple Silicon, llama.cpp on Intel Macs and Windows. LLMario also knows which model types your installed engine version can load.\n\n### 🔒 Private by default\n\nRuns offline once a model is downloaded. Local-only network access, no accounts, no telemetry, and messages are never written to logs.\n\n### ✅ Verified downloads\n\nEvery file is checked against Hugging Face checksums and pinned to an exact version. Models already on your computer are reused.\n\n### 🔌 Works with your apps\n\nAn OpenAI-compatible API on your own computer, so tools and scripts that speak OpenAI can use local models instead.\n\n## Models you can run\n\nA curated library of current open models, verified against Hugging Face. The app shows the full list with sizes and licenses.\n\nFull list with exact file names: [docs/MODELS.md](https://github.com/mannysinghx/LLMario/blob/main/docs/MODELS.md). Check each model's license before commercial use.\n\n## Configure\n\n### In the app: Settings\n\n| Setting | What it does | \n|---|---|\n| Performance profile | **Latency** : one chat, fastest replies (default).**Balanced** : up to 4 requests at once, e.g. your chat plus other apps.**Throughput** : many parallel requests. | \n| Context length | How much conversation the model remembers. Longer uses more memory. | \n| Max reply length | Upper limit on tokens per answer. Raise it for \"thinking\" models. | \n| Temperature | 0 = predictable, higher = more creative. | \n| System prompt | Instructions applied to every chat. | \n| Chat history | Stored only on this computer. Turn it off or clear it anytime. | \n\n### Advanced: config file\n\nOptional. Lives at `~/.llmario/config.toml` (Windows: `%USERPROFILE%\\.llmario\\config.toml`) and applies to the app, the command line and the API.\n\n```\n[server]\nport = 11500              # local API port\n\n[runtime]\nprofile = \"latency\"       # latency | balanced | throughput\ncontext = 16384           # tokens per conversation\nmemory_limit_gb = 24      # cap LLMario's memory use\n\n[backends]\nprefer = \"mlx\"            # mlx | llamacpp\n```\n\nSee the effective settings with `llmario config`. Models are stored in `~/.llmario/models`.\n\n## For developers\n\n### Command line\n\n```\nllmario doctor                    # your hardware, engines, what fits\nllmario model catalog coding      # browse the library\nllmario model pull qwen3.5-9b     # best version for your machine\nllmario run qwen3.5-9b \"Hello!\"   # chat in the terminal\nllmario serve                     # start the local API\nllmario bench -m qwen3.5-9b       # measure speed and memory\n```\n\n### OpenAI-compatible API\n\n`llmario serve` listens on `http://127.0.0.1:11500/v1`, on your computer only unless you enable remote access with an API key.\n\n``` python\nfrom openai import OpenAI\nclient = OpenAI(base_url=\"http://127.0.0.1:11500/v1\", api_key=\"local\")\nr = client.chat.completions.create(\n    model=\"qwen3.5-9b\",\n    messages=[{\"role\": \"user\", \"content\": \"Hello!\"}])\nprint(r.choices[0].message.content)\n```\n\n## Questions\n\n## Is it really private?\n\nYes. Models run on your own computer. After a model is downloaded, LLMario works offline. The API listens only on your computer by default, messages are never logged, and there is no telemetry.\n\n## Which computers work?\n\n**Mac:** macOS 13 or later. Apple Silicon (M1 and newer) is fastest and can use both engines; Intel Macs use llama.cpp. **Windows (preview):** Windows 10 or 11, 64-bit, with llama.cpp. LLMario uses an NVIDIA graphics card when it finds one; otherwise models run on the processor, which is slower. Memory decides model size: 8 GB suits small models (1–4B), 16 GB mid-size (up to ~9B), and 32 GB+ larger ones. The app shows exactly what fits.\n\n## A model says \"too large\" or \"needs a newer engine\"\n\n**Too large**: pick a smaller model, or lower the context length in Settings. **Needs a newer engine**: the model uses a new architecture your installed engine doesn't support yet. Update it with `brew upgrade llama.cpp` or `pip3 install -U mlx-lm` on a Mac, or `winget upgrade ggml.llamacpp` on Windows, or choose the model's other engine version.\n\n## \"Engine not installed\" / nothing happens when I chat\n\nInstall an engine (step 1), then reopen LLMario. **Models → Engines** shows what LLMario found. `llmario doctor` gives full details.\n\n## Windows says \"Windows protected your PC\"\n\nThe Windows preview isn't code-signed yet, so Microsoft SmartScreen shows this for new apps. Click **More info** → **Run anyway**. Only do this for the installer downloaded from this site or the [GitHub releases page](https://github.com/mannysinghx/LLMario/releases). You can check the download against its `.sha256` file with `Get-FileHash` in PowerShell.\n\n## macOS says the app can't be opened\n\nOfficial releases are notarized by Apple and open normally. If you received a build from someone else, prefer the GitHub release or install from source.\n\n## Does it use Ollama or send data to a server?\n\nNo. LLMario runs the open-source engines llama.cpp and MLX directly. It only goes online to download a model you choose from Hugging Face.\n\n## How do I remove it?\n\n**Mac:** delete LLMario from Applications. **Windows:** Settings → Apps → LLMario → Uninstall. Models and settings are in `~/.llmario` (Windows: `%USERPROFILE%\\.llmario`); delete that folder to free the space. Remove a source-built command-line tool with `cargo uninstall llmario`.", "url": "https://wpnews.pro/news/llmario", "canonical_source": "https://www.llmario.com", "published_at": "2026-10-02 04:55:41+00:00", "updated_at": "2026-10-02 05:15:48.754020+00:00", "lang": "en", "topics": ["ai-products", "ai-tools", "large-language-models", "artificial-intelligence"], "entities": ["LLMario", "Manny Singh", "Qwen", "Gemma", "Llama", "gpt-oss", "MLX", "llama.cpp"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/llmario", "markdown": "https://wpnews.pro/news/llmario.md", "text": "https://wpnews.pro/news/llmario.txt", "jsonld": "https://wpnews.pro/news/llmario.jsonld"}}