cd /news/ai-products/llmario · home › topics › ai-products › article
[ARTICLE · art-143686] src=llmario.com ↗ pub= topic=ai-products verified=true sentiment=↑ positive

LLMario

Manny Singh released LLMario v0.2.0, a free Apache-2.0 open-source desktop app that runs open models such as Qwen, Gemma, Llama and gpt-oss locally on macOS 13+ (Apple Silicon and Intel) and Windows 10/11 x64 in preview. LLMario detects the user's chip and memory, estimates model memory requirements and refuses to load a model that will not fit, then drives MLX on Apple Silicon or llama.cpp elsewhere, with no accounts, cloud or telemetry. The Windows installer is not yet code-signed, so Windows may display a "Windows protected your PC" warning.

read7 min views1 publishedOct 2, 2026
LLMario
Image: source

New Now on Windows (preview) →

LLMario is a free, open-source app for chatting with open models like Qwen, Gemma, Llama and gpt-oss on your own computer. It picks the right engine for your hardware, checks the model fits in memory before it, and never sends your messages anywhere.

Free · Apache-2.0 · macOS 13+ (Apple Silicon and Intel) · Windows 10/11 x64 (preview) · command line and local API included · all downloads and checksums

How it works #

LLMario doesn't reinvent AI engines. It runs proven open-source engines for you, correctly configured for your machine.

You ask

Chat in the LLMario window, use the llmario command, or connect your own apps through the built-in local API.

LLMario plans

It detects your chip and memory, estimates what the model needs, and refuses with a clear reason if it won't fit, instead of freezing your computer.

The right engine runs it

MLX (Apple's engine, usually fastest on Apple Silicon) or llama.cpp (runs everywhere, including Windows). LLMario starts, watches and stops it for you.

Answers stay on your computer

Replies stream back as the model writes them. No accounts, no cloud, no telemetry. Your messages are never logged.

Get started in 3 steps #

1### Install an engineEngines do the actual AI work. Install one (or both): Recommended: install both. MLX is usually fastest; llama.cpp runs the widest range of models.

brew install llama.cpp
pip3 install mlx-lm

No Homebrew? Get it at brew.sh .Windows uses llama.cpp. Run this in PowerShell or Command Prompt, then restart LLMario so it finds the engine:

winget install ggml.llamacpp

No winget? Download a Windows zip from llama.cpp releases and add its folder to your PATH. 2. 2### Install LLMarioDownload LLMario for macOS (.dmg, Apple Silicon and Intel), open it, and drag LLMario into Applications. It is signed and notarized by Apple, so it opens without warnings.Download for macOSDownload LLMario for Windows (10 and 11, 64-bit) and run the installer. It installs for your account only, with no admin rights needed. Windows 10 may also install Microsoft's WebView2 component; Windows 11 already has it.Download for WindowsPreview: the installer isn't code-signed yet, so Windows may show "Windows protected your PC". ClickMore info →Run anyway .Just the command-line tool? Download llmario for Windows (.zip) , unzip it anywhere and runllmario.exe .Builds and installs LLMario.app and thellmario command. NeedsRust and Xcode Command Line Tools (xcode-select --install ). The script checks both first.

git clone https://github.com/mannysinghx/LLMario.git && cd LLMario && ./scripts/install-macos.sh

Just the command-line tool: add --cli-only . Check prerequisites without installing:--check .On Windows, build the command-line tool with Rust and the Visual Studio Build Tools: cargo install --path crates/cli --locked in the cloned folder. 3. 3### Pick a model and chat

  1. Open LLMario : from Launchpad or Spotlight on a Mac, or the Start menu on Windows.
  2. Click Models →Library . Every model shows what it's good at, its size, and whether it fits your computer.
  3. Click Download on one marked recommended. A small model likeQwen3.5 4B orGemma 4 E4B is a good first choice.
  4. Select it in the top bar and start typing. The first message loads the model, which takes a few seconds. Already have a model file? Drag a .gguf file (or, on a Mac, an MLX model folder) onto the window. It's used where it is, never copied.
  5. Open

What you get #

🧭 Model library

37 current open-model families with exact Hugging Face names. Search by task (chat, reasoning, coding, agents, multilingual, long context, small) and see what fits your computer before down.

🧠 Memory check first

LLMario estimates weights plus conversation memory before and explains the limit, with a smaller setting that would fit.

⚙️ Right engine, automatically

MLX on Apple Silicon, llama.cpp on Intel Macs and Windows. LLMario also knows which model types your installed engine version can load.

🔒 Private by default

Runs offline once a model is downloaded. Local-only network access, no accounts, no telemetry, and messages are never written to logs.

✅ Verified downloads

Every file is checked against Hugging Face checksums and pinned to an exact version. Models already on your computer are reused.

🔌 Works with your apps

An OpenAI-compatible API on your own computer, so tools and scripts that speak OpenAI can use local models instead.

Models you can run #

A curated library of current open models, verified against Hugging Face. The app shows the full list with sizes and licenses.

Full list with exact file names: docs/MODELS.md. Check each model's license before commercial use.

Configure #

In the app: Settings

Setting What it does
Performance profile Latency : one chat, fastest replies (default).Balanced : up to 4 requests at once, e.g. your chat plus other apps.Throughput : many parallel requests.
Context length How much conversation the model remembers. Longer uses more memory.
Max reply length Upper limit on tokens per answer. Raise it for "thinking" models.
Temperature 0 = predictable, higher = more creative.
System prompt Instructions applied to every chat.
Chat history Stored only on this computer. Turn it off or clear it anytime.

Advanced: config file

Optional. Lives at ~/.llmario/config.toml (Windows: %USERPROFILE%\.llmario\config.toml) and applies to the app, the command line and the API.

[server]
port = 11500              # local API port

[runtime]
profile = "latency"       # latency | balanced | throughput
context = 16384           # tokens per conversation
memory_limit_gb = 24      # cap LLMario's memory use

[backends]
prefer = "mlx"            # mlx | llamacpp

See the effective settings with llmario config. Models are stored in ~/.llmario/models.

For developers #

Command line

llmario doctor                    # your hardware, engines, what fits
llmario model catalog coding      # browse the library
llmario model pull qwen3.5-9b     # best version for your machine
llmario run qwen3.5-9b "Hello!"   # chat in the terminal
llmario serve                     # start the local API
llmario bench -m qwen3.5-9b       # measure speed and memory

OpenAI-compatible API

llmario serve listens on http://127.0.0.1:11500/v1, on your computer only unless you enable remote access with an API key.

from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:11500/v1", api_key="local")
r = client.chat.completions.create(
    model="qwen3.5-9b",
    messages=[{"role": "user", "content": "Hello!"}])
print(r.choices[0].message.content)

Questions #

Is it really private? #

Yes. Models run on your own computer. After a model is downloaded, LLMario works offline. The API listens only on your computer by default, messages are never logged, and there is no telemetry.

Which computers work? #

Mac: macOS 13 or later. Apple Silicon (M1 and newer) is fastest and can use both engines; Intel Macs use llama.cpp. Windows (preview): Windows 10 or 11, 64-bit, with llama.cpp. LLMario uses an NVIDIA graphics card when it finds one; otherwise models run on the processor, which is slower. Memory decides model size: 8 GB suits small models (1–4B), 16 GB mid-size (up to ~9B), and 32 GB+ larger ones. The app shows exactly what fits.

A model says "too large" or "needs a newer engine" #

Too large: pick a smaller model, or lower the context length in Settings. Needs a newer engine: the model uses a new architecture your installed engine doesn't support yet. Update it with brew upgrade llama.cpp or pip3 install -U mlx-lm on a Mac, or winget upgrade ggml.llamacpp on Windows, or choose the model's other engine version.

"Engine not installed" / nothing happens when I chat #

Install an engine (step 1), then reopen LLMario. Models → Engines shows what LLMario found. llmario doctor gives full details.

Windows says "Windows protected your PC" #

The Windows preview isn't code-signed yet, so Microsoft SmartScreen shows this for new apps. Click More info → Run anyway. Only do this for the installer downloaded from this site or the GitHub releases page. You can check the download against its .sha256 file with Get-FileHash in PowerShell.

macOS says the app can't be opened #

Official releases are notarized by Apple and open normally. If you received a build from someone else, prefer the GitHub release or install from source.

Does it use Ollama or send data to a server? #

No. LLMario runs the open-source engines llama.cpp and MLX directly. It only goes online to download a model you choose from Hugging Face.

How do I remove it? #

Mac: delete LLMario from Applications. Windows: Settings → Apps → LLMario → Uninstall. Models and settings are in ~/.llmario (Windows: %USERPROFILE%\.llmario); delete that folder to free the space. Remove a source-built command-line tool with cargo uninstall llmario.

── more in #ai-products 4 stories · sorted by recency
── more on @llmario 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/llmario] indexed:0 read:7min 2026-10-02 · —