cd /news/large-language-models/run-a-4b-llm-model-locally-on-your-m… · home topics large-language-models article
[ARTICLE · art-93370] src=gist.github.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Run a 4B LLM model locally on your mac

A developer has published a step-by-step guide for running a 4B parameter LLM locally on a Mac, using Mozilla's llamafile and a fine-tuned Qwen3 model from Hugging Face. The process involves downloading llamafile v0.10.5 and the Qwen3-4B GGUF file, creating a shell script, and removing macOS quarantine attributes to run the model as a local server on port 8080.

read3 min views1 publishedAug 12, 2026

| Run a 4B LLM model locally on your mac, this has been tested on MacOS Tahoe | | | Step 1: | | | Lookup "llamafile github" or go to https://github.com/mozilla-ai/llamafile | | | Download the latest release, as of this it is llamafile v0.10.5 | | | Step 2: | | | Go to https://huggingface.co/models | | | Search for qween3 | | | Click on pramodlohra/Qween3_4B_thinking_finetune | | | Go to Files and versions | | | Download qwen3-4b-thinking-2507.Q4_K_M.gguf | | | Step 3: | | | Create a folder called qwen | | | Copy llamafile v0.10.5 and qwen3-4b-thinking-2507.Q4_K_M.gguf into the folder | | | Step 4: | | | In this folder, create a file called run_qwen.command | | | Write the following lines in it | | | #!/bin/bash | |

| cd "$(dirname "$0")" | |
| ./llamafile-0.10.5 --server --model ./qwen3-4b-thinking-2507.Q4_K_M.gguf | |

| Save the run_qwen.command file | | | Step 5: | | | Run chmod for llamafile llamafile-0.10.5 and run_qwen.command | | | chmod 755 llamafile-0.10.5 | | | chmod 755 run_qwen.command | | | Step 6: | | | Remove the quarantine flag from running this in the terminal | | | xattr -d com.apple.quarantine run_qwen.command | | | xattr -d com.apple.quarantine llamafile-0.10.5 | | | You can check whether quarantine is still present with: | | | xattr -l run_qwen.command | | | xattr -l llamafile-0.10.5 | | | If com.apple.quarantine doesn't appear, you're good. | | | NOTE: Don't disable Gatekeeper globally with commands such as sudo spctl --master-disable. | | | There's absolutely no need for that here; removing the quarantine attribute from these specific files is much safer. | | | Now you are ready to click on the run_qwen.command file and run the model locally | | | If everything is done correctly, this is what would show up in the terminal. | | | % /Users/santa/Downloads/qwen/run_qwen.command ; exit; | |

| 0.00.000.673 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg) | |
| 0.00.000.921 W srv llama_server: ----------------- | |

| 0.00.000.926 W srv llama_server: CORS is set to allow all origins ('*') and no API key is set | | | 0.00.000.926 W srv llama_server: this can be a security risk (cross-origin attacks) | | | 0.00.000.927 W srv llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655 | | | 0.00.000.927 W srv llama_server: ----------------- | | | 0.00.001.045 W srv llama_server: sandbox: disabled in GPU mode | | | 0.00.002.429 I srv load_model: model './qwen3-4b-thinking-2507.Q4_K_M.gguf' | | | 0.00.429.537 W load: control-looking token: 128247 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden | | | 0.04.590.135 I srv load_model: initializing, n_slots = 4, n_ctx_slot = 14080, kv_unified = 'true' | | | 0.04.600.756 I srv llama_server: model loaded | | | 0.04.600.778 I srv llama_server: listening on http://127.0.0.1:8080 | | | You can start using the model on 127.0.0.1:8080, Enjoy! |

── more in #large-language-models 4 stories · sorted by recency
── more on @mozilla 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/run-a-4b-llm-model-l…] indexed:0 read:3min 2026-08-12 ·