# Run a 4B LLM model locally on your PC, this has been tested on Windows 10

> Source: <https://gist.github.com/barshasantak/7ea564a5e5ddaf7e955c85d73360eb94>
> Published: 2026-08-12 23:08:24+00:00

| Run a 4B LLM model locally on your PC, this has been tested on Windows 10 | |
| Step 1: | |
| Lookup "llamafile github" or go to https://github.com/mozilla-ai/llamafile | |
| Download the latest release. As of this writing, the latest version is llamafile v0.10.5 | |
| The file would be llamafile-0.10.5 | |
| Rename the file to llamafile-0.10.5.exe | |
| Step 2: | |
| Go to https://huggingface.co/models | |
| Search for qween3 | |
| Click on pramodlohra/Qween3_4B_thinking_finetune | |
| Go to Files and versions | |
| Download qwen3-4b-thinking-2507.Q4_K_M.gguf | |
| Step 3: | |
| Create a folder called qwen | |
| Copy llamafile-0.10.5.exe and qwen3-4b-thinking-2507.Q4_K_M.gguf into this folder | |
| Step 4: | |
| In this folder, create a file called run_qwen.cmd | |
| Write the following lines in it | |
| @echo off | |
| cd /d "%~dp0" | |
| llamafile-0.10.5.exe --server --model "qwen3-4b-thinking-2507.Q4_K_M.gguf" | |
| pause | |
| Save the run_qwen.cmd file | |
| Now you are ready to click on the run_qwen.command file and run the model locally | |
| If everything is done correctly, something like this is what would show up in the command window | |
| 0.00.000.673 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the `-lv N` CLI arg) | |
| 0.00.000.921 W srv llama_server: ----------------- | |
| 0.00.000.926 W srv llama_server: CORS is set to allow all origins ('*') and no API key is set | |
| 0.00.000.926 W srv llama_server: this can be a security risk (cross-origin attacks) | |
| 0.00.000.927 W srv llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655 | |
| 0.00.000.927 W srv llama_server: ----------------- | |
| 0.00.001.045 W srv llama_server: sandbox: disabled in GPU mode | |
| 0.00.002.429 I srv load_model: loading model './qwen3-4b-thinking-2507.Q4_K_M.gguf' | |
| 0.00.429.537 W load: control-looking token: 128247 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden | |
| 0.04.590.135 I srv load_model: initializing, n_slots = 4, n_ctx_slot = 14080, kv_unified = 'true' | |
| 0.04.600.756 I srv llama_server: model loaded | |
| 0.04.600.778 I srv llama_server: listening on http://127.0.0.1:8080 | |
| You can start using the model on 127.0.0.1:8080, Enjoy! |
