{"slug": "compared-llama3-2-1b-vs-llama3-2-3b-memory-footprint", "title": "Compared llama3.2:1b vs llama3.2:3b Memory Footprint", "summary": "Ollama's llama3.2:3b model consumes 2.5 GB of memory (2.47 GB RSS) versus 1.5 GB (1.24 GB RSS) for the 1b model, showing that tripling parameters roughly doubles memory footprint. The test also revealed that the larger model without a custom system prompt reverts to default kubectl-based answers, confirming that domain-specific behavior requires prompt engineering regardless of model size.", "body_md": "**Context:** The number in a model name like `1b` or `3b` refers to parameters — roughly, the tunable values inside the model that encode what it’s learned. More parameters generally means better reasoning and more nuanced answers, at the cost of more memory and slower responses. One thing that trips people up early with Ollama: typing `ollama run <model>` drops you into an interactive chat session (marked by the `>>>` prompt), which is a different context from your regular shell. Commands like `ollama ps` only work back in a normal terminal prompt — typed inside the chat session, they get sent to the model as a question instead of running as a command.\n\n**Ran:** Pulled `llama3.2:3b` (the 3-billion-parameter sibling of Entry 01’s `1b` model), loaded it into memory, and captured the same `ollama ps` / `ps aux | grep ollama` numbers for a direct comparison. Along the way, typed `ollama ps` inside the chat session by mistake — got a confused response from the model instead of the process table, a good real-world example of the shell-vs-chat distinction above. Then asked the same test question from [Entry 02](https://pipelineandprompts.com/posts/02-oc-cli-mentor-system-prompt/) (“how do I check the status of pods in my namespace?”) — this time against the plain `llama3.2:3b` model, not the constrained `oc-mentor` build from Entry 02.\n\n**Result:**\n\n|  | `1b` (Entry 01) | `3b` | \n|---|---|---|\n| `ollama ps` size | 1.5 GB | 2.5 GB | \n| Process RSS | ~1.24 GB | ~2.47 GB | \n| Parameters | 1B | 3B | \n\nTripling the parameter count roughly doubled the memory footprint — not a 1:1 scaling, which is worth remembering when estimating resource requests for larger models.\n\nOn the question test: since this run used the plain `3b` model rather than the `oc-mentor` Modelfile from Entry 02, the answer came back as a verbose, multi-option explanation using `kubectl` — not `oc`, and not the single-command format Entry 02 enforced. That’s not a knock on the bigger model; it’s a reminder that the constrained, single-command behavior from Entry 02 came from the system prompt, not from model size. A bigger base model without that constraint just reverts to its default training bias (which, unsurprisingly, leans `kubectl` over `oc`).\n\n**Takeaway:** More parameters bought roughly 2x memory for 3x the parameter count — a useful data point for future sizing — but it didn’t buy domain-specific behavior on its own. Getting oc-only, single-command answers still requires the system prompt from Entry 02, regardless of model size.\n\n``` bash\n$ ollama ps\nNAME           ID              SIZE      PROCESSOR    CONTEXT    UNTIL\nllama3.2:3b    a80c4f17acd5    2.5 GB    100% GPU     4096       4 minutes from now\n\n$ ps aux | grep ollama\nflyers  94626  0.2  15.1  438021552  2531232  ??  S  llama-server --model ... -c 4096\nflyers   1905  0.0   0.4  436904528    60864  ??  S  ollama serve\n```\n\n*Correction (Aug 19, 2026): the `1b` baseline referenced from Entry 01 is Q8_0 quantization, not an unspecified default — see [Entry 04](https://pipelineandprompts.com/posts/04-quantization-q4-q8-fp16/) for the full breakdown across quantization levels.*", "url": "https://wpnews.pro/news/compared-llama3-2-1b-vs-llama3-2-3b-memory-footprint", "canonical_source": "https://pipelineandprompts.com/posts/03-1b-vs-3b-memory-comparison/", "published_at": "2026-08-14 00:00:00+00:00", "updated_at": "2026-09-07 17:30:36.814087+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure"], "entities": ["Ollama", "llama3.2:3b", "llama3.2:1b"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/compared-llama3-2-1b-vs-llama3-2-3b-memory-footprint", "markdown": "https://wpnews.pro/news/compared-llama3-2-1b-vs-llama3-2-3b-memory-footprint.md", "text": "https://wpnews.pro/news/compared-llama3-2-1b-vs-llama3-2-3b-memory-footprint.txt", "jsonld": "https://wpnews.pro/news/compared-llama3-2-1b-vs-llama3-2-3b-memory-footprint.jsonld"}}