cd /news/large-language-models/ollama-for-managing-local-language-m… · home › topics › large-language-models › article
[ARTICLE · art-142496] src=kdnuggets.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Ollama for Managing Local Language Models: A KDnuggets Cheat Sheet

KDnuggets published a cheat sheet on managing local language models with Ollama, the tool that pulls model weights, runs an HTTP server on port 11434, and exposes an OpenAI-shaped endpoint on the user's own machine. The guide highlights the `ollama ps` command, whose PROCESSOR column shows anything under 100% GPU as spillover to CPU, and notes that on macOS environment variables must be set via `launchctl setenv` because the desktop app never reads `export` lines in a `.zshrc`. It also covers passing a JSON schema as `format` to constrain decoding versus the looser `"json"` string, plus disk-management commands, the `/api/chat` and `/api/embed` endpoints, the `/v1/` compatibility layer, Modelfiles, and environment variables controlling model load duration and concurrency.

read2 min views1 publishedSep 30, 2026

Ollama pulls model weights, keeps an HTTP server on port 11434, and hands any client an OpenAI-shaped endpoint pointed at your own machine. Learn how to manage, configure, and optimize using Ollama right here.

Running a language model on your own hardware has become straightforward enough that the interesting problems have moved elsewhere. Ollama pulls model weights, keeps an HTTP server on port 11434, and hands any client an OpenAI-shaped endpoint pointed at your own machine. The good news? Getting that far only takes one command.

What follows is a different kind of question: whether or not the model plus its context still fits in the memory you have. ollama ps is the command that answers it. Alongside what models are currently resident, it displays a PROCESSOR column, and anything under 100% GPU means part of the model has spilled to CPU and generation has slowed to a crawl. It also shows the context that has been allocated, which may not be the number you had expected or asked for. Knowing how and where to defuse local model sizing decisions resolve quickly when you serve with Ollama.

You can download our latest cheat sheet to keep this info handy as you build and experiment with Ollama.

You'll want to know how to interact with your local OS for easy management as well. For example, the desktop app is launched by the system rather than your shell, so it never sees export lines in a .zshrc. Configuration that looks correct in a terminal simply has no effect. Variables have to be set through launchd via launchctl setenv, at least on macOS. This is the single most common reason a context length or a model directory refuses to change, though everything looks correct to a newcomer.

Chances are you want to know the nuances of dealing with structured output in Ollama-served models. Passing a JSON schema as format constrains decoding to that shape, so the reply parses every time rather than most of the time. However, the bare string "json" is the looser version: valid JSON, but no promise about which keys arrive.

The rest of the cheat sheet rounds out the fundamentals. There is the disk-management commands, the endpoints, /api/chat and /api/embed, plus the /v1/ compatibility layer that lets an existing OpenAI client switch to localhost otherwise unchanged. And there are Modelfiles for saving a base model with your own defaults, along with the environment variables governing how long models stay loaded and how many run at once.

Don't even think about it. Download the cheat sheet now, and get those optimized local models up and running for fun and profit.

── more in #large-language-models 4 stories · sorted by recency
── more on @ollama 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ollama-for-managing-…] indexed:0 read:2min 2026-09-30 · —