cd /news/large-language-models/graphrag-with-ollama-in-10-minutes-t… · home › topics › large-language-models › article
[ARTICLE · art-141082] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

GraphRAG with Ollama in 10 Minutes: The Local-First Quickstart

A developer published a 10-minute quickstart for running GraphRAG entirely locally using Ollama and the open-source Chaos Cypher stack, requiring no cloud API keys. The walkthrough covers pulling the qwen3:30b-instruct model, launching the containerized service via Docker, uploading a document, and watching entity extraction build a knowledge graph that answers questions with click-through citations to specific source chunks.

by read3 min views2 publishedSep 28, 2026

You don't need a cloud API key to try GraphRAG. Ten minutes, one document, and a laptop with a decent GPU (or none at all, for the first few steps) is enough to see a full knowledge graph form and answer a question with a citation you can click through to the source page.

This is the fast version. If you want the reasoning behind running everything locally -- privacy, VRAM presets, multi-GPU load balancing -- that's in Build a Private AI Knowledge Graph That Never Leaves Your Machine. This post is the "just show me" version: install, upload, watch the graph build, ask a question, done.

By the end of this walkthrough you'll have:

Nothing leaves your machine. No OpenAI key, no Anthropic key, no usage bill.

Grab Ollama from ollama.com for your platform, then pull the default chat/extraction model in a terminal:

ollama pull qwen3:30b-instruct

This is the long pole of the whole setup -- roughly 18-20 GB. Kick it off now and let it run in the background while you do the next steps; indexing and search don't need it, only extraction and chat do.

If you're on a smaller GPU or just want to see the pipeline move faster for this first run, point ollama_chat_model at a smaller qwen3 tag instead -- you can always switch to the full model afterward from Settings.

docker run -d --name chaoscypher \
  -p 80:80 \
  -p 443:443 \
  -v chaoscypher-data:/data \
  ghcr.io/chaoscypherinc/chaoscypher:latest

On Linux Docker Engine (not Docker Desktop), add --add-host=host.docker.internal:host-gateway so the container can reach Ollama on your host -- Docker Desktop resolves this automatically.

Open http://localhost. A startup page shows each service (Nginx, Cortex, Valkey, Neuron) coming online -- usually 30-60 seconds. Set a username and password on first run, and you land on the Dashboard.

Go to Sources in the sidebar and drag in a PDF, DOCX, or text file -- a 10-20 page document is a good first run so you can watch the whole pipeline complete in a couple of minutes.

Indexing starts immediately (chunking + embedding for search) and finishes in about 30 seconds per 100 pages -- this step downloads a small embedding model (~600 MB) the first time, but doesn't need Ollama at all. Once indexing finishes, a Review dialog proposes an extraction domain (technical, medical, legal, and so on) detected from the content. Click Confirm to queue entity extraction.

If the Ollama model pull from Step 1 is still running, the source will sit at extracting until it finishes -- indexing and search still work while you wait.

Once extraction commits, open Graph in the sidebar. You'll see nodes for the people, organizations, concepts, and events the model found in your document, connected by edges representing the relationships it inferred between them. For a 100-page document on a 30B-class model, expect roughly 5-10 minutes of extraction time.

Click any node to see its properties, its connections, and the source text it was extracted from -- that link back to source evidence is the same mechanism the chat citations use in the next step.

Open Chat, start a new conversation, and ask something specific about the document you just uploaded. Chaos Cypher retrieves the relevant chunks and graph context, hands them to your local model, and streams back an answer with citations attached to the sentences that came from your source.

Click a citation and it jumps straight to the exact chunk in the source document the sentence was grounded in -- not a link to "the document," the specific paragraph. That's the difference between an AI answer you have to take on faith and one you can verify in a click.

That's the whole loop: install, upload, extract, ask, verify. A few directions once you're comfortable with it:

In plain English: you can run a real knowledge graph, end to end, on your own machine, in about the time it takes to read this post.

── more in #large-language-models 4 stories · sorted by recency
── more on @ollama 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/graphrag-with-ollama…] indexed:0 read:3min 2026-09-28 · —