{"slug": "local-llm-observability", "title": "Local LLM Observability", "summary": "LLMxRay 0.6.0 launched as a free, local observability tool for Ollama models, adding a Cache Lab feature that the project reports delivered 3.5x faster prefill in its test. The tool runs via `npx llmxray` or a Docker image on port 5174 and provides real-time token streaming with confidence coloring, latency percentiles (P50/P95/P99), a 7-category error classifier, and side-by-side comparison of up to 4 model slots. It requires Ollama running locally with at least one model pulled, such as llama3.2, and includes a Protocol Observatory for comparing Ollama native, OpenAI-compatible, and Anthropic-compatible APIs.", "body_md": "**See what your AI is actually doing.**\n\n  Real-time token streaming, quality analysis, performance profiling, and cost tracking\n\n  for local LLMs. No cloud. No API keys. No cost.\n\n  🌐 **English** •\n  [Français](https://github.com/LogneBudo/llmxray/blob/master/README.fr.md) •\n  [中文](https://github.com/LogneBudo/llmxray/blob/master/README.zh-CN.md) •\n  [العربية](https://github.com/LogneBudo/llmxray/blob/master/README.ar.md) •\n  [Srpski](https://github.com/LogneBudo/llmxray/blob/master/README.sr.md)\n\n[Quick Start](#quick-start) •\n  [Features](#features) •\n  [Screenshots](#screenshots) •\n  [Who Is This For](#who-is-this-for) •\n  [Changelog](https://github.com/LogneBudo/llmxray/blob/master/CHANGELOG.md)\n\n**Chat Diagnostics** — Watch tokens arrive one by one, coloured by speed. See where the model hesitates, and what it was unsure about.\n\n**Cache Lab** *(new in 0.6.0)* — Your system prompt is probably being recomputed from scratch every turn. Find out in 30 seconds, and measure the fix. *(3.5x faster prefill in our test.)*\n\n**Compare** — The same prompt against two models, or two temperatures, side by side. Stop guessing which setting was actually better.\n\n**Surgical Benchmark** — Score models on real token logprobs, not on whether the answer happened to look right.\n\n**Protocol Observatory** — See how the same local model replies through Ollama's native, OpenAI-compatible and Anthropic-compatible APIs, and exactly where they disagree.\n\n**Knowledge Base** — Drop in PDFs and documents, chunk them, and see what RAG actually retrieves before the model ever sees it.\n\nPlus embeddings, cost tracking, analytics, a tool builder, and a local history database. [Full feature list below.](#features)\n\n**One command. 30 seconds.**\n\n```\nnpx llmxray\n```\n\nOr with Docker:\n\n```\ndocker run -p 5174:5174 djovaneli/llmxray\n```\n\nOpen **[http://localhost:5174](http://localhost:5174)** and start chatting. That's it.\n\n**Prerequisite:** [Ollama](https://ollama.com/download) running locally with at least one model pulled (`ollama pull llama3.2`).\n\nYou run a local LLM. You chat with it. But what actually happened?\n\n- How fast was each token? Which ones was the model confident about?\n- Is the response quality degrading over long conversations?\n- What would this have cost if you ran it in the cloud?\n- Is the model repeating itself? Refusing? Generating gibberish?\n- How does temperature 0.3 compare to 0.9 on the *same* prompt?\n\n**LLMxRay answers all of these, visually, in real time, for free.**\n\nChat with any Ollama model and watch tokens arrive with **confidence coloring** — each token is tinted based on generation speed. Supports markdown, multi-turn conversations, file attachments, vision models, and slash commands. For reasoning models, set the thinking budget per conversation — off, model's choice, or an explicit low / medium / high / max effort.\n\nEvery response is automatically analyzed. Colored badges appear only when something is wrong:\n\n- **Repetition** — excessive repeated phrases (4-gram analysis)\n- **Refusal** — \"as an AI language model\" and 7 other patterns\n- **Gibberish** — high non-ASCII ratio\n- **Empty** — fewer than 10 words\n- **Truncation** — hit the token limit without finishing\n\nUp to **4 slots** with independent model, temperature, and system prompt. Features include side-by-side streaming, word-level diff highlighting, metrics comparison, and one-click presets (Temperature Sweep, Deterministic Pair, Language Compare with Token Tax visualization).\n\n- **Latency percentiles** (P50/P95/P99) for duration and TTFT\n- **Error intelligence** — 7-category classifier with timeline\n- **Usage heatmap** — 7x24 grid of your active hours\n- **Settings impact** — temperature vs tokens/sec scatter plots\n- **Cold vs warm start** tracking with model load history\n\nToken usage per model/day with estimated cloud-equivalent pricing. See what you're *saving* by running locally.\n\nTest model knowledge with multi-choice question suites. Uses real logprobs via OpenAI-compatible endpoint for accurate confidence measurement. Build custom suites visually or let AI generate them from a topic.\n\nEmbed text, visualize vectors, measure cosine similarity. Request a narrower output vector to see what Matryoshka truncation costs in similarity. Build a local knowledge base from PDFs, DOCX, and CSV — chunked, embedded, and searchable. All stored in IndexedDB. Zero cost.\n\nDrag-and-drop node canvas for building tool definitions. Bidirectional code sync (edit nodes or TypeScript — both update). Probe APIs, auto-generate schemas, test with live execution.\n\nCode completion for Qwen-Coder, CodeLlama, Codestral, DeepSeek-Coder, and StarCoder. Two textareas (prefix / suffix), the model fills the gap. Uses Ollama's `suffix` field on `/api/generate`. Stitched preview shows the result as it would appear in your editor.\n\nFind out why your prompt misses the model’s KV cache, and measure what it costs every turn. A local model reuses its cache only while the prompt still matches from the very first token, so a single timestamp near the top forfeits everything below it. The lab finds the values that change between turns, shows the exact point where reuse dies, and then **measures** — sending each layout twice with a changed value, against your own daemon — what moving them to the end actually saves. Measured on a real 324-token prompt: **4 tokens reused and 64.6 ms of prefill with the timestamp at the front, 290 reused and 18.6 ms with it at the back. 3.5x faster, same words.** Requires Ollama 0.33.3+.\n\nFire the same prompt through Ollama's three serving protocols — **native** `/api/chat`, **OpenAI-compat** `/v1/chat/completions`, and **Anthropic-compat** `/v1/messages` — in parallel against your local model. Side-by-side streaming, per-protocol metrics, and an envelope-diff tab that shows how each protocol frames finish reasons, token counts, and error envelopes. No cloud, no API keys — all three endpoints are local on `localhost:11434`.\n\nCurate training data from your conversations. Tag, review, and export as JSONL for fine-tuning.\n\nEvery experiment (benchmarks, comparisons, chats, training pairs) is automatically archived in a queryable IndexedDB database with filters, trends, exports, and retention policies.\n\nFull translations in English, French, Serbian (Latin + Cyrillic), Chinese, and Arabic. RTL layout support. Community scaffolds for Hebrew and Japanese.\n\nTested and verified against **Ollama 0.34.x** (verified on 0.34.0, September 2026). LLMxRay uses these Ollama endpoints:\n\n| Endpoint | Used for | \n|---|---|\n| `/api/chat` | Streaming chat (NDJSON, with `tools` ,`think` effort levels,`format` schema) | \n| `/api/generate` | Generation + Fill-in-the-Middle via `suffix` | \n| `/api/tags` | Model list + capabilities, context length, and embedding width | \n| `/api/show` | Parameters, template, license, and architecture metadata | \n| `/api/embed` | Vector embeddings for RAG, with optional `dimensions` truncation | \n| `/api/pull` ,`/api/delete` ,`/api/ps` ,`/api/version` | Model management + status | \n| `/v1/chat/completions` | OpenAI-compat path used by Surgical Benchmark for real logprobs and usage totals | \n| `/v1/messages` | Anthropic-compat path used by Protocol Observatory | \n\n**Compatible with:** Ollama 0.20 and newer (older versions work for chat/generate but lack `think` and JSON-schema `format`). **Recommended:** Ollama 0.33.3+ — prompt-cache reuse is reported (`prompt_eval_cached_count`, and `usage.prompt_tokens_details.cached_tokens` on the OpenAI-compatible endpoint), so prefill throughput is measured over the tokens actually evaluated. From 0.32: capabilities and context length arrive with the model listing, `think` accepts graded effort levels, and embeddings accept a `dimensions` width.\n\n| **Chat with token streaming and confidence** | **Model comparison — side by side** | \n| **Session deep dive — metrics and timing** | **Benchmark with confidence radar** | \n| **Embeddings — cosine similarity** | **System monitor — hardware and Ollama status** | \n\n| You are... | LLMxRay helps you... | \n|---|---|\n| **Developer** | Debug prompts, profile latency, compare models, inspect tool calls, track costs | \n| **Researcher** | Run controlled experiments with consistent settings across models and temperatures | \n| **Student / Educator** | Explore model behavior visually — built-in Educators Kit with 9 interactive modules | \n| **AI team lead** | Understand quality trends, error patterns, and resource usage across your local fleet | \n\n```\nnpx llmxray\nnpx llmxray --port 3000\nnpx llmxray --ollama-url http://192.168.1.50:11434\ndocker run -p 5174:5174 djovaneli/llmxray\ndocker run -p 5174:5174 -e OLLAMA_URL=http://host.docker.internal:11434 djovaneli/llmxray\ngit clone https://github.com/LogneBudo/llmxray.git\ncd llmxray\nnpm install\nnpm run dev     # http://localhost:5173\n```\n\n| Layer | Technology | \n|---|---|\n| Framework | Vue 3.5 + Composition API | \n| Language | TypeScript 5.9 (strict) | \n| Build | Vite 7.3 | \n| Styling | Tailwind CSS 4.2 | \n| State | Pinia 3 (store-per-concern) | \n| Charts | Chart.js 4, D3.js 7 | \n| Canvas | Vue Flow (visual node editor) | \n| Code Editor | CodeMirror 6 | \n| Storage | IndexedDB (browser-native) | \n| LLM Backend | Ollama (local) | \n\n**Streaming** — Reads Ollama NDJSON via `fetch()` + `ReadableStream`. Tokens update the UI reactively through Pinia stores.\n\n**Token confidence** — Approximated from inter-token latency (faster = more confident). Clearly labeled as approximation. Benchmarks use real logprobs via OpenAI-compatible endpoint.\n\n**Store-per-concern** — Each domain has its own Pinia store: tokens, sessions, metrics, reasoning, comparison, embeddings, quality, cost, and more.\n\n**Hardware detection** — Custom Vite plugin queries the OS directly (PowerShell/proc/sysctl) for accurate hardware specs.\n\n| Command | What it does | \n|---|---|\n| `npm run dev` | Dev server (port 5173) | \n| `npm run build` | Type-check + production build | \n| `npm run test` | Unit tests (Vitest) | \n| `npm run test:e2e` | End-to-end (Playwright) | \n\nContributions welcome! See [CONTRIBUTING.md](https://github.com/LogneBudo/llmxray/blob/master/CONTRIBUTING.md) for setup and guidelines.\n\n**Community translations especially welcome** — scaffold files ready for Hebrew and Japanese.\n\n**LLMxRay** is a trademark of Ivan Stankovic ([LogneBudo](https://github.com/LogneBudo)). See [TRADEMARK.md](https://github.com/LogneBudo/llmxray/blob/master/TRADEMARK.md).\n\n**If LLMxRay helps you understand your AI better, consider giving it a star.**\n\n  It helps others discover the project.", "url": "https://wpnews.pro/news/local-llm-observability", "canonical_source": "https://github.com/LogneBudo/llmxray", "published_at": "2026-09-30 13:11:36+00:00", "updated_at": "2026-09-30 13:20:04.774313+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "mlops", "developer-tools", "ai-infrastructure"], "entities": ["LLMxRay", "Ollama", "OpenAI", "Anthropic", "llama3.2", "Docker"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/local-llm-observability", "markdown": "https://wpnews.pro/news/local-llm-observability.md", "text": "https://wpnews.pro/news/local-llm-observability.txt", "jsonld": "https://wpnews.pro/news/local-llm-observability.jsonld"}}