{"slug": "self-hosted-langfuse-tracing-7-of-my-ai-agents-and-clickhouse-logging-itself", "title": "Self-hosted Langfuse: tracing 7% of my AI agents, and ClickHouse logging itself", "summary": "A developer self-hosting Langfuse discovered that tracing was enabled on only one of ten agent profiles, capturing roughly 7% of traffic — 618 events across 79 traces in four weeks — while the coding agent alone made 522 model calls in 14 days. A 25-minute failed turn that produced a fabricated article summary was traced to a stale long-lived session that never loaded the reply skill, not to the web tools, which together took only 4.6 seconds versus 1,517 seconds of model time. The gap went unnoticed because the tracing plugin fails open when API keys are absent from a profile's secret scope, producing no error or warning.", "body_md": "On Discord I told my agent `post` followed by a comment id. That's the approval step for a dev.to reply it had drafted: read the draft, say \"post\", done.\n\nNothing got posted. The turn took **25 minutes** and came back with a confident, detailed summary of the article. It was made up.\n\nI self-host Langfuse to trace my agents, so I went and looked at the trace. Then I checked what else it had seen. It had been watching about 7% of my agents' traffic.\n\nHere is the trace of that one turn:\n\n```\nHermes turn                      1526.3 s\n  LLM call 1   qwen3.5-64k        569.4 s\n  Tool: web_search                  0.8 s\n  Tool: web_search                  0.9 s\n  LLM call 2   qwen3.5-64k        207.4 s\n  Tool: web_extract                 2.9 s\n  LLM call 3   qwen3.5-64k        740.2 s\n```\n\nTools: 4.6 seconds in total. Model: 1,517 seconds. The local model spent 25 minutes thinking its way to the wrong answer, with three short trips to the web in between.\n\nWithout the trace I'd have blamed the web tools. A turn that searches the web and takes 25 minutes *feels* like a network problem. The trace says the network was the fastest thing in the room.\n\nThe actual cause had nothing to do with speed. The Discord conversation was one long-lived session, started days before the reply skill existed. The agent reads its list of skills once, at session start. So as far as that session knew, there was no skill for posting a reply. Asked to \"post\" something it had no tool for, it did what models do and improvised: searched the web for the article, extracted the page, and wrote a plausible summary instead.\n\nThe fix was a fresh session. But it made me curious about what else the traces could tell me.\n\nPulling latency per model from the same Langfuse data (the default profile, all history):\n\n| Model | Calls | p50 | Slowest | \n|---|---|---|---|\n| qwen/qwen3.7-flash | 92 | 4.3 s | 180.5 s | \n| qwen3.5-64k (local, 6 GB GPU) | 91 | 125.2 s | 900.3 s | \n| deepseek-v4-flash | 65 | 6.7 s |  | \n| mistral-nemo (local) | 24 | 106.6 s | 1,802.9 s | \n| hermes3-64k | 15 | 2.5 s |  | \n\nOn the tool side, `terminal` ran 135 times with a median of 0.6 s, and the slowest single tool call was a `search_files` at 120.6 s.\n\nNone of this is shocking. A local model on a 6 GB card is slow, and cheap cloud models are fast. But \"slow\" and \"125 seconds median, about 30 times slower than the fastest cloud option\" are different sentences, and only one of them helps decide which job runs where.\n\nThen I noticed the call counts. 91 calls for the local model. That seemed low for a system I lean on all day.\n\nLangfuse held **618 events across 79 traces in four weeks**. That's about three traces a day.\n\nThe agent keeps its own session database, so I counted from the other side. The last 14 days alone:\n\nTracing was enabled on exactly **one of ten** profiles: the default chat one, with 91 calls. That's roughly 7% of the traffic.\n\nThe untraced ones were the ones doing the work:\n\n| Profile | Model calls (14 days) | Traced | \n|---|---|---|\n| coder | 522 | no | \n| sec | 390 | no | \n| itadmin | 172 | no | \n| writer | 120 | no | \n| default (chat) | 91 | yes | \n| reviewer | 55 | no | \n| scout | 29 | no | \n| others | small | no | \n\nThe coding agent alone made more than five times as many calls as the only profile I was watching.\n\nThe tracing plugin is enabled per profile, and it reads its Langfuse API keys through that profile's own secret scope. That's deliberate: profile B's traces shouldn't be shipped with profile A's keys.\n\nWhen I set Langfuse up, I enabled the plugin and added the keys on the default profile, and moved on. Every specialist profile then ran without tracing. No error. No warning in a log. The plugin fails open by design (no keys means its hooks quietly do nothing), which is right for a plugin, and exactly why nobody noticed.\n\nA monitoring tool that's switched off looks identical to one that has nothing to report.\n\nThere was a second, smaller gap. One-shot jobs launched in \"safe mode\" skip plugins entirely, and the reply drafter ran that way because it's quicker. That one is closed now: the drafter sends its own traces through the Langfuse SDK, so it no longer depends on the plugin.\n\nLangfuse v3 onwards stores traces in ClickHouse. Postgres holds users and projects, Redis handles the queue, MinIO stores blobs. So no, you can't drop ClickHouse to save resources; it is the trace store.\n\nClickHouse was using **6.03 GiB** of disk. The actual Langfuse trace data in it was **2.2 MiB**.\n\nThe rest was ClickHouse's own diagnostic logging:\n\n| System table | Size | Rows | \n|---|---|---|\n| `system.trace_log` | 3.38 GiB | 174.6 million | \n| `system.text_log` | 1.19 GiB |  | \n| `system.part_log` | 492 MiB |  | \n| `system.asynchronous_metric_log` | 485 MiB | 949 million | \n| `system.metric_log` | 461 MiB |  | \n\nThat's roughly 2,700 times more bytes about ClickHouse than about my agents.\n\nThose tables aren't written once and forgotten. They're written continuously. The node this runs on keeps its guests on a single 5400 rpm HDD, and it already has a roughly 15-minute IO storm after every reboot. The Langfuse container is the one I stop to relieve it. So the trace store I'd barely been using was grinding that disk all day, recording its own profiler samples.\n\nFor each of the ten profiles:\n\n`environment` to the profile name.\nThe third step is the one that matters for later. Traces carry no profile tag of their own, so without it every trace lands in one undifferentiated pile. With it, \"cost and latency per agent role\" is a filter in the UI rather than a spreadsheet exercise.\n\nThen I restarted the gateway and sent a one-word probe on two profiles. Traces arrived tagged `legal` (on a cloud flash model) and `writer` (on the local qwen3.5-64k). That's the test I should have run the first time.\n\nClickHouse lets you override server config with a file dropped into `/etc/clickhouse-server/config.d/`. This one removes twelve system log tables, keeps `query_log` with a 7-day TTL (it's genuinely useful when a query is slow), and turns the server logger down to warnings:\n\n```\n<clickhouse>\n  <trace_log remove=\"1\"/>\n  <text_log remove=\"1\"/>\n  <metric_log remove=\"1\"/>\n  <asynchronous_metric_log remove=\"1\"/>\n  <part_log remove=\"1\"/>\n  <processors_profile_log remove=\"1\"/>\n  <query_thread_log remove=\"1\"/>\n  <query_views_log remove=\"1\"/>\n  <query_metric_log remove=\"1\"/>\n  <opentelemetry_span_log remove=\"1\"/>\n  <session_log remove=\"1\"/>\n  <latency_log remove=\"1\"/>\n  <query_log>\n    <ttl>event_date + INTERVAL 7 DAY DELETE</ttl>\n  </query_log>\n  <logger><level>warning</level></logger>\n</clickhouse>\n```\n\nIt's mounted read-only into that directory through the compose file, and the container was recreated. Checks afterwards:\n\n`trace_log` row count unchanged over 30 seconds, so it has stopped writing\nOne honest caveat: removing a log table from config stops ClickHouse writing to it, but it doesn't delete what's already there. The old 5.9 GiB of log tables was still on disk. I dropped those deliberately, as a separate step, rather than bundling a destructive delete into a config change. The Langfuse container went from 15 GB of disk used to 7.7 GB, with the trace data intact.\n\nIt would have been easy to conclude Langfuse was dead weight: a few gigabytes of disk for 79 traces.\n\nBut the estate runs around ten agent profiles: a chief of staff that I chat with, a coder, security, IT admin, writer, scout, reviewer, legal, analyst and a few more. They run on a mix of local models on a 6 GB GPU and cheap paid cloud models with a fallback chain. The questions I actually have about it are ones only traces answer:\n\nThe rule I run the estate by is \"scripts gather facts, models never do\". Models get to summarise what scripts measured; they never go and measure it themselves. Langfuse is the fact-gatherer for the agents themselves. It just can't gather facts from agents it isn't connected to.\n\nA small script now reads the agents' own session records and Langfuse side by side, and every Monday it posts a report: per role, the turns, model calls, tokens, cost, errors and the share that was traced; per model, p50 and p95 latency and cost; how turn time splits between model and tools; the five slowest turns; and tool errors. It's built to answer these questions:\n\n`system.parts` grouped by table. If `asynchronous_metric_log` dwarf your actual data, a `config.d` override fixes it in a few lines.\nOn 22 September the tracing server's address changed, and all 12 specialist agents' configs were repointed at the new one. When it moved back on 23 September, only the main config was updated. Every specialist silently stopped tracing. I found and fixed it on 24 September.\n\nNo error, no warning, nothing in a log. It's the same fail-open silence this whole post is about, and the reason point 1 above isn't a one-off check.\n\n🤖 *Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.*", "url": "https://wpnews.pro/news/self-hosted-langfuse-tracing-7-of-my-ai-agents-and-clickhouse-logging-itself", "canonical_source": "https://dev.to/c1-anderson/self-hosted-langfuse-tracing-7-of-my-ai-agents-and-clickhouse-logging-itself-3e80", "published_at": "2026-10-01 08:00:15+00:00", "updated_at": "2026-10-01 08:14:09.835388+00:00", "lang": "en", "topics": ["ai-agents", "mlops", "ai-tools", "large-language-models"], "entities": ["Langfuse", "ClickHouse", "Discord", "dev.to", "qwen3.5-64k", "deepseek-v4-flash", "mistral-nemo", "hermes3-64k"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/self-hosted-langfuse-tracing-7-of-my-ai-agents-and-clickhouse-logging-itself", "markdown": "https://wpnews.pro/news/self-hosted-langfuse-tracing-7-of-my-ai-agents-and-clickhouse-logging-itself.md", "text": "https://wpnews.pro/news/self-hosted-langfuse-tracing-7-of-my-ai-agents-and-clickhouse-logging-itself.txt", "jsonld": "https://wpnews.pro/news/self-hosted-langfuse-tracing-7-of-my-ai-agents-and-clickhouse-logging-itself.jsonld"}}