Self-hosted Langfuse: tracing 7% of my AI agents, and ClickHouse logging itself A developer self-hosting Langfuse discovered that tracing was enabled on only one of ten agent profiles, capturing roughly 7% of traffic — 618 events across 79 traces in four weeks — while the coding agent alone made 522 model calls in 14 days. A 25-minute failed turn that produced a fabricated article summary was traced to a stale long-lived session that never loaded the reply skill, not to the web tools, which together took only 4.6 seconds versus 1,517 seconds of model time. The gap went unnoticed because the tracing plugin fails open when API keys are absent from a profile's secret scope, producing no error or warning. On Discord I told my agent post followed by a comment id. That's the approval step for a dev.to reply it had drafted: read the draft, say "post", done. Nothing got posted. The turn took 25 minutes and came back with a confident, detailed summary of the article. It was made up. I self-host Langfuse to trace my agents, so I went and looked at the trace. Then I checked what else it had seen. It had been watching about 7% of my agents' traffic. Here is the trace of that one turn: Hermes turn 1526.3 s LLM call 1 qwen3.5-64k 569.4 s Tool: web search 0.8 s Tool: web search 0.9 s LLM call 2 qwen3.5-64k 207.4 s Tool: web extract 2.9 s LLM call 3 qwen3.5-64k 740.2 s Tools: 4.6 seconds in total. Model: 1,517 seconds. The local model spent 25 minutes thinking its way to the wrong answer, with three short trips to the web in between. Without the trace I'd have blamed the web tools. A turn that searches the web and takes 25 minutes feels like a network problem. The trace says the network was the fastest thing in the room. The actual cause had nothing to do with speed. The Discord conversation was one long-lived session, started days before the reply skill existed. The agent reads its list of skills once, at session start. So as far as that session knew, there was no skill for posting a reply. Asked to "post" something it had no tool for, it did what models do and improvised: searched the web for the article, extracted the page, and wrote a plausible summary instead. The fix was a fresh session. But it made me curious about what else the traces could tell me. Pulling latency per model from the same Langfuse data the default profile, all history : | Model | Calls | p50 | Slowest | |---|---|---|---| | qwen/qwen3.7-flash | 92 | 4.3 s | 180.5 s | | qwen3.5-64k local, 6 GB GPU | 91 | 125.2 s | 900.3 s | | deepseek-v4-flash | 65 | 6.7 s | | | mistral-nemo local | 24 | 106.6 s | 1,802.9 s | | hermes3-64k | 15 | 2.5 s | | On the tool side, terminal ran 135 times with a median of 0.6 s, and the slowest single tool call was a search files at 120.6 s. None of this is shocking. A local model on a 6 GB card is slow, and cheap cloud models are fast. But "slow" and "125 seconds median, about 30 times slower than the fastest cloud option" are different sentences, and only one of them helps decide which job runs where. Then I noticed the call counts. 91 calls for the local model. That seemed low for a system I lean on all day. Langfuse held 618 events across 79 traces in four weeks . That's about three traces a day. The agent keeps its own session database, so I counted from the other side. The last 14 days alone: Tracing was enabled on exactly one of ten profiles: the default chat one, with 91 calls. That's roughly 7% of the traffic. The untraced ones were the ones doing the work: | Profile | Model calls 14 days | Traced | |---|---|---| | coder | 522 | no | | sec | 390 | no | | itadmin | 172 | no | | writer | 120 | no | | default chat | 91 | yes | | reviewer | 55 | no | | scout | 29 | no | | others | small | no | The coding agent alone made more than five times as many calls as the only profile I was watching. The tracing plugin is enabled per profile, and it reads its Langfuse API keys through that profile's own secret scope. That's deliberate: profile B's traces shouldn't be shipped with profile A's keys. When I set Langfuse up, I enabled the plugin and added the keys on the default profile, and moved on. Every specialist profile then ran without tracing. No error. No warning in a log. The plugin fails open by design no keys means its hooks quietly do nothing , which is right for a plugin, and exactly why nobody noticed. A monitoring tool that's switched off looks identical to one that has nothing to report. There was a second, smaller gap. One-shot jobs launched in "safe mode" skip plugins entirely, and the reply drafter ran that way because it's quicker. That one is closed now: the drafter sends its own traces through the Langfuse SDK, so it no longer depends on the plugin. Langfuse v3 onwards stores traces in ClickHouse. Postgres holds users and projects, Redis handles the queue, MinIO stores blobs. So no, you can't drop ClickHouse to save resources; it is the trace store. ClickHouse was using 6.03 GiB of disk. The actual Langfuse trace data in it was 2.2 MiB . The rest was ClickHouse's own diagnostic logging: | System table | Size | Rows | |---|---|---| | system.trace log | 3.38 GiB | 174.6 million | | system.text log | 1.19 GiB | | | system.part log | 492 MiB | | | system.asynchronous metric log | 485 MiB | 949 million | | system.metric log | 461 MiB | | That's roughly 2,700 times more bytes about ClickHouse than about my agents. Those tables aren't written once and forgotten. They're written continuously. The node this runs on keeps its guests on a single 5400 rpm HDD, and it already has a roughly 15-minute IO storm after every reboot. The Langfuse container is the one I stop to relieve it. So the trace store I'd barely been using was grinding that disk all day, recording its own profiler samples. For each of the ten profiles: environment to the profile name. The third step is the one that matters for later. Traces carry no profile tag of their own, so without it every trace lands in one undifferentiated pile. With it, "cost and latency per agent role" is a filter in the UI rather than a spreadsheet exercise. Then I restarted the gateway and sent a one-word probe on two profiles. Traces arrived tagged legal on a cloud flash model and writer on the local qwen3.5-64k . That's the test I should have run the first time. ClickHouse lets you override server config with a file dropped into /etc/clickhouse-server/config.d/ . This one removes twelve system log tables, keeps query log with a 7-day TTL it's genuinely useful when a query is slow , and turns the server logger down to warnings: