cd /news/developer-tools/nobody-alerts-on-silence-wiring-sent… · home topics developer-tools article
[ARTICLE · art-91909] src=dev.to ↗ pub= topic=developer-tools verified=true sentiment=· neutral

Nobody Alerts on Silence: Wiring Sentry Into an LLM Pipeline

A developer wired Sentry error monitoring into the LLM pipeline of TextStack, an open-source .NET reader, after discovering that a user's PDF fell onto a CPU-only Ollama container, causing 390% CPU usage for an hour with zero alerts. The fix includes route resolution that records intent, alert throttling, and a per-book re-check that caught a hole in the initial fix. The first 24 hours in production caught HTTP 429 errors on the paid translation and explanation endpoints.

read6 min views1 publishedAug 11, 2026

This is a submission for DEV's Summer Bug Smash: Clear the Lineup powered by Sentry.

🔨 #bugsmash, week by week:

[a 390% CPU hour nobody noticed],[a state machine with no exit],[a backup that leaked 156 GB]. Week four is the finale: I wired monitoring into the pipeline that produced all three — and its best catch was itself.

TextStack is an open-source reader for technical books, built in .NET: an ASP.NET Core API, a background Worker, PostgreSQL + pgvector, React on top. The LLM pipeline does translation, word explanations, "Ask this book" RAG, and three production agents (Enrichment, Librarian, Tutor), routed between a local Ollama and OpenAI by a config-driven router. The code is public: github.com/mrviduus/textstack.

Three weeks ago a user's PDF fell through my LLM router onto a CPU-only Ollama container instead of GPT-4.1, and my CPU sat at 390% for an hour. Zero exceptions. Zero error logs. Zero alerts. And when I went to see what my existing observability had recorded, the answer was nothing at all: the OTLP exporter pointed at an Aspire dashboard container that is profile-gated and doesn't run in production. Every span my services had ever produced in prod had been fired into a closed socket.

Observability you never read is indistinguishable from observability you never installed.

The one-line config fix was submission #1. This submission is the fix for the class of bug — a system that has no way to make a sound when it does the wrong thing successfully:

Four PRs, all merged to main; 1,363 unit tests, full CI green:

23505

on reading-progress upserts, real users losing their place in booksThe router now says why. Route resolution was a ??

chain that produced a string — identical whether an operator deliberately routed a task or it fell off the end onto the default. That chain doesn't just fail to record intent; it destroys it. So it returns two things now:

private RouteDecision ResolveRoute(string? featureTag)
{
    var matched = RegistryKey(featureTag) ?? ConfigRouteKey(featureTag);
    return matched is not null
        ? new RouteDecision(matched, RouteReason.RouteMatched)
        : new RouteDecision(config["Ai:DefaultProvider"] ?? "openai",
                            RouteReason.DefaultFallback);
}

Every LLM call tags its span with ai.task

, ai.provider.resolved

, and ai.provider.reason

= route_matched | default_fallback

. "Which model answered this, and did anyone choose it on purpose?" is now a trace query instead of a CPU graph.

Alert arithmetic. pdf.parse

resolves a route once per page with parallelism six — my first version would have turned the original incident into 106 identical Sentry events. Every alarm goes through a throttle keyed on (task, provider, reason)

: first hit fires immediately, then one event per hour per distinct problem. The unit test literally counts to 106 and asserts one claim.

No silent fallback, ever — in either direction. When the breaker finds Ollama dead, tasks are skipped and stay queued; nothing auto-switches to a paid provider, because that converts an outage into unbounded spend. Provider choice stays 100% config-driven.

And the first live run found a hole in my own fix. The startup probe opens the circuit on a one-minute backoff; the backfill worker wakes after a two-minute start delay — by then the circuit is legitimately half-open, and my single up-front gate waved the whole batch through. A per-book re-check turned 38 calls into one:

Metadata backfill: enriching 38 user books
Metadata backfill: aborting after 0 enriched / 1 failed — provider 'ollama'
  is unavailable; the remaining candidates stay queued

Tests check what you imagined; a live run checks what's there.

Error Monitoring — what the first 24 hours in production caught:

HTTP 429 (insufficient_quota)

on /translate

and /explain

— the entire paid surface had been failing for readers for twelve hours. No version of my logs would have surfaced that before a user complained.PUT /me/progress

threw 23505: duplicate key value violates unique constraint

ten times in four hours: a textbook read-then-insert race (session heartbeat + sendBeacon

on unload + second device), milliseconds wide, invisible in tests. Fixed in Custom tags (ai.task

, ai.provider

, ai.failure

, agent.name

, agent.outcome

) go through an allowlist scrubber — every tag not explicitly blessed dies at the edge, so a future SetTag("prompt", userText)

can never leak. A Sentry issue answers "which feature, on which model, is broken?" without opening a trace.

Tracing covers agent runs and RAG indexing at 100% sampling (they're the reason I installed this), HTTP at 20%, health checks at 0%. I rejected the deprecated OTel bridge specifically because spans leaving through the OpenTelemetry SDK bypass BeforeSend

— my OTel pipeline carries raw client IPs and full SQL text that must never leave the box. A tiny TraceScope

dual-writes an Activity

and a Sentry span instead, so everything Sentry receives passes my scrubber.

Breadcrumbs caught my scrubber lying — twice. A live event's breadcrumb trail contained SQL: EF Core interpolates the query into the breadcrumb message, not the structured data

bag my scrubber nulled (and my unit tests were green the whole time, asserting exactly the wrong thing). Fixed by dropping EF command breadcrumbs outright — then production found the same leak in a second channel: EF logs a failed command at Error

level and Sentry's ILogger

integration promotes it to an event, SQL in the message again. A scrubber written against one egress path will be bypassed by the next one. Both doors are closed in #446, and dropping loses no signal — the exception middleware already reports the same failure with the SQLSTATE and constraint name, no SQL.

Release + environment tags as forensics. The most interesting issue of the first day showed a dead Ollama starving a metadata pipeline: thirty events, tagged environment: Production

. I read it as an outage and started writing the fix. It was my laptop — a dev .env

with ASPNETCORE_ENVIRONMENT=Production

plus the production DSN I'd pasted in to verify the integration. What broke the spell was Sentry's own metadata: linux-arm64

runtime on an x86_64 prod, and a release

tag pointing at a commit that had never been deployed. An environment tag is a claim a process makes about itself, not a fact. Now SENTRY_RELEASE

comes from the GIT_SHA

build arg — every CI-built image has one, no dotnet run

ever does — and a Production claim without a release gets renamed production-unverified

(#448).

And the meta-lesson that justified the whole exercise: I verified the integration by sending real events at the real DSN and reading the captured payloads in the UI — that's how the leaks, the inferred-geo surprise, and the middleware capture path all surfaced. A monitoring system whose first act is to indict itself is one you can start trusting.

What's the most embarrassing thing your monitoring has ever caught — and was it in the code, or in you?

I build TextStack, an open-source reader for technical books, in .NET. The full write-up lives on my blog. github.com/mrviduus/textstack

── more in #developer-tools 4 stories · sorted by recency
── more on @sentry 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/nobody-alerts-on-sil…] indexed:0 read:6min 2026-08-11 ·