Cheap LLM API Gateway Telemetry: One Key, Moderation Quality, and Cost A developer proposes a gateway-boundary telemetry schema for LLM-based moderation, arguing that teams should choose an API gateway by the evidence it lets them collect rather than by headline rate. The approach puts every backend behind one small TypeScript interface and emits normalized quality, latency, token, cache, batch, and region fields per classification attempt, so routes can be compared only when they used the same labeled report set. The winning route is defined as the fastest one that clears a per-class quality floor and satisfies data-location policy. TL;DR: Pick a gateway for moderation by the evidence it lets you collect, not by its cheapest-looking rate. Put every backend behind one small TypeScript interface, emit the same quality, latency, token, cache, batch, and region fields, then compare only runs that used the same labeled report set. The winning route is the fastest one that clears your quality floor and satisfies your data-location policy. A single key is convenient. Comparable telemetry is the real control surface. A gaming report classifier has an awkward job. It should move obvious spam, harassment, cheating, and benign reports into useful queues before human review, but a fast wrong label wastes moderator attention. A slow perfect label leaves the queue growing. This makes quality versus latency the primary decision, while cost is a constraint rather than the scoreboard. The before/after mental model is short. Before: send traffic through one credential, inspect a blended invoice, and debate which backend feels fast. After: attach an experiment ID to a frozen evaluation set, record normalized events at the gateway boundary, and decide from paired quality and latency distributions. That turns a vague comparison into an operational test your team can repeat. Start with one event per classification attempt. Keep the raw player report out of the event; an opaque case ID is enough to join against a restricted evaluation table. Record the requested backend separately from the backend actually used, because a fallback can otherwise make a healthy chart lie. For quality, retain the predicted class, the labeled class, and whether a human overrode the prediction. Overall accuracy alone is weak for moderation. A system can look good by getting a common benign class right while missing a rare, high-impact class. Compute per-class precision and recall from the joined evaluation data, then choose a quality floor that reflects the review policy. The gateway event supplies the join keys; the evaluation pipeline supplies the judgment. Latency needs at least two clocks. Time to first event shows how quickly a streamed response begins. Time to completion shows when the classification is actually usable. Server-Sent Events are a one-way server-to-client stream with named events, IDs, and reconnect behavior; MDN also documents the text/event-stream response type. Those mechanics make SSE useful for progress delivery, but the browser receiving bytes does not mean the moderation label is complete. Measure both boundaries. Token fields require similar care. Save estimated input tokens before dispatch and actual usage when the backend returns it. Do not silently substitute one for the other. Add the pricing snapshot ID used by your internal calculation instead of baking a volatile unit price into application code. A cache hit, a cache write, and an uncached request are different observations. So are online and batch work. This is the trap. A compact event can carry the evidence without becoming a transcript: type Region = "eu" | "us"; type Mode = "online" | "batch"; type Label = "spam" | "harassment" | "cheating" | "benign"; type Outcome = | { status: "completed"; predictedClass: Label } | { status: "failed"; errorCode: string }; interface ModerationRun { schemaVersion: 1; experimentId: string; caseId: string; requestedBackend: string; resolvedBackend: string; region: Region; mode: Mode; outcome: Outcome; labeledClass?: Label; startedAt: string; firstEventMs?: number; completedMs: number; estimatedInputTokens: number; actualInputTokens?: number; actualOutputTokens?: number; cacheStatus: "hit" | "write" | "miss" | "ineligible"; pricingSnapshotId: string; } The optional fields matter. A non-streamed or failed call may have no first event. A preflight estimate may exist before actual usage arrives. Missing is honest; zero is a measurement. Version 1 of this event deliberately separates four moderation labels, two regions, two execution modes, and four cache states. The 15 000 millisecond timeout in the runner is an example policy boundary, not a universal recommendation: set it from the queue's service objective, record it with the experiment, and keep it identical across candidates. Otherwise a route given 30 seconds will appear more reliable than one stopped after 10 seconds, even though the test created the difference. That specific configuration mismatch is easy to miss because both results can still produce tidy success-rate charts. Keep backend-specific payload conversion inside adapters. The experiment runner should see one request and one result shape, regardless of what sits behind the gateway. This is where one gateway credential earns its keep: application code authenticates once, while routing policy and upstream credentials remain outside the classifier. Do not let that convenience erase backend identity from logs. interface ClassifyInput { experimentId: string; caseId: string; report: string; backend: string; region: "eu" | "us"; mode: "online" | "batch"; } interface ClassifyResult { label: Label; resolvedBackend: string; actualInputTokens?: number; actualOutputTokens?: number; cacheStatus: "hit" | "write" | "miss" | "ineligible"; } interface GatewayAdapter { classify input: ClassifyInput, signal: AbortSignal : Promise