A Guide to 4 NestJS Error Tracking Boundaries Beyond HTTP Interceptors and Filters A developer outlined an approach to NestJS production error tracking that uses one normalized failure envelope across four execution boundaries — HTTP, cron, queue worker, and outbound AI-agent step — rather than relying solely on HTTP exception filters and interceptors. The design assigns local context to each boundary while a shared reporter handles normalization, deduplication keys, and redaction, with latency and cost tracked as measurements separate from exceptions. The stated invariant is that one failed unit of work produces exactly one canonical failure event, enforced via a capture marker or stable event key. For NestJS production error tracking, use one normalized failure envelope at every execution boundary, then measure latency and estimated cost separately from exceptions. An HTTP filter and interceptor cover only the request path; a media AI agent loop also needs explicit capture around scheduled work, queue processing, and outbound model steps without turning every slow call, retry, or expected rejection into the same noisy alert. The deciding constraint is signal quality. An HTTP exception filter can see request failures, but the production system also runs scheduled ingestion, queue-based transcoding, and model calls that may outlive the request. The architecture decision is therefore to capture at four boundaries: HTTP, cron, queue worker, and outbound agent step. Each boundary owns local context; one shared reporter owns normalization, deduplication keys, and redaction. This is deliberately narrower than "log everything." Latency and cost are measurements. Exceptions are failed outcomes. They can share correlation fields without sharing alert policy. The invariant is simple: one failed unit of work produces one canonical failure event, even if the error crosses several layers. Give the event a stable operation ID, execution kind, attempt number, media asset ID, agent-step name, duration, cost estimate, and a normalized exception class. Do not attach raw prompts, generated scripts, access tokens, or full media URLs; high-cardinality payloads make search expensive and can leak material that never belonged in telemetry. The failure boundary matters more than the class name. A controller can translate a domain rejection into an HTTP response, a cron runner can mark a scheduled scan failed, a queue consumer can decide whether an attempt is retryable, and an outbound agent wrapper can record provider latency. If all four call the reporter independently for the same thrown value, the dashboard counts four incidents. If none claims ownership, the worker can fail silently while the HTTP graph stays green. Use a capture marker or stable event key to enforce exactly one report per attempt. This is an application-level deduplication rule, not a claim that the transport delivers exactly once. Telemetry delivery can fail, processes can terminate between capture and flush, and queues can redeliver work; durable business state must never depend on the error tracker accepting an event. Failure modes should be named because vague dashboards invite vague responses: 404 for a missing media asset pages the same team as a failed queue execution. Keep it boring. Treat framework hooks as adapters around a shared capture contract. The filter is the HTTP adapter; an interceptor measures request completion and can attach timing, but it should not become the universal exception owner. Cron callbacks and queue processors need explicit boundary wrappers because they execute outside the HTTP lifecycle. The outbound AI-agent step needs its own measurement wrapper so a successful but slow or costly call remains a metric rather than a fabricated error. | Capture point | Context it owns | Useful signal | Main limitation | |---|---|---|---| | HTTP exception filter | route, method, status, request correlation | uncaught request failure | cannot observe detached cron or worker execution | | HTTP interceptor | end-to-end request duration and outcome | latency distribution | request completion can precede background failure | | cron boundary | schedule name, run ID, scheduled time | failed media scan or aggregation run | no request context exists unless propagated explicitly | | queue worker boundary | job ID, attempt, queue operation | terminal or retryable processing failure | redelivery can inflate counts without a stable key | | agent-step wrapper | model operation, duration, usage-derived estimate | latency and estimated cost by step | an estimate must be labeled as such and reconciled elsewhere | The table has five rows because the agent-step wrapper is a measurement boundary nested inside one of the four execution contexts, not a fifth top-level runtime. That distinction prevents a common accounting error: summing queue duration and model duration as if they were independent units of end-to-end latency. Store timestamps for spans, then derive critical-path latency from their relationships. Don't add durations blindly. Alert only on outcomes that need action. For HTTP, an expected client rejection can remain searchable without paging. For a worker, record each attempt but alert on the terminal outcome or on a retry-rate rule. A 429 from a downstream dependency may justify retry metadata and a latency sample; it does not prove that the media job is permanently lost. The precise terminal rule depends on queue policy, so the dashboard must expose attempt and terminal-state fields rather than infer them from exception text. The following Python reference is intentionally framework-neutral. In NestJS, the filter, interceptor, scheduler method, and worker processor are thin adapters that assemble this envelope and call the same contract. Keeping the contract independent of a vendor SDK makes it testable with an in-memory sink and prevents transport details from leaking into domain code. python from dataclasses import asdict, dataclass from hashlib import sha256 from time import monotonic from typing import Callable, Literal, Protocol, TypeVar ExecutionKind = Literal "http", "cron", "queue", "agent step" T = TypeVar "T" @dataclass frozen=True class FailureEvent: operation id: str execution kind: ExecutionKind operation: str attempt: int duration ms: int exception class: str message: str terminal: bool @property def event key self - str: raw = f"{self.operation id}:{self.execution kind}:{self.operation}:{self.attempt}" return sha256 raw.encode "utf-8" .hexdigest class EventSink Protocol : def emit self, event: dict - None: ... def run boundary , sink: EventSink, operation id: str, execution kind: ExecutionKind, operation: str, attempt: int, terminal: bool, work: Callable , T , - T: started = monotonic try: return work except Exception as error: event = FailureEvent operation id=operation id, execution kind=execution kind, operation=operation, attempt=attempt, duration ms=round monotonic - started 1000 , exception class=type error . name , message=str error :500 , terminal=terminal, payload = asdict event payload "event key" = event.event key sink.emit payload raise The reporter rethrows because capture must not change business control flow. A NestJS HTTP filter still maps the exception to the intended response; a worker still follows its retry policy; a cron runner still records its run outcome. The sink should accept a structured dictionary, redact before serialization, and avoid storing arbitrary exception objects whose fields vary across libraries. Test this as a matrix, not as a happy-path snapshot. For each of the four contexts, verify success emits no failure, a thrown exception emits one event, a repeated attempt gets a distinct key, the same exception crossing an inner adapter is not captured twice, and sensitive fields never appear. Then test the latency and estimated-cost measurements independently. A failure counter cannot tell you that an agent loop is healthy but too slow, while a cost total cannot tell you which step failed before producing usable media. Take a hypothetical media asset with operation ID asset-42 . Its first queue attempt enters the agent step, receives a retryable downstream response, and exits without producing the final transcript; the second attempt completes. The event stream should show one queue operation with two attempts, one captured exception keyed to attempt one, and separate duration samples for each agent call. It should not show two unrelated media failures, duplicate the first exception at both the agent and worker adapters, or add both agent durations to the second attempt's critical path. This small fixture forces the test to answer the questions dashboards usually obscure: what was attempted, which boundary owned the failure, whether the work eventually completed, and which measurements belong in latency and cost analysis rather than an alert count. Rollout also deserves an explicit boundary. A release toggle can shadow-write normalized events to a test sink while the established path remains authoritative; compare event counts and grouping keys, then change ownership one context at a time. Feature toggles add carrying cost and should have an owner and removal condition, particularly when they alter operational behavior rather than user-facing features. A global interceptor is attractive because it centralizes code and naturally measures HTTP latency. The catch is that its lifecycle is the request lifecycle. It is not suitable as the sole production capture point when cron jobs and queue workers continue without an HTTP request, and pretending otherwise creates the most dangerous dashboard: one that looks complete. Reject it as the universal owner, not as a tool. Stick with a global interceptor when the application is strictly request-response, has no detached work, and only needs HTTP outcome timing plus a consistent correlation field. Likewise, a small internal service with a single scheduler may be adequately served by a local try / except boundary; introducing a generalized envelope there can cost more cognitive load than it removes. Your mileage may vary once retries, multiple teams, or cross-process correlation enter the design. I'm not sure any fixed alert threshold can be correct for both an interactive editing request and an overnight media indexing run. The evidence needed is workload-specific: latency objectives, queue delay, retry policy, and the operational cost of a missed failure. Start with separate views by execution kind and terminal state, then tune alerts from observed distributions without rewriting the capture contract. The decision record is therefore compact: adapters own context, the reporter owns shape, the sink owns delivery, and business state owns truth. That split gives a media team enough detail to compare agent-loop latency and estimated cost while keeping error tracking focused on failures that someone can act on.