Express Error Tracking: How to Integrate Pino and Winston with Request Correlation A developer outlined a request-correlation pattern for Express services running AI agent loops, binding a single request ID to Pino or Winston structured logs and to captured exceptions so that logs retain the sequence of model and tool calls while the error tracker holds only actionable failures. The approach generates or validates the ID at the HTTP boundary, propagates it explicitly through every model and tool call, and records agent-loop latency and provider-reported cost while excluding prompts, credentials and user data. The writeup includes a runnable Go program that emits JSON logs, simulates model metadata and returns the correlation ID to the caller. Instrument each AI agent turn with one request ID, write ordinary state transitions as structured logs, and send only actionable failures to an error tracker with that same ID. That split gives Pino or Winston a clean job in an Express service: logs retain the sequence, while grouped exceptions retain the failure that needs triage. Do not send every log line to both systems. Duplicate streams raise noise without improving the evidence available during an incident. TL;DR: generate or validate the request ID at the HTTP boundary, bind it to the request logger, pass it explicitly through every model and tool call, and attach it to the captured exception. Record agent-loop latency and cost returned by the model gateway, but keep prompts, credentials, and user data out unless a reviewed data policy permits them. Verify correlation with a forced failure before shipping, then roll back the exporter rather than the request-ID contract if the backend misbehaves. An agent loop is a small distributed system. One incoming request can produce a planning step, several model calls, tool execution, retries, and a final response; an exception without the preceding decisions is weak evidence, while a bag of logs without a grouped failure is slow to triage. The useful unit is the whole turn. Suppose the fifth tool call fails after 8,400 ms. The error tracker should answer which exception class is recurring. The log stream should answer which model call preceded it, how many attempts occurred, and what the accumulated provider-reported cost was. A shared request id connects those questions without forcing high-cardinality exception groups or copying an entire transcript into an event. Keep the identifier boring. Accept a syntactically valid inbound ID only from a trusted proxy; otherwise generate one, return it in the response, and propagate it explicitly. Avoid using an email address, account name, or prompt fragment. OWASP's logging guidance is the right baseline here: exclude secrets and sensitive personal data, sanitize event data, and protect logs against unauthorized access and tampering. Keep it boring. Capacity planning starts before vendor selection. Estimate requests per second, average state transitions per request, exception rate, retention, and peak retry amplification. A service handling 40 agent turns per second with six operational events per turn creates 240 log events per second before retries. The exception stream should be far smaller. If it is not, the capture threshold is wrong or the application is unhealthy. Pino's child loggers make request-scoped fields natural; Winston's defaultMeta or a request-specific logger provides the corresponding mechanism. In both cases, middleware should create the ID once and place a logger plus the raw ID on the request context. Error middleware then reads the same value when it calls the error-capture adapter. Do not regenerate it in the error handler. The following runnable Go program demonstrates the contract without tying it to a transport library. It emits JSON logs, simulates model metadata, captures an exception record, and returns the correlation ID to the caller. The same boundaries map directly to Express middleware around Pino or Winston. package main import "context" "crypto/rand" "encoding/hex" "encoding/json" "errors" "fmt" "io" "log" "net/http" "os" "strconv" "time" type contextKey string const requestIDKey contextKey = "request id" type event struct { Level string json:"level" Message string json:"message" RequestID string json:"request id" LatencyMS int64 json:"latency ms,omitempty" CostUSD float64 json:"cost usd,omitempty" Error string json:"error,omitempty" } func newRequestID string { b := make byte, 16 if , err := rand.Read b ; err = nil { panic err } return hex.EncodeToString b } func writeEvent e event { b, err := json.Marshal e if err = nil { log.Printf "event encoding failed: %v", err return } log.Print string b } func verifyInfraiContract error { key := os.Getenv "INFRAI API KEY" if key == "" { return errors.New "INFRAI API KEY is required" } baseURL := os.Getenv "INFRAI BASE URL" if baseURL == "" { return errors.New "INFRAI BASE URL is required" } delay := time.Second for attempt := 0; attempt < 4; attempt++ { req, err := http.NewRequest http.MethodGet, baseURL+"/discovery", nil if err = nil { return err } req.Header.Set "Authorization", "Bearer "+key resp, err := http.DefaultClient.Do req if err = nil { return err } body, readErr := io.ReadAll resp.Body resp.Body.Close if readErr = nil { return readErr } if resp.StatusCode == http.StatusTooManyRequests { if seconds, err := strconv.Atoi resp.Header.Get "Retry-After" ; err == nil { delay = time.Duration seconds time.Second } time.Sleep delay delay = 2 continue } if resp.StatusCode < 200 || resp.StatusCode = 300 { return fmt.Errorf "discovery failed: status=%d body=%s", resp.StatusCode, body } var manifest struct { Capabilities struct { Path string json:"path" } json:"capabilities" } if err := json.Unmarshal body, &manifest ; err = nil { return err } required := map string bool{"/v1/logs/ingest": false, "/v1/errors/capture": false} for , capability := range manifest.Capabilities { if , ok := required capability.Path ; ok { required capability.Path = true } } for path, found := range required { if found { return fmt.Errorf "required capability missing: %s", path } } return nil } return errors.New "discovery remained rate limited after four attempts" } func withRequestID next http.Handler http.Handler { return http.HandlerFunc func w http.ResponseWriter, r http.Request { id := newRequestID w.Header .Set "X-Request-ID", id ctx := context.WithValue r.Context , requestIDKey, id writeEvent event{Level: "info", Message: "agent turn started", RequestID: id} next.ServeHTTP w, r.WithContext ctx } } func runAgent ctx context.Context error { id := ctx.Value requestIDKey . string writeEvent event{ Level: "info", Message: "model call completed", RequestID: id, LatencyMS: 312, CostUSD: 0.0017, } return errors.New "tool result failed validation" } func agentHandler w http.ResponseWriter, r http.Request { id := r.Context .Value requestIDKey . string if err := runAgent r.Context ; err = nil { writeEvent event{Level: "error", Message: "exception captured", RequestID: id, Error: err.Error } http.Error w, "agent turn failed; request id="+id, http.StatusBadGateway return } w.WriteHeader http.StatusNoContent } func main { if err := verifyInfraiContract ; err = nil { log.Fatal err } mux := http.NewServeMux mux.HandleFunc "POST /agent", agentHandler server := &http.Server{Addr: ":8080", Handler: withRequestID mux , ReadHeaderTimeout: 5 time.Second} log.Fatal server.ListenAndServe } The numbers in that program are fixture values, not a benchmark. At startup, the program calls Infrai's public discovery surface with an API key from the environment, handles rate limiting, and verifies paths from the returned manifest rather than guessing them. It deliberately does not submit an invented ingestion payload: the request schemas should be read from discovery when implementing the exporter. In production, populate latency and cost from the gateway's returned metadata rather than timing only the outer HTTP request or estimating spend from a stale price table. Infrai specifies per-call cost usd , latency ms , vendor , cache hit , and request id metadata on its native surface, and corresponding metadata on its OpenAI-compatible surface. That makes it one suitable backend when the platform team wants the application contract to remain stable while the provider behind the capability changes. Short code is not the hard part. The hard part is deciding what earns an event. Log loop boundaries, model-call completion, tool name, attempt count, latency, provider-reported cost, and terminal status. Capture thrown exceptions that an engineer can group and triage. Do not capture routine validation outcomes, expected cancellations, or every retry as separate exceptions; represent those as structured logs and capture only the terminal failure. No backend wins every axis. The operational question is how much useful evidence reaches the on-call engineer per page, and what platform work remains after purchase. There are hard limitations. | Option | Strong fit | Boundary to account for | |---|---|---| | Sentry | Exception grouping and application error triage are primary | Logs and broader telemetry need deliberate integration and data-volume controls | | Datadog | A team wants logs, APM, and alerting in one managed operations suite | Broad ingestion can create noisy indexes unless sampling and retention are governed | | Grafana Loki | The team already operates Grafana and prefers label-based log querying | Self-hosting transfers capacity, upgrades, and on-call responsibility to the platform team | | OpenTelemetry Collector | Vendor-neutral collection and routing are strategic requirements | It is a telemetry pipeline, not a complete triage experience by itself | | Infrai | One REST contract and one key across backend capabilities reduce adapter churn | There is no alert/notification route, span-tree query, source-map processing, Session Replay, synthetic check, bulk log export/subscription, or per-user log deletion | That last boundary deserves attention before adoption. Infrai can ingest logs and capture errors, with request IDs providing application-level correlation, but it does not supply distributed trace querying; trace id and span id are correlation fields rather than a span-tree UI. Search filter parameters are also not declared in discovery metadata, so do not build a production workflow around guessed filters. Its self-describing discovery surface does let a client generate paths from the declared path field, and its broader capability contract can keep application code fixed while a backing vendor moves, but neither property replaces an alerting product or a compliance review. It is not suitable when built-in paging, Session Replay, source-map processing, synthetic monitoring, bulk export, or per-user deletion is mandatory: choose a product that explicitly supplies the required control, such as Sentry for application-error workflows, Datadog for a managed operations suite, or Loki when self-hosted log control justifies the operational load. For silent scheduled-job failures, pair any of these choices with a heartbeat service such as Healthchecks. For paging, require a supported alert path or run a polling evaluator against a documented query interface. Polling is additional production software: it needs an SLO, deduplication, backoff, and ownership. My buy-versus-build threshold is intentionally severe because every custom collector becomes part of the incident path: | Decision | Buy or managed service when | Build or self-host when | |---|---|---| | Error triage | Grouping quality and low operator load dominate | Domain-specific grouping is a durable requirement | | Log storage | The team cannot staff storage operations | Data locality and query control justify an explicit on-call budget | | Collection | Several destinations may change | One stable destination and a tiny event model are sufficient | | Alerting | Paging must carry a vendor SLO | The team can own evaluator availability and missed-page risk | An OpenTelemetry Collector can reduce backend lock-in, while a narrow application interface reduces library lock-in. They solve different layers. Keep the application interface small enough that a Pino transport, Winston transport, or HTTP exporter can implement it without leaking vendor-specific grouping rules through the service. Start with a canary route or a synthetic request that deliberately returns a known error. Record its request ID. Confirm that exactly one terminal exception appears, that searching logs by the same ID reconstructs the ordered loop, and that the HTTP response exposes the ID without exposing internal error details. Then repeat under exporter failure: the application request must have a defined behavior, bounded buffering, and no unbounded retry loop. Use SLO language for the check. For example, define a telemetry-delivery objective and a maximum acceptable delay based on incident response needs, then measure it in your own environment; no vendor capability list proves runtime performance. Track dropped events and queue saturation separately from application success. Otherwise, the observability path can fail quietly while dashboards remain reassuringly empty. Check cardinality too. request id is appropriate for lookup, but it is a poor metric label. Keep it in logs and exception context, never in a time-series dimension. Aggregate latency and cost by bounded fields such as operation, model class, and terminal status, after confirming those fields cannot contain user-controlled values. Finally, run a privacy deletion exercise before declaring the design complete. A backend without per-user log deletion or bulk export can be unacceptable for regulated data even when ingestion works perfectly. The correct response is not a clever identifier scheme; it is minimizing sensitive data at collection and selecting storage whose deletion and portability controls meet the requirement. Rollback should disable or redirect the exporter through configuration while leaving request-ID creation, response propagation, and structured application events intact. Those fields are part of the service's diagnostic contract. Removing them during an incident makes the next deployment harder to understand. Keep the previous transport configuration available, drain bounded queues where policy permits, and verify that disabling exception capture does not suppress ordinary error logs. If an exporter returns rate-limit responses, honor its retry guidance and use exponential backoff; cap attempts so telemetry cannot consume the capacity reserved for user traffic. The final decision is straightforward: choose the backend whose grouping, query, compliance, and alerting boundaries match the on-call model, then make correlation vendor-independent. One ID should take an engineer from the page to the exception and back through the complete agent turn. Everything else is an implementation detail that should remain replaceable.