cd /news/ai-agents/how-to-combine-production-failure-al… · home › topics › ai-agents › article
[ARTICLE · art-145556] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

How to Combine Production Failure Alerts, Logs, Metrics, and Traces — 2026

A developer outlines an architecture for alerting on AI agent failures by treating each agent run as a single auditable operation rather than paging on individual error logs or spans. The design keeps metrics low-cardinality, stores request_id, trace_id, region, model usage and attempt data in structured logs, and uses a single idempotent polling worker to evaluate closed time windows for EU and US traffic, writing a durable decision record before notifying anyone. The stated goal is exactly-once alert effects despite at-least-once polling and delivery, with thresholds set by service owners from their error budget and traffic shape.

by read8 min views10 publishedOct 5, 2026

Instrument each AI agent run as one auditable operation, page on a sustained failure ratio or latency breach, and use request and trace identifiers only to investigate the page. The deciding constraint is signal quality: a media SaaS cannot treat every failed tool call, slow model turn, or polling retry as a distinct incident without training its on-call engineers to ignore alerts.

Noise wins otherwise.

TL;DR: keep metrics low-cardinality, put request_id, trace_id, region, model usage, and attempt data in structured logs, and let a single idempotent polling worker evaluate completed time windows for EU and US traffic. The worker should write a durable decision record before notifying anyone. This gives retries exactly-once alert effects, even though polling and delivery remain at-least-once operations.

The architecture decision begins with invariants, not dashboards. A media workflow may retrieve a brief, call a model, invoke a rights-checking tool, and revise a draft several times; the unit that matters is the whole run, while an individual step remains necessary evidence. An alert therefore represents a population-level condition, not a single span.

Four invariants govern the design:

region, operation, and outcome; request and trace identifiers stay out of labels because each new label combination creates another time series. The failure boundaries follow from those rules. A Next.js entry point or Node.js API process records run outcomes but never decides whether to page. The metrics system aggregates health but does not carry forensic identity. Logs carry identity and usage detail, while traces explain the critical path. Finally, the polling worker owns policy evaluation and notification deduplication. These boundaries matter when a regional ingest path is late: an incomplete window must be marked incomplete, rather than quietly interpreted as a healthy zero. They also prevent a presentation-tier deployment from becoming the owner of a durable operational decision merely because it received the original request.

This is the same discipline used for a ledger: preserve the event, make the derived decision reproducible, and separate an attempted side effect from its committed result. Exactly once is an outcome to engineer, not a transport promise.

The useful comparison is not “logs versus metrics.” It is where alert truth lives and how much noise each arrangement admits.

Option Alert trigger Investigation context Noise and correctness boundary
Alert on every error log Individual event Already attached Fast, but retries and expected tool failures can multiply pages
Evaluate traces Sampled or complete run paths Rich causal sequence Strong diagnosis, but alert completeness depends on trace retention and sampling choices
Poll bounded metric windows Aggregated outcomes Joined afterward through stored identifiers Stable paging signal, but requires late-data handling and durable deduplication

The decision here is the third option. Poll a closed window, compare failed / completed and latency against policy, then attach a small set of exemplar request IDs for investigation. A threshold should be configured by the service owner from its error budget and traffic shape; inventing a universal percentage would create false precision. Require a minimum completed-run count as well, because 1 failure among 2 runs says something different from 500 failures among 1,000. I choose this trade-off because the page should describe user-visible outcomes, while the supporting evidence should still preserve each attempt.

I initially prefer event-driven alert evaluation for its immediacy. Under this design, however, individual agent steps are the wrong accounting grain: retries, fallbacks, and multi-turn loops turn one user-visible outcome into several error-shaped events. The rejected option remains valid for integrity violations that must never be aggregated away, such as a malformed audit record or an impossible state transition. Those should emit a separate, direct operational signal.

No alert path should depend on joining every log line during an incident. The decision record needs enough evidence to stand alone, while the IDs provide a route into deeper context.

This polling design has a real limitation: it is not suitable when a single integrity violation requires an immediate stop, and its detection delay is at least the chosen window plus evaluation lag. In that case, use a direct integrity alarm alongside the aggregate policy, not in place of the durable reconciliation path. It is also a poor fit for traffic so sparse that no window reaches a meaningful denominator; a synthetic transaction or an explicit per-run control is the more honest alternative.

The worker below deliberately uses interfaces rather than a specific telemetry backend. It evaluates the last closed 5-minute window per region, refuses incomplete data, inserts a deterministic decision key, and sends only when that insert wins. In production, InsertDecision and an outbox write belong in one database transaction; a dispatcher can then retry delivery without losing the audit trail.

package alerting

import (
    "context"
    "fmt"
    "time"
)

type Window struct {
    Start, End time.Time
}

type Aggregate struct {
    Completed   uint64
    Failed      uint64
    P95Latency  time.Duration
    Complete    bool
    EvidenceIDs []string
}

type Decision struct {
    Key, PolicyVersion, Region, Outcome string
    Window                              Window
    Completed, Failed                   uint64
    P95Latency                          time.Duration
    EvidenceIDs                         []string
}

type Store interface {
    Aggregate(context.Context, string, Window) (Aggregate, error)
    InsertDecision(context.Context, Decision) (inserted bool, err error)
    EnqueueNotification(context.Context, string) error
}

type Policy struct {
    Version        string
    MinimumRuns    uint64
    FailureRatio   float64
    LatencyCeiling time.Duration
}

func Evaluate(ctx context.Context, store Store, region string, w Window, p Policy) error {
    a, err := store.Aggregate(ctx, region, w)
    if err != nil {
        return fmt.Errorf("aggregate %s: %w", region, err)
    }
    if !a.Complete || a.Completed < p.MinimumRuns {
        return nil
    }

    failureRatio := float64(a.Failed) / float64(a.Completed)
    breached := failureRatio >= p.FailureRatio || a.P95Latency >= p.LatencyCeiling
    if !breached {
        return nil
    }

    key := fmt.Sprintf("%s/%s/%d", p.Version, region, w.Start.UTC().Unix())
    d := Decision{
        Key: key, PolicyVersion: p.Version, Region: region, Outcome: "alert",
        Window: w, Completed: a.Completed, Failed: a.Failed,
        P95Latency: a.P95Latency, EvidenceIDs: a.EvidenceIDs,
    }
    inserted, err := store.InsertDecision(ctx, d)
    if err != nil {
        return fmt.Errorf("record decision %s: %w", key, err)
    }
    if !inserted {
        return nil
    }
    return store.EnqueueNotification(ctx, key)
}

One detail is intentionally strict: the denominator is Completed, not “requests observed.” A run that is still awaiting a tool result has no terminal outcome yet. Counting it as success suppresses failures; counting it as failure pages on ordinary latency. The aggregator should classify terminal outcomes once and retain the source event identifier so replay cannot increment the same run twice.

Replay is normal.

Keep the deterministic key free of volatile evidence IDs. Policy version, region, and window start define the decision identity; the database enforces uniqueness on that key. This makes a worker crash after insertion boring. On restart, the same evaluation loses the insert race and does not enqueue another notification.

There is a sharper transaction issue in the sample interface: if inserting the decision succeeds and enqueueing fails, a naive retry sees the existing decision and returns without a message. The implementation contract must therefore make decision insertion and outbox enqueue atomic, or have the retry repair a decision whose outbox row is absent. The compact interface keeps storage mechanics out of the example, but the invariant is non-negotiable.

Metrics answer “is the population unhealthy?” Logs and traces answer “which run, and why?” Preserve that division. A counter might use operation=agent_run, region=eu|us, and outcome=success|failure; a latency histogram can use the same bounded dimensions. Never add request_id, trace_id, user, prompt, document title, or model response as metric labels. Prometheus explicitly warns that each label set creates a new time series and advises caution with high-cardinality values.

Cost attribution needs an auditable input rather than a mutable price embedded in an alert. Record input and output usage units, model class, currency context if applicable, and the effective rate-card version in the run ledger. Then compute cost as a derived field under that version. The alert policy can detect a change in usage per completed article without claiming that a particular amount is universally expensive.

For EU and US processing, keep region bounded and explicit at run creation. Do not infer it later from an IP address or whichever collector received the log. The alert worker evaluates each region separately before any global rollup, so a large healthy region cannot mask a smaller unhealthy one. It should also record source freshness in the decision. Late telemetry is a data-quality state.

Short logs are enough. A terminal record can contain run_id, request_id, trace_id, region, outcome, attempt, usage units, timestamps, and an error class; prompts and generated copy do not belong in routine operational evidence because they increase exposure without improving the initial page. Retention and access rules should follow the organization’s actual compliance obligations. No universal retention period can be inferred from this architecture.

Test the evaluator with fixed windows and a fake store before deployment. The cases with the highest value are duplicate polling, a crash after the decision commit, late aggregates, zero completions, the exact threshold boundary, and simultaneous workers evaluating the same key. Property tests can generate duplicate and reordered run events, then assert that each run contributes once and each breached policy window produces at most one decision.

Deployment should begin in record-only mode. Compare decisions with terminal run data, inspect the false-positive classes, and version any policy change; do not silently rewrite an old decision under new thresholds. Once notifications are enabled, reconcile three counts on a schedule: breached windows, durable decisions, and outbox deliveries. Any mismatch is itself an operational defect, because the paging system has become unable to prove what it did.

Keep two service-level views. The first measures the media workflow: terminal failure ratio, completion latency, and usage per completed run. The second measures the observer: aggregate freshness, worker evaluation lag, decision-write failures, and undelivered outbox rows. Avoid making the second path recursively page through itself without a separate failure boundary.

The final rule is plain. Page from bounded, completed-window evidence; diagnose with request and trace identity; and preserve every alert decision as a versioned audit record. That arrangement gives the on-call engineer fewer, better claims to investigate while keeping enough evidence to challenge each claim later.

── more in #ai-agents 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-to-combine-produ…] indexed:0 read:8min 2026-10-05 · —