Logging Runbook for Small SaaS — Docker Cron Search With Safe Rollbacks A developer published a Go-based logging runbook for small SaaS teams that centralizes web, worker, and cron output into a single JSON event schema with fields like service, job_name, request_id, trace_id, latency_ms, and cost_usd, aimed at AI agent loops that cross process boundaries. The approach uses an allowlisted event struct, idempotency keys to prevent duplicate writes on retry, Retry-After handling for 429 responses, and exponential backoff, while pairing scheduled jobs with external heartbeat monitors since a cron process that never starts emits no failure event. The writeup warns that centralization fixes the search boundary but not semantics, and that health data such as prompts, model responses, and patient-linked attributes should be excluded or masked before ingestion. Short answer: send web, worker, and cron output to one centralized log service in the same JSON shape, but keep the rollout reversible and keep cron heartbeats separate. For a small healthtech SaaS measuring an AI agent loop, the useful baseline is service , job name , level , env , request id , trace id , span id , latency ms , and cost usd ; without that contract, a search product merely centralizes inconsistent text. This is a reasonable low-complexity alternative to operating a full ELK stack. It is not a complete observability system. Logs can answer what ran and what it reported, while a Healthchecks-style monitor must answer whether a scheduled job ran at all, and alert delivery needs either a separate monitor or a polling script when the logging service has no threshold rules or notification routing. An AI agent loop crosses several process boundaries: an HTTP request accepts work, a worker calls one or more models, and a cron job may reconcile usage later. If each process names duration, cost, and correlation fields differently, incident response becomes a manual join performed under pressure. The first capacity-planning question is equally awkward: did p95 latency rise because loops took more steps, or because one step became slower? Centralization fixes the search boundary, not the semantics. Emit one event per meaningful state transition, preserve the same request id across the web and worker processes, and record job name for scheduled work. trace id and span id are useful correlation handles, but they do not turn a log search interface into distributed tracing; there is no span-tree query here. Be strict about health data. OWASP's logging guidance calls out data that should usually be excluded, masked, sanitized, hashed, or encrypted. In a healthtech system, prompts, model responses, access tokens, session identifiers, and patient-linked attributes should not drift into logs because a convenient JSON serializer captured an entire object. This deserves a schema review before ingestion starts, particularly when the service has no per-user deletion interface, bulk export, or subscription API and does not expose retention or cold-storage configuration. Silence is different. A cron process that never starts emits no failure event, so no log query can discover it from the absent record alone. Pair each important schedule with an external heartbeat monitor and define its grace period from the job's actual completion SLO, not from wishful timing. The following Go program sends a compact allowlisted event to the centralized ingestion route. The same schema can be used by an HTTP service, queue worker, or cron process. It reads the key from the environment, sets an idempotency key so a retry cannot duplicate the write, honors Retry-After on a 429 response, uses exponential backoff otherwise, and returns real error bodies instead of treating every response as success. package main import "bytes" "context" "encoding/json" "fmt" "io" "log" "net/http" "os" "strconv" "strings" "time" type AgentEvent struct { Timestamp time.Time json:"timestamp" Service string json:"service" JobName string json:"job name,omitempty" Level string json:"level" Env string json:"env" RequestID string json:"request id" TraceID string json:"trace id,omitempty" SpanID string json:"span id,omitempty" Event string json:"event" LatencyMS int64 json:"latency ms" CostUSD float64 json:"cost usd" LoopSteps int json:"loop steps" ModelRoute string json:"model route" } type ingestRequest struct { Logs AgentEvent json:"logs" } func retryDelay resp http.Response, attempt int time.Duration { if value := resp.Header.Get "Retry-After" ; value = "" { if seconds, err := strconv.Atoi value ; err == nil { return time.Duration seconds time.Second } if retryAt, err := http.ParseTime value ; err == nil { if delay := time.Until retryAt ; delay 0 { return delay } } } return time.Second time.Duration 1<