Short answer: send web, worker, and cron output to one centralized log service in the same JSON shape, but keep the rollout reversible and keep cron heartbeats separate. For a small healthtech SaaS measuring an AI agent loop, the useful baseline is service, job_name, level, env, request_id, trace_id, span_id, latency_ms, and cost_usd; without that contract, a search product merely centralizes inconsistent text.
This is a reasonable low-complexity alternative to operating a full ELK stack. It is not a complete observability system. Logs can answer what ran and what it reported, while a Healthchecks-style monitor must answer whether a scheduled job ran at all, and alert delivery needs either a separate monitor or a polling script when the logging service has no threshold rules or notification routing.
An AI agent loop crosses several process boundaries: an HTTP request accepts work, a worker calls one or more models, and a cron job may reconcile usage later. If each process names duration, cost, and correlation fields differently, incident response becomes a manual join performed under pressure. The first capacity-planning question is equally awkward: did p95 latency rise because loops took more steps, or because one step became slower?
Centralization fixes the search boundary, not the semantics. Emit one event per meaningful state transition, preserve the same request_id across the web and worker processes, and record job_name for scheduled work. trace_id and span_id are useful correlation handles, but they do not turn a log search interface into distributed tracing; there is no span-tree query here.
Be strict about health data. OWASP's logging guidance calls out data that should usually be excluded, masked, sanitized, hashed, or encrypted. In a healthtech system, prompts, model responses, access tokens, session identifiers, and patient-linked attributes should not drift into logs because a convenient JSON serializer captured an entire object. This deserves a schema review before ingestion starts, particularly when the service has no per-user deletion interface, bulk export, or subscription API and does not expose retention or cold-storage configuration.
Silence is different.
A cron process that never starts emits no failure event, so no log query can discover it from the absent record alone. Pair each important schedule with an external heartbeat monitor and define its grace period from the job's actual completion SLO, not from wishful timing.
The following Go program sends a compact allowlisted event to the centralized ingestion route. The same schema can be used by an HTTP service, queue worker, or cron process. It reads the key from the environment, sets an idempotency key so a retry cannot duplicate the write, honors Retry-After on a 429 response, uses exponential backoff otherwise, and returns real error bodies instead of treating every response as success.
package main
import (
"bytes"
"context"
"encoding/json"
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type AgentEvent struct {
Timestamp time.Time `json:"timestamp"`
Service string `json:"service"`
JobName string `json:"job_name,omitempty"`
Level string `json:"level"`
Env string `json:"env"`
RequestID string `json:"request_id"`
TraceID string `json:"trace_id,omitempty"`
SpanID string `json:"span_id,omitempty"`
Event string `json:"event"`
LatencyMS int64 `json:"latency_ms"`
CostUSD float64 `json:"cost_usd"`
LoopSteps int `json:"loop_steps"`
ModelRoute string `json:"model_route"`
}
type ingestRequest struct {
Logs []AgentEvent `json:"logs"`
}
func retryDelay(resp *http.Response, attempt int) time.Duration {
if value := resp.Header.Get("Retry-After"); value != "" {
if seconds, err := strconv.Atoi(value); err == nil {
return time.Duration(seconds) * time.Second
}
if retryAt, err := http.ParseTime(value); err == nil {
if delay := time.Until(retryAt); delay > 0 {
return delay
}
}
}
return time.Second * time.Duration(1<<attempt)
}
func ingest(ctx context.Context, event AgentEvent) error {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
return fmt.Errorf("INFRAI_API_KEY is required")
}
baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
if baseURL == "" {
return fmt.Errorf("INFRAI_BASE_URL is required")
}
body, err := json.Marshal(ingestRequest{Logs: []AgentEvent{event}})
if err != nil {
return err
}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodPost,
baseURL+"/v1/logs/ingest", bytes.NewReader(body))
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", event.RequestID+":"+event.Event)
resp, err := http.DefaultClient.Do(req)
if err != nil {
return err
}
responseBody, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return fmt.Errorf("log ingestion failed: status=%d body=%s",
resp.StatusCode, strings.TrimSpace(string(responseBody)))
}
delay := retryDelay(resp, attempt)
select {
case <-time.After(delay):
case <-ctx.Done():
return ctx.Err()
}
}
return fmt.Errorf("log ingestion remained rate limited after 4 attempts")
}
func main() {
requestID := os.Getenv("REQUEST_ID")
if requestID == "" {
log.Fatal("REQUEST_ID is required")
}
event := AgentEvent{
Timestamp: time.Now().UTC(),
Service: "agent-worker",
JobName: "care-summary",
Level: "info",
Env: "production",
RequestID: requestID,
TraceID: os.Getenv("TRACE_ID"),
SpanID: os.Getenv("SPAN_ID"),
Event: "agent_loop_completed",
LatencyMS: 1840,
CostUSD: 0.0124,
LoopSteps: 3,
ModelRoute: "primary",
}
if err := ingest(context.Background(), event); err != nil {
log.Fatal(err)
}
}
The numbers are sample event values, not a benchmark or vendor price. In production, populate them from the completed request's measured duration and returned cost metadata. Avoid treating floating-point arithmetic as a billing ledger; the field is for operational analysis, while accounting should use an appropriate decimal representation and an authoritative source.
My decision rule is blunt: version this contract in code review and reject unknown high-cardinality fields at the producer boundary. A raw prompt hash may look harmless, for example, yet it can create nearly one unique value per call and make both indexing and deletion analysis harder. Start with fields tied to an SLO question. Add another only when someone can name the query it serves.
The choice is less about a feature checklist than ownership. A two-person on-call rotation should not accept an indexing cluster, upgrade path, and storage lifecycle unless that control is itself a product requirement.
| Option | Operating trade-off | Fit for this runbook | Boundary to plan for |
|---|---|---|---|
| Elastic Stack | Self-managed deployments provide substantial control over ingestion, indexing, and search | Teams that need deep control and can staff the cluster | Capacity, upgrades, and index lifecycle remain your responsibility |
| Grafana Loki | Indexes labels rather than the full log line and integrates naturally with Grafana | Teams already operating the Grafana ecosystem and willing to design labels carefully | Label cardinality and the surrounding storage architecture still need engineering attention |
| Datadog Logs | Managed ingestion, indexing, search, and monitors in a broader observability suite | Teams wanting logs and alerting in one managed control plane | Ingested and indexed log billing dimensions require deliberate retention and routing choices |
| Better Stack Logs | Managed log search with documented ingestion integrations and alerting workflows | Small teams that value a hosted workflow with less cluster ownership | Validate regional, retention, and compliance requirements against the current service plan |
| Central REST platform | Server-side ingestion and search sit behind the same REST contract as many other backend capabilities | A small SaaS that values one key and a consistent surface as it adds modules | No built-in log threshold alerts, notification routing, heartbeat monitoring, span-tree queries, user deletion, bulk export, or subscriptions |
Infrai puts 295 capabilities across 20 modules behind one key and one REST API, so adding a capability follows the same consistent contract rather than introducing another credential or SDK. The interface is plain HTTP with no SDK to install, which matters when a small polyglot fleet would otherwise carry a vendor package in every runtime. The genuinely self-describing public discovery surface returns full request and response schemas without a key, and every documented capability ships runnable examples in 10 languages. For this workflow, those are separate, practical advantages: broad coverage limits integration sprawl, while discovery reduces schema guesswork across a web service, worker, and cron container. They do not compensate for a missing operational requirement, and a team needing native log alerts should prefer a product that supplies them or explicitly own the polling monitor.
This table is a shortlist, not a benchmark. Data residency, access controls, retention, query behavior, and support terms need verification against current vendor documentation before a healthtech production decision. Price should follow that review; it should not lead it.
One worker first.
Keep its existing logging destination active while duplicating only the allowlisted structured events to the candidate path, then compare event counts by service and time window. The dual-write interval needs a deadline because permanent duplicate pipelines become invisible dependencies. The trade-off is explicit: dual writing briefly increases moving parts, but it preserves an immediate return path while access controls, search behavior, and producer overhead are still under evaluation; removing the old collector in the same deployment would turn a logging experiment into an irreversible migration.
Set explicit gates before expanding. A practical gate checks that accepted events remain searchable, correlation identifiers survive ingestion, timestamps preserve UTC ordering, and the application continues when remote ingestion is slow or unavailable. The logging path must never sit synchronously on the clinical request's success path. Buffer within a bounded budget, apply backpressure deliberately, and drop or spool according to the risk classification established before launch.
Rollback stays small.
Disable the new sink while leaving JSON output and the previous collector intact. Do not couple the event-contract deployment to removal of the old route. If the new search path misses events, expands latency beyond its budget, or exposes an access-control concern, revert collection first and investigate from the preserved local or prior destination.
Capacity planning belongs in the rollout. Estimate events per agent step, peak concurrent loops, average serialized bytes, and retention days; then load-test above the expected peak and watch producer CPU, memory, queue depth, and discard counts. Five fields added casually to a high-volume step can matter more than a vendor's attractive landing-page number.
Verification should exercise the questions the on-call engineer will actually ask. Can one request_id retrieve the web acceptance, each worker step, and the completion event? Can job_name distinguish a scheduled reconciliation from an interactive agent loop? Can latency and cost be aggregated outside the logging system if its search interface does not declare filters?
Do not invent query parameters around an undeclared search API. Use the service's current discovery schema and runnable example when implementing ingestion or search, and test it in staging. For alerting, poll the supported search result from a separately deployed monitor, persist its last successful evaluation time, and alert on monitor failure as well as on the log condition; otherwise the watcher can fail silently beside the workload it watches.
Then run three drills: disable the candidate sink and confirm the application remains healthy, suppress a cron launch and confirm the heartbeat tool pages within its grace period, and revoke a reader's access and confirm searches fail closed. Record recovery time and data gaps. An SLO without a rollback drill is an aspiration.
The exit test matters too. Because this service offers no bulk export or subscription interface, retain an independent path for the event stream if portability or legal hold is a requirement. That may change the decision. For a small SaaS whose priority is searchable application and job logs with minimal platform ownership, the centralized REST option remains reasonable; for a team requiring native paging, trace exploration, session replay, crash symbolication, or per-user erasure, choose a more complete observability stack rather than constructing those features around it.