Telemetry Budgets for One-Key Unified Image Generation APIs Scoring Candidates Across Multiple Models A developer outlines an architecture decision record for a unified image-generation API layer that places a narrow internal contract in front of model adapters, requiring bounded telemetry and evidence separation before any upstream provider is selected. The design records a stable request ID, adapter ID, normalized outcome, latency, image count, and metering units while keeping prompts, images, and candidate data out of routine logs, and it rejects silent fallback so a failed adapter cannot change policy regime, image dimensions, or cost attribution. The author argues that "one key" describes credential handling only, and that OpenAI, Claude, and Gemini must each pass discovery and contract tests before becoming routable. Put a narrow internal image-generation contract in front of model adapters, and give its telemetry a budget before choosing any upstream API. The deciding constraint is provider portability: a shared key or endpoint reduces integration work, but it does not make output semantics, retention, or observability portable. Short answer: record a stable request ID, adapter ID, normalized outcome, latency, image count, and metering units when the upstream reports them. Keep prompts and images out of routine logs. Compare OpenAI, Claude, Gemini, and any alternative by running the same contract tests, then retain only the measurements needed to make a routing or audit decision. This architecture decision record applies to a customer-support hiring workflow. A reviewer scores candidates against a job rubric, while an image generator produces a standardized role-play card from approved, non-candidate text. The score and generated card share a workflow ID, but they do not share a payload. That separation matters: provider switching belongs in the media boundary, while hiring evidence belongs in the application boundary. The first invariant is a deliberately small contract. The application submits an approved prompt, an aspect-ratio class, an output count, and an idempotency token. It receives internal asset IDs plus a normalized status. Provider-specific model names, revised-prompt fields, safety metadata, and delivery URLs remain inside the adapter. If a new upstream cannot represent a required input or output, capability negotiation rejects the request before generation; the adapter must not silently reinterpret it. The second invariant is evidence separation. Candidate rubric scores, reviewer notes, and hiring decisions stay in the system of record. The image service sees a workflow ID that is useful for correlation but carries no candidate name, resume excerpt, score, or protected characteristic. The generated role-play card is presentation material, not evidence for the employment decision. The third invariant is bounded telemetry. A trace or log event may identify the internal adapter, contract version, result class, attempt number, duration bucket, image count, and a coarse metering quantity. It must not use raw prompt text, request IDs, asset IDs, error messages, or candidate IDs as metric labels. Those values create an unbounded series space or duplicate sensitive content into another retention system. These are hard boundaries. A timeout after an accepted request is an ambiguous outcome and requires reconciliation by idempotency token. An unsupported capability is a preflight rejection, not an upstream failure. A policy rejection is not retried against another provider unless policy equivalence has been established in the contract. A malformed response is quarantined before an asset becomes visible to the reviewer. No silent fallback. That rule prevents a superficially successful request from changing its policy regime, image dimensions, or cost attribution merely because one adapter failed. It costs flexibility. This approach has limitations and trade-offs. It is not suitable when the team needs immediate access to every provider-specific control, when experimental response fields are the object of study, or when one upstream is already a permanent organizational dependency. A direct provider adapter is the better boundary in those cases. The unified layer is justified only when switching and comparable governance matter more than exposing the widest surface area. Names are poor selection criteria. “One key” describes credential handling; it says nothing about whether three adapter targets expose equivalent image outputs. The words OpenAI, Claude, and Gemini in a requirements document should therefore be treated as requested integration labels, not proof of a common capability. Each target has to pass discovery and contract tests before it becomes routable. | Adapter target | Admission rule | Portability evidence | Telemetry boundary | |---|---|---|---| | OpenAI-labelled integration | Enable only an image-output configuration accepted by the contract probe | Golden-request schema, normalized result, and asset validation pass | Internal adapter ID; upstream model detail belongs in a low-cardinality lookup | | Claude-labelled integration | Reject image-output routing unless discovery and tests establish the required output capability | An explicit unsupported result is preferable to substituting text output | Record capability rejection, not prompt or candidate data | | Gemini-labelled integration | Enable only the tested image-output configuration, independent of other Gemini capabilities | The same fixture, dimensions policy, and asset checks pass | Preserve reported usage separately from estimated allocation | | Another text-to-image integration | Add it through the same adapter interface; do not widen the application contract for one provider | Replay the complete fixture set and rollback test | Assign a stable adapter ID and the same outcome taxonomy | This table does not rank products, and it intentionally makes no claim that a brand label guarantees a particular modality. It defines the evidence an engineering team must collect. A useful evaluation fixture contains approved role-play text, expected output count, permitted media types, minimum and maximum byte bounds, and a deterministic list of validations. Pixel equality is not a sensible invariant for generative output; contract conformance is. The team should deploy a new adapter dark first: accept no user traffic, run the fixed fixtures, validate parsing and storage, and compare normalized telemetry. Next, allow a small, explicitly assigned cohort rather than a hidden fallback. Promotion requires successful reconciliation of ambiguous outcomes and a tested rollback to the prior routing configuration. The application-facing request should expose business constraints, not a vendor's parameter vocabulary. This proposed contract uses a reserved example domain and curl so the boundary is visible without implying a commercial endpoint: curl --fail-with-body --request POST \ --url https://images.example/api/generations \ --header 'Authorization: Bearer ${IMAGE GATEWAY KEY}' \ --header 'Content-Type: application/json' \ --header 'Idempotency-Key: hiring-demo-0187-card-01' \ --data '{ "contract version": "2026-01", "workflow id": "hiring-demo-0187", "purpose": "support-roleplay-card", "prompt": "A neutral help-desk workspace used for a customer support role-play exercise", "aspect ratio": "landscape", "output count": 1 }' The prompt is approved scenario copy, not candidate input. The idempotency key makes retries refer to one intended generation, while workflow id joins operational events to the application workflow without becoming a metric label. The gateway chooses only among adapters that declared compatibility with contract version 2026-01 and passed the fixture suite. The response should contain internal asset references, not an upstream URL that leaks provider choice into the caller. Persist the raw provider response only when a defined audit or debugging need justifies its retention; otherwise, store the normalized fields and a short-lived restricted diagnostic record. An upstream error body can contain echoed input, so copying it wholesale into a general log defeats prompt minimization. That is deliberate. Error handling follows the boundary established above. Definite preflight rejections return immediately. Definite failures can be retried only under the idempotency policy. Ambiguous timeouts enter reconciliation. Success is published after media-type, byte-bound, and decode checks, followed by durable storage of the internal asset. Observability volume is arithmetic, not atmosphere. Suppose the service emits 8 events per generation, receives 50,000 generation requests per day, stores an average of 700 bytes per event after encoding, and retains them for 30 days. The uncompressed planning quantity is: 8 × 50,000 × 700 × 30 = 8,400,000,000 bytes , or 8.4 GB using decimal units. That number is an example calculation, not a benchmark or a vendor bill. Index overhead, replication, compression, and derived data are deliberately excluded because they depend on the selected telemetry system. The point is to make every proposed event earn its retention period. Cardinality needs the same treatment. With 4 adapters, 6 normalized outcomes, 3 deployment environments, and 10 latency buckets, the full Cartesian upper bound for one histogram family is 4 × 6 × 3 × 10 = 720 bucket combinations before standard histogram series are considered. Adding 50,000 daily workflow IDs as a label changes the order of magnitude for no operational benefit. Put workflow IDs in sampled traces or indexed diagnostic events, where access and retention can be controlled, rather than in metrics. Sampling should follow decisions. Keep all capability rejections, policy rejections, malformed responses, and ambiguous outcomes for the short diagnostic window because each can change routing or correctness. Sample routine successes at a declared rate for trace analysis, but retain aggregated counters for every request. If success traces are sampled at 1%, inverse weighting can estimate counts, yet it cannot reconstruct rare detail that was never retained. Do not use sampled traces as the billing ledger. Metering also needs two fields: reported usage and internal allocation estimate. A provider-reported quantity is evidence from that adapter; an estimate is a planning value derived by the gateway. Combining them into one number hides uncertainty. Financial reconciliation should use immutable request-level records with restricted access and a retention policy distinct from debugging logs. Prompts deserve their own decision. Hashing a prompt does not automatically make it a safe metric label: unique hashes still produce high cardinality, and predictable prompt spaces may permit guessing. A bounded template ID answers the useful question—what class of role-play card was requested—without retaining the text in every event. The rejected option is a transparent proxy that logs complete requests and responses, exposes provider model names to the client, and forwards every upstream field. It is attractive during a short evaluation because engineers can inspect everything and adopt new provider options immediately. It is also the wrong steady-state boundary for this workflow. Application code becomes coupled to each upstream schema, provider switching reaches into callers, and telemetry volume grows with image payloads and verbose response bodies. There is a valid use case for that design: an isolated, access-controlled model laboratory with synthetic prompts, disposable outputs, brief retention, and no candidate data. Full-fidelity capture can accelerate adapter discovery there. The promotion path should convert what was learned into fixtures and normalized fields, then disable raw capture before real hiring workflows arrive. Vector search is separate. Embeddings convert inputs into vectors, and pgvector adds vector similarity search to Postgres; neither source establishes an image-generation contract. If the hiring application later retrieves rubric passages semantically, that subsystem should use its own data classification, retention math, and evaluation set rather than being folded into image telemetry because both happen to involve AI. The final selection rule is compact: choose only adapters that pass the same capability, idempotency, asset-validation, reconciliation, and rollback tests under the internal contract. Then select routing using measured quality and operational evidence from the approved scenario. A single credential can simplify secret distribution, but portability comes from the boundary, tests, and telemetry budget.