{"slug": "graceful-degradation-for-ai-features", "title": "Graceful Degradation for AI Features", "summary": "A developer argues that AI features should be designed with explicit degradation contracts, mapping each failure mode to a defined product behavior rather than letting a model outage disable an entire workspace. The piece lays out five modes — explicit unavailable state, deterministic non-AI path, delayed processing, reduced functionality, and disclosed stale result — and recommends HTTP status codes such as 503 with Retry-After, 429, and 401 to signal them at the boundary.", "body_md": "An AI feature does not have to return an answer during an outage.\n\nIf it cannot meet its normal contract, a clear unavailable state may be the best response. Other features can fall back to ordinary application code, accept the work for later, keep an independent capability available, or show an older result with a visible warning.\n\nDecide that behavior as part of the product contract. A catch block can detect a timeout, but it cannot decide what the user should still be able to do.\n\nConsider a support workspace that can retrieve policy documents and prepare a reply for a ticket. The model call is one part of a larger user flow:\n\n``` php\nopen the ticket\n    -> review customer history\n    -> retrieve approved policy\n    -> prepare an AI draft\n    -> edit and save the reply\n```\n\nIf the model is unavailable, the human support agent should still be able to open the ticket, inspect its history, and write a reply manually, provided those capabilities do not depend on the model. Disabling the entire workspace because one optional capability failed would turn a dependency incident into a larger product outage.\n\nDifferent dependency failures call for different behavior across user flows. A live suggestion can disappear without much harm. A long-running document analysis may be worth queueing. An answer that must be grounded in current policy should stop when the application cannot obtain that policy evidence, even if the model still responds.\n\nStart with what the user needs to accomplish, then decide how the flow should behave during each failure.\n\nFor each important flow, define:\n\nWrite those decisions down as the degradation contract for the flow.\n\nStart with five possible product behaviors, then add failover machinery where the product needs it.\n\n| Mode | Use it when | Example | \n|---|---|---|\n| Explicit unavailable state | No safe or useful substitute exists | Disable AI reply generation while leaving the ticket editor available | \n| Deterministic non-AI path | Rules, templates, or ordinary application code can still complete a narrower job | Insert an approved acknowledgement template instead of generating a reply | \n| Delayed processing | The result remains useful after the current request ends | Accept a document summary job and expose its status | \n| Reduced functionality | One independent part of the feature still has value | Show retrieved policy documents without generating a synthesis | \n| Disclosed stale result | A previous result remains relevant within a defined age and input version | Show an earlier draft for the same ticket revision and label when it was produced | \n\nThese modes can coexist in one product. The workspace might display an earlier draft, offer an approved template, and accept a new drafting job at the same time. Give each operation a clear outcome and show which workspace capabilities are currently available.\n\nIf the feature cannot produce a result that meets its normal contract, say so. Keep the unaffected application available and make the boundary visible. \"AI reply drafting is temporarily unavailable\" is more useful than a spinner that eventually disappears or a generic error at the top of the whole page.\n\nDo not replace a grounded answer with an ungrounded one just to avoid showing an error. If the application cannot obtain the required policy evidence from an approved source, it cannot prepare that reply. A failed retriever may leave another approved source available. A responding model does not make missing evidence safe to ignore.\n\nAt an HTTP boundary, a temporary inability to handle the operation may map to `503 Service Unavailable`. Send `Retry-After` only when the application can suggest a meaningful interval. An arbitrary fixed interval can make clients retry together. Clients can use jitter to spread their attempts. A per-client rate limit may call for `429 Too Many Requests`. Invalid caller credentials call for `401 Unauthorized` with a `WWW-Authenticate` challenge (provided those capabilities do not depend on the model). Valid credentials without sufficient permission generally call for `403 Forbidden`. Choose the status from the failure and the API contract, not from the fact that generation did not finish.\n\nDo not blindly forward an upstream provider's status code. Map dependency failures according to your own API contract so callers can distinguish their rate limits from temporary service-capacity problems.\n\nA non-AI path works when the product already has a narrower, predictable way to help.\n\nFor the support example, the application might offer an approved acknowledgement template. Its text and selection rules can be reviewed and tested. The template does less than a generated draft, and it does not need to imitate one.\n\nDo not present a canned sentence as if the assistant wrote a case-specific answer. Label the alternative for what it is and let the user decide whether it helps.\n\nDeterministic paths are especially useful when AI improves speed but does not own the underlying job. They are much harder to add later if the user interface and application service assume that every successful flow must contain model output.\n\nA document analysis, report, or batch classification can remain useful after the interactive request ends. The application can durably accept the work, return an operation identifier, and process it when the dependency recovers. An HTTP API commonly represents this with `202 Accepted` and a status resource.\n\nPersist the job and its input identity before returning `202`. If dispatch uses a separate queue, write an outbox entry in the same transaction as the job or use another reliable dispatch path. Saving the job and then publishing a message leaves a failure window.\n\nEnforce submission idempotency keys atomically, scoped to the caller and operation and bound to the inputs. A retry with the same key and inputs returns the existing job. Different inputs are a conflict. Message delivery can repeat, so workers need an atomic claim and idempotent result persistence. For external calls, use the other system's idempotency support where available. Otherwise, define how to handle an uncertain outcome before retrying.\n\nDefine how long the job may wait, when it expires, whether the user can cancel it, and what terminal failure looks like. Check the ticket revision, policy version, and required permissions before execution and again before making the result available. If required permissions were revoked during generation, do not expose the result. If relevant inputs changed, reject the draft or apply the product's stale-result policy. Guard the result write with an expected-version condition where the data shares a store. Synchronous generation has the same race.\n\nAuthorize status polling, cancellation, and result retrieval against current permissions. A job ID is not proof of access. Revoked ticket access must also block a completed result.\n\n[RFC 9110](https://www.rfc-editor.org/rfc/rfc9110.html#section-15.3.3) says `202` means processing was accepted, not that it will eventually succeed. An in-memory task can disappear during a restart.\n\nDo not queue every failed chat message by default. A reply that arrives 20 minutes after the conversation moved on may be worse than an immediate unavailable state.\n\nIf retrieval works but the model does not, the application may still show retrieved source documents after the usual relevance, freshness, and authorization checks. If required evidence is unavailable, the application may keep unrelated writing assistance available while disabling policy-grounded answers. If optional enrichment fails, the core transaction may continue without it.\n\nManual editing and policy search are separate workspace capabilities. Keep them available when their own dependencies work. Policy search may itself use embeddings or reranking, so it will not survive every AI outage.\n\nThe reply-preparation operation can keep retrieval, generation, validation, and persistence together. Separate unrelated ticket viewing and editing from that operation so a drafting failure does not take down the workspace. The user interface also needs to distinguish \"the page failed\" from \"this capability is temporarily unavailable\".\n\nReduced mode must preserve the original safety and authorization rules. Dependency failure is not permission to skip access checks, output validation, required audit records, or tenant isolation.\n\nFor a previous support draft, a time-to-live is not enough. The application should also know that the ticket revision, policy version, tenant, and relevant user context still match. It must recheck current authorization before returning the draft as an accepted previous result. If the ticket changed after the draft was produced, a three-minute-old draft may already be wrong.\n\nReturn the time the result was generated and label it as an earlier result. Define a maximum age and the events that invalidate it. Sensitive or consequential answers may not permit stale use at all.\n\nThis is stricter than caching a public, slow-changing reference value. Generated content can contain private context and can depend on inputs that are not obvious from the final text. A cache key based only on the user's question is rarely enough.\n\nFor the support flow, a decision table might look like this:\n\n| Failed capability | What remains available | Chosen behavior | Why | \n|---|---|---|---|\n| Model generation | Ticket view, manual editor, templates, policy search | Offer a template or manual path. Optionally queue a draft | The support agent can still complete the job without pretending a generated reply exists | \n| Required policy evidence unavailable | Ticket view, manual editor, non-policy tools | Disable grounded reply generation | The application cannot support current policy claims without approved evidence | \n| Durable job acceptance unavailable | Interactive paths | Do not offer delayed processing | If the application cannot durably accept and later dispatch the job, it must not return `202` . | \n| Previous-draft store | Normal live generation | Continue without stale-result mode | Cached output is an optional convenience, not a critical dependency | \n| Live authorization required but unavailable | Operations that can establish authorization with the required guarantees | Stop operations that require a live decision | A locally validated token may suffice for some operations, but current permissions or revocation checks may be required for others | \n\nIf the application cannot establish authorization with the guarantees an operation requires, that operation must stop.\n\nUse these questions when choosing a mode:\n\nIf the team cannot answer the second question, default to unavailable until it can.\n\nCallers should not have to parse exception messages to learn whether work was queued, a template was returned, or the feature is unavailable.\n\nModel reply preparation as an application outcome:\n\n```\npublic enum ReplyMode\n{\n    Generated,\n    ApprovedTemplate,\n    Queued,\n    PreviousDraft,\n    Unavailable\n}\n\npublic enum DegradationReason\n{\n    None,\n    ModelUnavailable,\n    RetrievalUnavailable,\n    RequiredEvidenceUnavailable,\n    RequestDeadlineExceeded\n}\n\npublic sealed record PrepareReplyResult(\n    ReplyMode Mode,\n    DegradationReason Reason,\n    string? Content,\n    Guid? JobId,\n    DateTimeOffset? GeneratedAt,\n    string? InputVersion);\n```\n\nThe type is intentionally limited to reply preparation. `ReducedFunctionality` is a state of the wider workspace, not another kind of reply. Ticket viewing, policy search, and drafting each need their own availability state. `RetrievalUnavailable` means a required retrieval dependency failed. `RequiredEvidenceUnavailable` means retrieval ran but no approved source provided enough evidence. An approved template chosen during normal work has `Reason.None`. One offered because generation failed has `Reason.ModelUnavailable`.\n\n`PreviousDraft` means more than \"cache hit\". It tells the endpoint and user interface to show the age and reduced-mode notice. Here, `InputVersion` is a composite identifier for the ticket revision, policy version, tenant, and relevant user context used to produce the draft. Compare those inputs and recheck current authorization before returning it as an accepted result. `Queued` requires a job identifier and must not contain a pretend final result. `Unavailable` is a handled outcome, not an unexpected exception.\n\nProduction code can use factory methods or separate result types to prevent invalid combinations. The mode should cross the service boundary explicitly.\n\nMap it deliberately at the edge:\n\n`Generated`, `ApprovedTemplate`, and an accepted `PreviousDraft` can return a normal response with mode metadata.`Queued` can return `202 Accepted` with a status URL.`Unavailable` can return a typed application error or `503` for temporary service inability, depending on the API contract. Handle per-client rate limits with `429` and invalid credentials with `401` at the appropriate boundary.\n`RequestDeadlineExceeded` means this request ran out of time. A token limit, provider rate limit, or monetary budget needs a different reason and may require a different response.\n\nThe user-facing message should describe what changed and what remains possible. Keep provider names and circuit state out of the UI. Record structured failure information in telemetry, and sanitize exception details so prompts, credentials, or customer data do not leak into logs.\n\nA request can be technically successful while the product is degraded. If a template returns with HTTP 200, an availability dashboard that counts only status codes will miss the model outage.\n\nRecord enough structured data to distinguish the paths:\n\nDo not copy prompts, ticket content, or generated replies into telemetry to explain the mode. Stable identifiers and outcome categories should be enough for operational analysis.\n\nTrack the share of requests using each mode and reason together. Normal template use is not degraded traffic. If templates replace generation for 30 percent of requests for two days, the model incident is still visible even though those requests returned successfully.\n\nForce each dependency outcome with test doubles or fault injection. Verify that:\n\nAlso test combinations. The model may be unavailable while the queue is full. Retrieval may recover after a job was accepted but before it runs. A previous result may exist after the user's access was revoked.\n\nThese cases are where a fallback can become a data leak or a false promise. When the model recovers, release accumulated jobs at a bounded rate rather than sending the whole backlog at once.\n\nDesign a degraded mode when the AI capability supports a larger user job and the product can preserve useful, honest behavior without it. Good candidates include drafting assistance, optional summarization, enrichment, recommendations, and long-running analysis that can finish later.\n\nReturn an explicit unavailable result when every substitute would violate the feature's promise, required current evidence cannot be obtained, authorization cannot be confirmed, or delayed output would no longer help. Stopping one capability can be the behavior that keeps the rest of the product trustworthy.\n\nDo not treat a second model as an invisible replacement. Another model can change output quality, structured-output behavior, tool use, latency, and cost. That is a separate policy decision, not a transparent implementation detail.\n\nPick one user flow and write down what remains possible when each dependency fails. Put the resulting behavior in the application contract, show it in the UI, and test it under forced failure.", "url": "https://wpnews.pro/news/graceful-degradation-for-ai-features", "canonical_source": "https://dev.to/lukaswalter/graceful-degradation-for-ai-features-406b", "published_at": "2026-10-01 15:30:00+00:00", "updated_at": "2026-10-01 15:44:38.051714+00:00", "lang": "en", "topics": ["ai-products", "ai-tools", "developer-tools"], "entities": [], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/graceful-degradation-for-ai-features", "markdown": "https://wpnews.pro/news/graceful-degradation-for-ai-features.md", "text": "https://wpnews.pro/news/graceful-degradation-for-ai-features.txt", "jsonld": "https://wpnews.pro/news/graceful-degradation-for-ai-features.jsonld"}}