# Mid-stream failover made my chat API answer the same prompt twice — switch models before the first token or not at all

> Source: <https://dev.to/imapphelp/mid-stream-failover-made-my-chat-api-answer-the-same-prompt-twice-switch-models-before-the-first-13d0>
> Published: 2026-10-03 15:33:10+00:00

A client's transcript came in looking like a slip of the tongue from a language model: three cut-off sentences about caching, then the exact same question answered again, in a slightly different voice, spliced together as one continuous message.

The culprit was failover logic in my chat router. I run an OpenAI-compatible `/v1/chat/completions` endpoint with 50+ models behind one integration ([https://x402.freeq.one/tools/llm_chat.html](https://x402.freeq.one/tools/llm_chat.html)), and the eco tier picks the cheapest healthy provider for each request. One night an upstream started handing out `429` s — but only after accepting the connection and streaming about forty tokens. My router treated `429` as retryable, re-dispatched the prompt to the next provider on the list, and appended the new stream to the old one. The client's SDK happily glued both halves into a single message.

The fix is a rule I now enforce in the stream state machine: **failover is only legal in the `zero tokens sent` state.** Connection refused, a `401` or `404` on handshake, a `429` before any content delta — those are safe to reroute, and the client never notices. After the first content delta, a silent model swap is worse than a plain failure: you've already committed to one voice, one context window, one set of capabilities, and splicing a second model onto a truncated first half produces text that reads like two ghosts sharing one keyboard.

So after the first token, the options narrow to: end the stream with `finish_reason` set plus an error payload in metadata, or emit an explicit error event. Never append. I also added a flushed-bytes counter per request, so logs can prove which state a failure landed in — "429 after 0 bytes" and "429 after 40 bytes" are separate alert categories now, and only the first one reroutes.

Since the change: zero double-answer reports. If you run any OpenAI-compatible shim in front of multiple providers, classify upstream errors by stream state before you write the retry loop. Most proxy bugs aren't in model selection — they're in transitions between states you didn't think needed to be states.
