cd /news/ai-infrastructure/mid-stream-failover-made-my-chat-api… · home › topics › ai-infrastructure › article
[ARTICLE · art-144501] src=dev.to ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Mid-stream failover made my chat API answer the same prompt twice — switch models before the first token or not at all

A developer running an OpenAI-compatible /v1/chat/completions endpoint with 50+ models behind one integration traced a client's duplicated, spliced-together answer to failover logic that re-dispatched a prompt after an upstream returned a 429 mid-stream, appending the second provider's output to the truncated first. The fix restricts failover to the zero-tokens-sent state, so errors arriving after the first content delta end the stream with an error payload instead of rerouting, and adds a flushed-bytes counter to distinguish "429 after 0 bytes" from "429 after 40 bytes" in alerts. The developer reports zero double-answer reports since the change.

by read2 min views1 publishedOct 3, 2026

A client's transcript came in looking like a slip of the tongue from a language model: three cut-off sentences about caching, then the exact same question answered again, in a slightly different voice, spliced together as one continuous message.

The culprit was failover logic in my chat router. I run an OpenAI-compatible /v1/chat/completions endpoint with 50+ models behind one integration (https://x402.freeq.one/tools/llm_chat.html), and the eco tier picks the cheapest healthy provider for each request. One night an upstream started handing out 429 s — but only after accepting the connection and streaming about forty tokens. My router treated 429 as retryable, re-dispatched the prompt to the next provider on the list, and appended the new stream to the old one. The client's SDK happily glued both halves into a single message.

The fix is a rule I now enforce in the stream state machine: failover is only legal in the zero tokens sent state. Connection refused, a 401 or 404 on handshake, a 429 before any content delta — those are safe to reroute, and the client never notices. After the first content delta, a silent model swap is worse than a plain failure: you've already committed to one voice, one context window, one set of capabilities, and splicing a second model onto a truncated first half produces text that reads like two ghosts sharing one keyboard.

So after the first token, the options narrow to: end the stream with finish_reason set plus an error payload in metadata, or emit an explicit error event. Never append. I also added a flushed-bytes counter per request, so logs can prove which state a failure landed in — "429 after 0 bytes" and "429 after 40 bytes" are separate alert categories now, and only the first one reroutes.

Since the change: zero double-answer reports. If you run any OpenAI-compatible shim in front of multiple providers, classify upstream errors by stream state before you write the retry loop. Most proxy bugs aren't in model selection — they're in transitions between states you didn't think needed to be states.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/mid-stream-failover-…] indexed:0 read:2min 2026-10-03 · —