AI API call hangs in production, then 504: timeouts and budgets A developer analyzed 1,662 public posts from builders whose apps broke at or after launch and verified 215 recent cases, finding that AI model and API failures account for 4% of them, with six verified cases sharing a pattern of outbound or AI calls lacking timeouts, budgets, and fallbacks. The writeup shows that default SDK timeouts (10 minutes per attempt for OpenAI and Anthropic Node SDKs) exceed the limits of outer layers like Cloudflare's 125-second proxy timeout and Vercel's 300-second function limit, so requests are killed with 504/524 errors before application-level catch blocks or logging can run. The developer also demonstrates a local fake-provider test suite that reproduces never-answering, never-finishing agent, and spend-limit 429 scenarios without an API key. We read 1,662 public posts from builders whose apps broke at or after launch and verified 215 recent cases. AI model and API failures make up 4% of them. Six of the verified cases share one pattern: an outbound or AI call with no timeout, no budget and no fallback. In most of the six, nobody noticed for a while, because the app didn't log model calls or their outcomes. From the builder's side it shows up as one of these: All three have the same cause. The call has no limits of its own, so the only limits that apply are set by other layers. Here is a Next.js App Router route using the OpenAI Node SDK with its defaults. Most generated code ships in this shape. python // app/api/summary/route.ts import OpenAI from "openai"; const openai = new OpenAI ; // defaults: 10 min per attempt, 2 retries export async function POST req: Request { const { text } = await req.json ; try { const r = await openai.chat.completions.create { model: "gpt-5-mini", messages: { role: "user", content: Summarise this:\n${text} } , } ; return Response.json { summary: r.choices 0 .message.content } ; } catch { // the user sees this with a 200, and nothing is logged return Response.json { summary: "Summary unavailable right now." } ; } } You don't need a provider outage to see it fail. The SDK reads OPENAI BASE URL , so you can point it at a server that accepts connections and never replies: js terminal 1: a provider that never answers node -e "require 'http' .createServer = {} .listen 4010 " terminal 2 OPENAI BASE URL=http://127.0.0.1:4010/v1 OPENAI API KEY=test npm run dev terminal 3 curl -s -X POST localhost:3000/api/summary \ -H 'content-type: application/json' -d '{"text":"hello"}' -w '\n%{time total}s\n' curl waits for at least 15 minutes. On Vercel, the same request is killed at 300 seconds with a 504 FUNCTION INVOCATION TIMEOUT before the catch runs, so nothing is logged. Behind Cloudflare the user gets a 524 even sooner. Now make the stub answer 500 instead: node -e "require 'http' .createServer q, s = { s.statusCode = 500; s.end '{}' } .listen 4010 " . After a few seconds of retries the route returns 200 with the fallback sentence. Every monitor you have counts that as a success. Every layer between the browser and the model has its own clock. In the default setup the clock in your own code is the longest, so an outer layer always runs out first. | Layer | Default limit | What happens | |---|---|---| | Cloudflare proxy | 125 s with no response from the origin | 524 to the user | | Vercel function fluid compute | maxDuration 300 s | function terminated, 504 | | Node's fetch undici | 300 s waiting for response headers | connection error | | OpenAI and Anthropic Node SDKs | 10 minutes per attempt | APIConnectionTimeoutError | Here is what the official docs say about each part: max tokens , Anthropic's SDK raises its default timeout above 10 minutes, up to 60. maxDuration . Your catch never runs, and neither does the logging inside it. project spend limit exceeded . Because the SDKs retry every 429, each user's request retries an error that won't clear until the limit resets. Anthropic returns 400 invalid request error when a workspace spend limit is reached, and retry logic that only looks for 429 won't recognise it. This suite starts a fake provider on localhost with three behaviours: it never answers, it plays an agent that never finishes, or it returns a spend-limit 429. It imports the wrapper from the next section, and its assertions define what that wrapper has to do. It needs no API key and costs nothing to run. python // lib/ai-guards.test.ts import http from "node:http"; import type { AddressInfo } from "node:net"; import { afterAll, beforeAll, expect, test } from "vitest"; import { callModel, makeClient, newRun, runAgent } from "./ai"; process.env.OPENAI API KEY ??= "test"; let hits = 0; // One fake provider; the URL prefix picks its behaviour. const server = http.createServer req, res = { hits++; if req.url?.startsWith "/stall/" return; // accept, never answer res.setHeader "content-type", "application/json" ; if req.url?.startsWith "/spent/" { res.statusCode = 429; return res.end JSON.stringify { error: { code: "project spend limit exceeded", message: "limit reached" } } ; } res.end JSON.stringify { // "/loop/": a model that never says it is done id: "c1", object: "chat.completion", created: 0, model: "stub", choices: { index: 0, finish reason: "stop", message: { role: "assistant", content: "keep going" } } , usage: { prompt tokens: 500, completion tokens: 500, total tokens: 1000 }, } ; } ; const client = mode: string = makeClient http://127.0.0.1:${ server.address as AddressInfo .port}/${mode}/v1 ; const hi = { role: "user" as const, content: "hi" } ; beforeAll = new Promise