{"slug": "ai-api-call-hangs-in-production-then-504-timeouts-and-budgets", "title": "AI API call hangs in production, then 504: timeouts and budgets", "summary": "A developer analyzed 1,662 public posts from builders whose apps broke at or after launch and verified 215 recent cases, finding that AI model and API failures account for 4% of them, with six verified cases sharing a pattern of outbound or AI calls lacking timeouts, budgets, and fallbacks. The writeup shows that default SDK timeouts (10 minutes per attempt for OpenAI and Anthropic Node SDKs) exceed the limits of outer layers like Cloudflare's 125-second proxy timeout and Vercel's 300-second function limit, so requests are killed with 504/524 errors before application-level catch blocks or logging can run. The developer also demonstrates a local fake-provider test suite that reproduces never-answering, never-finishing agent, and spend-limit 429 scenarios without an API key.", "body_md": "We read 1,662 public posts from builders whose apps broke at or after launch and verified 215 recent cases. AI model and API failures make up 4% of them. Six of the verified cases share one pattern: an outbound or AI call with no timeout, no budget and no fallback. In most of the six, nobody noticed for a while, because the app didn't log model calls or their outcomes.\n\nFrom the builder's side it shows up as one of these:\n\nAll three have the same cause. The call has no limits of its own, so the only limits that apply are set by other layers.\n\nHere is a Next.js App Router route using the OpenAI Node SDK with its defaults. Most generated code ships in this shape.\n\n``` python\n// app/api/summary/route.ts\nimport OpenAI from \"openai\";\n\nconst openai = new OpenAI(); // defaults: 10 min per attempt, 2 retries\n\nexport async function POST(req: Request) {\n  const { text } = await req.json();\n  try {\n    const r = await openai.chat.completions.create({\n      model: \"gpt-5-mini\",\n      messages: [{ role: \"user\", content: `Summarise this:\\n${text}` }],\n    });\n    return Response.json({ summary: r.choices[0].message.content });\n  } catch {\n    // the user sees this with a 200, and nothing is logged\n    return Response.json({ summary: \"Summary unavailable right now.\" });\n  }\n}\n```\n\nYou don't need a provider outage to see it fail. The SDK reads `OPENAI_BASE_URL`, so you can point it at a server that accepts connections and never replies:\n\n``` js\n# terminal 1: a provider that never answers\nnode -e \"require('http').createServer(() => {}).listen(4010)\"\n\n# terminal 2\nOPENAI_BASE_URL=http://127.0.0.1:4010/v1 OPENAI_API_KEY=test npm run dev\n\n# terminal 3\ncurl -s -X POST localhost:3000/api/summary \\\n  -H 'content-type: application/json' -d '{\"text\":\"hello\"}' -w '\\n%{time_total}s\\n'\n```\n\ncurl waits for at least 15 minutes. On Vercel, the same request is killed at 300 seconds with a 504 `FUNCTION_INVOCATION_TIMEOUT` before the `catch` runs, so nothing is logged. Behind Cloudflare the user gets a 524 even sooner.\n\nNow make the stub answer 500 instead: `node -e \"require('http').createServer((q, s) => { s.statusCode = 500; s.end('{}') }).listen(4010)\"`. After a few seconds of retries the route returns 200 with the fallback sentence. Every monitor you have counts that as a success.\n\nEvery layer between the browser and the model has its own clock. In the default setup the clock in your own code is the longest, so an outer layer always runs out first.\n\n| Layer | Default limit | What happens | \n|---|---|---|\n| Cloudflare proxy | 125 s with no response from the origin | 524 to the user | \n| Vercel function (fluid compute) | `maxDuration` 300 s | function terminated, 504 | \n| Node's `fetch` (undici) | 300 s waiting for response headers | connection error | \n| OpenAI and Anthropic Node SDKs | 10 minutes per attempt | `APIConnectionTimeoutError` | \n\nHere is what the official docs say about each part:\n\n`max_tokens`, Anthropic's SDK raises its default timeout above 10 minutes, up to 60.`maxDuration`. Your `catch` never runs, and neither does the logging inside it.`project_spend_limit_exceeded`. Because the SDKs retry every 429, each user's request retries an error that won't clear until the limit resets. Anthropic returns 400 `invalid_request_error` when a workspace spend limit is reached, and retry logic that only looks for 429 won't recognise it.\nThis suite starts a fake provider on localhost with three behaviours: it never answers, it plays an agent that never finishes, or it returns a spend-limit 429. It imports the wrapper from the next section, and its assertions define what that wrapper has to do. It needs no API key and costs nothing to run.\n\n``` python\n// lib/ai-guards.test.ts\nimport http from \"node:http\";\nimport type { AddressInfo } from \"node:net\";\nimport { afterAll, beforeAll, expect, test } from \"vitest\";\nimport { callModel, makeClient, newRun, runAgent } from \"./ai\";\n\nprocess.env.OPENAI_API_KEY ??= \"test\";\nlet hits = 0;\n\n// One fake provider; the URL prefix picks its behaviour.\nconst server = http.createServer((req, res) => {\n  hits++;\n  if (req.url?.startsWith(\"/stall/\")) return; // accept, never answer\n  res.setHeader(\"content-type\", \"application/json\");\n  if (req.url?.startsWith(\"/spent/\")) {\n    res.statusCode = 429;\n    return res.end(JSON.stringify({ error: { code: \"project_spend_limit_exceeded\", message: \"limit reached\" } }));\n  }\n  res.end(JSON.stringify({ // \"/loop/\": a model that never says it is done\n    id: \"c1\", object: \"chat.completion\", created: 0, model: \"stub\",\n    choices: [{ index: 0, finish_reason: \"stop\", message: { role: \"assistant\", content: \"keep going\" } }],\n    usage: { prompt_tokens: 500, completion_tokens: 500, total_tokens: 1000 },\n  }));\n});\nconst client = (mode: string) =>\n  makeClient(`http://127.0.0.1:${(server.address() as AddressInfo).port}/${mode}/v1`);\nconst hi = [{ role: \"user\" as const, content: \"hi\" }];\n\nbeforeAll(() => new Promise<void>((done) => { server.listen(0, \"127.0.0.1\", done); }));\nafterAll(() => { server.closeAllConnections(); server.close(); });\n\ntest(\"a provider that never answers fails inside the deadline\", async () => {\n  const t0 = Date.now();\n  await expect(callModel(client(\"stall\"), newRun(), hi, 2_000)).rejects.toThrow();\n  expect(Date.now() - t0).toBeLessThan(3_000);\n});\n\ntest(\"an agent that never finishes stops at its step cap\", async () => {\n  const run = newRun(5, 1);\n  await expect(runAgent(client(\"loop\"), run, \"plan my week\")).rejects.toThrow(\"run_budget_exceeded\");\n  expect(run.steps).toBe(5);\n});\n\ntest(\"a spend-limit 429 surfaces its code after at most one retry\", async () => {\n  hits = 0;\n  await expect(callModel(client(\"spent\"), newRun(), hi)).rejects.toMatchObject({\n    status: 429, code: \"project_spend_limit_exceeded\",\n  });\n  expect(hits).toBeLessThanOrEqual(2);\n});\nnpm i -D vitest\nnpx vitest run lib/ai-guards\n```\n\nEach test takes a few seconds. Run the suite in CI next to your other tests.\n\nGive every call its own limits, shorter than every layer outside it, and make failures visible.\n\n``` python\n// lib/ai.ts\nimport OpenAI from \"openai\";\n\nexport type Run = { steps: number; maxSteps: number; spentUsd: number; capUsd: number };\nexport const newRun = (maxSteps = 1, capUsd = 0.05): Run => ({ steps: 0, maxSteps, spentUsd: 0, capUsd });\n\n// 20 s per attempt, one retry. The deadline in callModel caps the total.\nexport const makeClient = (baseURL?: string) => new OpenAI({ baseURL, timeout: 20_000, maxRetries: 1 });\n\n// Set these from your model's price page; at 0 the spend cap never trips.\nconst USD_PER_TOKEN = { input: 0, output: 0 };\n\nexport async function callModel(\n  client: OpenAI,\n  run: Run,\n  messages: OpenAI.Chat.ChatCompletionMessageParam[],\n  deadlineMs = 45_000,\n) {\n  if (run.steps >= run.maxSteps || run.spentUsd >= run.capUsd) throw new Error(\"run_budget_exceeded\");\n  run.steps++;\n  const res = await client.chat.completions.create(\n    { model: \"gpt-5-mini\", messages, max_completion_tokens: 800 },\n    { signal: AbortSignal.timeout(deadlineMs) }, // one clock for the whole call, retries included\n  );\n  const u = res.usage;\n  run.spentUsd += (u?.prompt_tokens ?? 0) * USD_PER_TOKEN.input + (u?.completion_tokens ?? 0) * USD_PER_TOKEN.output;\n  return res;\n}\n\nexport async function runAgent(client: OpenAI, run: Run, task: string) {\n  const messages: OpenAI.Chat.ChatCompletionMessageParam[] = [{ role: \"user\", content: task }];\n  for (;;) {\n    const res = await callModel(client, run, messages);\n    const text = res.choices[0].message.content ?? \"\";\n    if (text.includes(\"[done]\")) return text;\n    messages.push({ role: \"assistant\", content: text }, { role: \"user\", content: \"Continue.\" });\n  }\n}\n```\n\nThe route now fails loudly and leaves a record:\n\n``` js\n// app/api/summary/route.ts\nimport { callModel, makeClient, newRun } from \"@/lib/ai\";\n\nconst client = makeClient();\n\nexport async function POST(req: Request) {\n  const { text } = await req.json();\n  const t0 = Date.now();\n  try {\n    const r = await callModel(client, newRun(), [{ role: \"user\", content: `Summarise this:\\n${text}` }], 25_000);\n    console.info(JSON.stringify({ tool: \"summary\", status: \"ok\", ms: Date.now() - t0, usage: r.usage }));\n    return Response.json({ summary: r.choices[0].message.content });\n  } catch (err: any) {\n    console.error(JSON.stringify({ tool: \"summary\", status: \"error\", code: err?.code ?? err?.name, http: err?.status, ms: Date.now() - t0 }));\n    return Response.json({ error: \"summary_failed\" }, { status: 503 });\n  }\n}\n```\n\nPick `deadlineMs` below both your proxy's timeout and your function's `maxDuration`, with time left over to log and respond. If an answer really needs longer, stream it so bytes reach the proxy early, or return a job ID and have the client poll. With streaming, a 200 only means the stream started. Log the call as failed if the stream ends without a `finish_reason`.\n\nFor a per-user daily cap, add up the user's spend for today from the same log rows before creating the run, and pass what's left as `capUsd`. Once those rows are in a table, this query shows the failures you already have. It sorts by failure rate rather than spend, because failed calls barely move a cost chart:\n\n```\nselect tool, count(*) as calls,\n       avg((status <> 'ok')::int) as failure_rate\nfrom ai_calls\nwhere created_at > now() - interval '7 days'\ngroup by tool\norder by failure_rate desc;\n```\n\n`timeout` and `maxRetries` explicitly.`AbortSignal.timeout`), shorter than Gemmein (in beta) handles the per-user half of this for you: its AI tools spend or reserve a person's credits before the provider is called, and every call is recorded with who, tool, tokens, credits and outcome.\n\nWhich layer in your stack has the shortest clock today, and did you set it yourself?", "url": "https://wpnews.pro/news/ai-api-call-hangs-in-production-then-504-timeouts-and-budgets", "canonical_source": "https://dev.to/gemmein/ai-api-call-hangs-in-production-then-504-timeouts-and-budgets-1bel", "published_at": "2026-09-28 23:04:35+00:00", "updated_at": "2026-09-28 23:19:18.096202+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "developer-tools", "mlops"], "entities": ["OpenAI", "Anthropic", "Vercel", "Cloudflare", "Next.js", "Node.js", "undici", "Vitest"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-api-call-hangs-in-production-then-504-timeouts-and-budgets", "markdown": "https://wpnews.pro/news/ai-api-call-hangs-in-production-then-504-timeouts-and-budgets.md", "text": "https://wpnews.pro/news/ai-api-call-hangs-in-production-then-504-timeouts-and-budgets.txt", "jsonld": "https://wpnews.pro/news/ai-api-call-hangs-in-production-then-504-timeouts-and-budgets.jsonld"}}