cd /news/ai-agents/ai-api-call-hangs-in-production-then… · home › topics › ai-agents › article
[ARTICLE · art-141334] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↓ negative

AI API call hangs in production, then 504: timeouts and budgets

A developer analyzed 1,662 public posts from builders whose apps broke at or after launch and verified 215 recent cases, finding that AI model and API failures account for 4% of them, with six verified cases sharing a pattern of outbound or AI calls lacking timeouts, budgets, and fallbacks. The writeup shows that default SDK timeouts (10 minutes per attempt for OpenAI and Anthropic Node SDKs) exceed the limits of outer layers like Cloudflare's 125-second proxy timeout and Vercel's 300-second function limit, so requests are killed with 504/524 errors before application-level catch blocks or logging can run. The developer also demonstrates a local fake-provider test suite that reproduces never-answering, never-finishing agent, and spend-limit 429 scenarios without an API key.

by read7 min views1 publishedSep 28, 2026

We read 1,662 public posts from builders whose apps broke at or after launch and verified 215 recent cases. AI model and API failures make up 4% of them. Six of the verified cases share one pattern: an outbound or AI call with no timeout, no budget and no fallback. In most of the six, nobody noticed for a while, because the app didn't log model calls or their outcomes.

From the builder's side it shows up as one of these:

All three have the same cause. The call has no limits of its own, so the only limits that apply are set by other layers.

Here is a Next.js App Router route using the OpenAI Node SDK with its defaults. Most generated code ships in this shape.

// app/api/summary/route.ts
import OpenAI from "openai";

const openai = new OpenAI(); // defaults: 10 min per attempt, 2 retries

export async function POST(req: Request) {
  const { text } = await req.json();
  try {
    const r = await openai.chat.completions.create({
      model: "gpt-5-mini",
      messages: [{ role: "user", content: `Summarise this:\n${text}` }],
    });
    return Response.json({ summary: r.choices[0].message.content });
  } catch {
    // the user sees this with a 200, and nothing is logged
    return Response.json({ summary: "Summary unavailable right now." });
  }
}

You don't need a provider outage to see it fail. The SDK reads OPENAI_BASE_URL, so you can point it at a server that accepts connections and never replies:

node -e "require('http').createServer(() => {}).listen(4010)"

OPENAI_BASE_URL=http://127.0.0.1:4010/v1 OPENAI_API_KEY=test npm run dev

curl -s -X POST localhost:3000/api/summary \
  -H 'content-type: application/json' -d '{"text":"hello"}' -w '\n%{time_total}s\n'

curl waits for at least 15 minutes. On Vercel, the same request is killed at 300 seconds with a 504 FUNCTION_INVOCATION_TIMEOUT before the catch runs, so nothing is logged. Behind Cloudflare the user gets a 524 even sooner.

Now make the stub answer 500 instead: node -e "require('http').createServer((q, s) => { s.statusCode = 500; s.end('{}') }).listen(4010)". After a few seconds of retries the route returns 200 with the fallback sentence. Every monitor you have counts that as a success.

Every layer between the browser and the model has its own clock. In the default setup the clock in your own code is the longest, so an outer layer always runs out first.

Layer Default limit What happens
Cloudflare proxy 125 s with no response from the origin 524 to the user
Vercel function (fluid compute) maxDuration 300 s function terminated, 504
Node's fetch (undici) 300 s waiting for response headers connection error
OpenAI and Anthropic Node SDKs 10 minutes per attempt APIConnectionTimeoutError

Here is what the official docs say about each part:

max_tokens, Anthropic's SDK raises its default timeout above 10 minutes, up to 60.maxDuration. Your catch never runs, and neither does the logging inside it.project_spend_limit_exceeded. Because the SDKs retry every 429, each user's request retries an error that won't clear until the limit resets. Anthropic returns 400 invalid_request_error when a workspace spend limit is reached, and retry logic that only looks for 429 won't recognise it. This suite starts a fake provider on localhost with three behaviours: it never answers, it plays an agent that never finishes, or it returns a spend-limit 429. It imports the wrapper from the next section, and its assertions define what that wrapper has to do. It needs no API key and costs nothing to run.

// lib/ai-guards.test.ts
import http from "node:http";
import type { AddressInfo } from "node:net";
import { afterAll, beforeAll, expect, test } from "vitest";
import { callModel, makeClient, newRun, runAgent } from "./ai";

process.env.OPENAI_API_KEY ??= "test";
let hits = 0;

// One fake provider; the URL prefix picks its behaviour.
const server = http.createServer((req, res) => {
  hits++;
  if (req.url?.startsWith("/stall/")) return; // accept, never answer
  res.setHeader("content-type", "application/json");
  if (req.url?.startsWith("/spent/")) {
    res.statusCode = 429;
    return res.end(JSON.stringify({ error: { code: "project_spend_limit_exceeded", message: "limit reached" } }));
  }
  res.end(JSON.stringify({ // "/loop/": a model that never says it is done
    id: "c1", object: "chat.completion", created: 0, model: "stub",
    choices: [{ index: 0, finish_reason: "stop", message: { role: "assistant", content: "keep going" } }],
    usage: { prompt_tokens: 500, completion_tokens: 500, total_tokens: 1000 },
  }));
});
const client = (mode: string) =>
  makeClient(`http://127.0.0.1:${(server.address() as AddressInfo).port}/${mode}/v1`);
const hi = [{ role: "user" as const, content: "hi" }];

beforeAll(() => new Promise<void>((done) => { server.listen(0, "127.0.0.1", done); }));
afterAll(() => { server.closeAllConnections(); server.close(); });

test("a provider that never answers fails inside the deadline", async () => {
  const t0 = Date.now();
  await expect(callModel(client("stall"), newRun(), hi, 2_000)).rejects.toThrow();
  expect(Date.now() - t0).toBeLessThan(3_000);
});

test("an agent that never finishes stops at its step cap", async () => {
  const run = newRun(5, 1);
  await expect(runAgent(client("loop"), run, "plan my week")).rejects.toThrow("run_budget_exceeded");
  expect(run.steps).toBe(5);
});

test("a spend-limit 429 surfaces its code after at most one retry", async () => {
  hits = 0;
  await expect(callModel(client("spent"), newRun(), hi)).rejects.toMatchObject({
    status: 429, code: "project_spend_limit_exceeded",
  });
  expect(hits).toBeLessThanOrEqual(2);
});
npm i -D vitest
npx vitest run lib/ai-guards

Each test takes a few seconds. Run the suite in CI next to your other tests.

Give every call its own limits, shorter than every layer outside it, and make failures visible.

// lib/ai.ts
import OpenAI from "openai";

export type Run = { steps: number; maxSteps: number; spentUsd: number; capUsd: number };
export const newRun = (maxSteps = 1, capUsd = 0.05): Run => ({ steps: 0, maxSteps, spentUsd: 0, capUsd });

// 20 s per attempt, one retry. The deadline in callModel caps the total.
export const makeClient = (baseURL?: string) => new OpenAI({ baseURL, timeout: 20_000, maxRetries: 1 });

// Set these from your model's price page; at 0 the spend cap never trips.
const USD_PER_TOKEN = { input: 0, output: 0 };

export async function callModel(
  client: OpenAI,
  run: Run,
  messages: OpenAI.Chat.ChatCompletionMessageParam[],
  deadlineMs = 45_000,
) {
  if (run.steps >= run.maxSteps || run.spentUsd >= run.capUsd) throw new Error("run_budget_exceeded");
  run.steps++;
  const res = await client.chat.completions.create(
    { model: "gpt-5-mini", messages, max_completion_tokens: 800 },
    { signal: AbortSignal.timeout(deadlineMs) }, // one clock for the whole call, retries included
  );
  const u = res.usage;
  run.spentUsd += (u?.prompt_tokens ?? 0) * USD_PER_TOKEN.input + (u?.completion_tokens ?? 0) * USD_PER_TOKEN.output;
  return res;
}

export async function runAgent(client: OpenAI, run: Run, task: string) {
  const messages: OpenAI.Chat.ChatCompletionMessageParam[] = [{ role: "user", content: task }];
  for (;;) {
    const res = await callModel(client, run, messages);
    const text = res.choices[0].message.content ?? "";
    if (text.includes("[done]")) return text;
    messages.push({ role: "assistant", content: text }, { role: "user", content: "Continue." });
  }
}

The route now fails loudly and leaves a record:

// app/api/summary/route.ts
import { callModel, makeClient, newRun } from "@/lib/ai";

const client = makeClient();

export async function POST(req: Request) {
  const { text } = await req.json();
  const t0 = Date.now();
  try {
    const r = await callModel(client, newRun(), [{ role: "user", content: `Summarise this:\n${text}` }], 25_000);
    console.info(JSON.stringify({ tool: "summary", status: "ok", ms: Date.now() - t0, usage: r.usage }));
    return Response.json({ summary: r.choices[0].message.content });
  } catch (err: any) {
    console.error(JSON.stringify({ tool: "summary", status: "error", code: err?.code ?? err?.name, http: err?.status, ms: Date.now() - t0 }));
    return Response.json({ error: "summary_failed" }, { status: 503 });
  }
}

Pick deadlineMs below both your proxy's timeout and your function's maxDuration, with time left over to log and respond. If an answer really needs longer, stream it so bytes reach the proxy early, or return a job ID and have the client poll. With streaming, a 200 only means the stream started. Log the call as failed if the stream ends without a finish_reason.

For a per-user daily cap, add up the user's spend for today from the same log rows before creating the run, and pass what's left as capUsd. Once those rows are in a table, this query shows the failures you already have. It sorts by failure rate rather than spend, because failed calls barely move a cost chart:

select tool, count(*) as calls,
       avg((status <> 'ok')::int) as failure_rate
from ai_calls
where created_at > now() - interval '7 days'
group by tool
order by failure_rate desc;

timeout and maxRetries explicitly.AbortSignal.timeout), shorter than Gemmein (in beta) handles the per-user half of this for you: its AI tools spend or reserve a person's credits before the provider is called, and every call is recorded with who, tool, tokens, credits and outcome.

Which layer in your stack has the shortest clock today, and did you set it yourself?

── more in #ai-agents 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-api-call-hangs-in…] indexed:0 read:7min 2026-09-28 · —