{"slug": "my-google-ai-api-500-errors-stopped-being-scary-when-i-stopped-retrying-the", "title": "My Google AI API 500 errors stopped being scary when I stopped retrying the whole workflow", "summary": "A developer found that intermittent 500, 502, 503, and 429 errors from Google's Gemini and Vertex AI APIs were being amplified by retrying entire automation workflows rather than just the model call, causing duplicate side effects such as repeated CRM writes and Slack alerts. The recommended fix is to move the retry boundary to wrap only the model request, isolate the LLM call in its own sub-workflow or worker, and pair retries with idempotency keys tied to the business event rather than the execution attempt.", "body_md": "I knew something was wrong when one flaky Gemini call turned into:\n\nAt first it looked like a normal \"Google AI API is throwing random 500s\" problem.\n\nIt wasn’t.\n\nThe real bug was that we were retrying the entire automation instead of retrying the model call.\n\nThat distinction matters a lot once your workflow has side effects.\n\nIf you’re seeing intermittent `500`, `502`, `503`, or `429` errors from Gemini or Vertex AI, the fix usually is not “add more retries everywhere.” The fix is to move the retry boundary.\n\nIf you use the Gemini Python SDK, Google already retries transient failures by default.\n\nThat includes transient `429` and `5xx` responses, with exponential backoff.\n\nSo if your production automation is still blowing up after “we added retries,” one of these is probably true:\n\nThat second one is the expensive mistake.\n\nIf your workflow does this:\n\n...then one transient model failure becomes a duplicate generator.\n\nThe model error is annoying.\n\nThe replay damage is worse.\n\nA lot of teams treat `500` as \"Google is broken.\"\n\nSometimes that’s true.\n\nBut with Gemini and Vertex AI, `500` can also mean overload, dependency failures, shared-capacity pressure, or quota-related behavior that doesn’t show up as a clean `429`.\n\nThat’s why these incidents feel spooky in production:\n\nIf you only watch for explicit rate limits, you miss the actual pattern.\n\nYour \"random 500s\" may really be burst traffic, project-wide contention, or spend throttling wearing a different mask.\n\nIf a workflow has side effects, retrying the whole thing is the wrong default.\n\nThe retry boundary should sit around the model call, not around everything before and after it.\n\nThis is the rule I trust now:\n\nIf you don’t do that, you get the usual mess:\n\nAt that point it’s not really an LLM problem anymore.\n\nIt’s a workflow design problem.\n\nThe cleanest version of this is to isolate the LLM call into its own sub-workflow or worker.\n\nInstead of this:\n\nDo this instead:\n\nThat one change removes most of the blast radius.\n\nIf Gemini throws a transient `5xx`, only the model step retries.\n\nYour CRM write doesn’t happen twice.\n\nYour Slack alert doesn’t fire twice.\n\nYour upstream data fetch doesn’t get repeated for no reason.\n\nn8n is actually pretty good at this if you use the primitives it gives you.\n\nUseful pieces:\n\nA decent shape looks like this:\n\n`Error Trigger`\nThat turns a noisy crash into something debuggable.\n\nThis is where a lot of teams accidentally create their own outage.\n\nIf you call Gemini through direct REST, an n8n HTTP Request node, Make, Zapier, or a custom worker, you need to implement retry policy yourself.\n\nMinimum bar:\n\n```\nretry_on = [408, 429, 500, 502, 503, 504]\nmax_attempts = 4\nbase_delay_seconds = 1\nmax_delay_seconds = 60\nuse_jitter = True\n```\n\nAnd the retry should wrap only the model request.\n\nNot the whole business process.\n\n``` python\nimport random\nimport time\nimport requests\n\nRETRY_ON = {408, 429, 500, 502, 503, 504}\nMAX_ATTEMPTS = 4\nBASE_DELAY = 1\nMAX_DELAY = 60\n\ndef call_gemini_with_retry(url, headers, payload):\n    attempt = 0\n\n    while attempt < MAX_ATTEMPTS:\n        attempt += 1\n        response = requests.post(url, headers=headers, json=payload, timeout=60)\n\n        if response.status_code < 400:\n            return response.json()\n\n        if response.status_code not in RETRY_ON:\n            response.raise_for_status()\n\n        if attempt == MAX_ATTEMPTS:\n            response.raise_for_status()\n\n        delay = min(BASE_DELAY * (2 ** (attempt - 1)), MAX_DELAY)\n        jitter = random.uniform(0, delay * 0.25)\n        time.sleep(delay + jitter)\n```\n\nThat’s still not enough by itself.\n\nYou also need idempotency around whatever happens after the model returns.\n\nThis is the design I’d recommend to anyone running AI automations in production.\n\nGive each model request a stable operation ID tied to the business event.\n\nNot the execution attempt.\n\nFor example:\n\n```\nlead_enrichment:hubspot_contact_12345\nsupport_triage:zendesk_ticket_98765\ninvoice_review:invoice_2026_00412\n```\n\nIf the same job replays, your system should recognize it as the same operation.\n\nStore:\n\nIf you don’t log the exact request shape, replay becomes guesswork.\n\nDo not let a worker spin forever because one model is having a bad hour.\n\nAfter max attempts, route to one of these:\n\nFallback routing is not cheating.\n\nIt’s production engineering.\n\nA lot of “random instability” is really bursty traffic.\n\nIf your cron job wakes up and slams Gemini with a huge batch, shared-capacity systems can get weird fast.\n\nPaced workers beat spiky workers.\n\nQueues beat bursts.\n\nThis is another easy trap.\n\nGemini and Vertex AI limits are not always about a single request or a single API key.\n\nThey can be project-wide.\n\nSo if you have:\n\n...all hitting the same Google project, failures can look random unless you correlate them with project-wide traffic.\n\nThat means you should track at least:\n\nIf you only inspect one failing execution, you’ll miss the real cause.\n\nThis is one of those boring implementation details that decides whether your week stays calm.\n\n| Option | What happens when Gemini gets flaky | \n|---|---|\n| Gemini API via official SDK | Safer defaults. Built-in transient retry behavior. Less custom work. | \n| Gemini API via direct REST or n8n HTTP Request | You own retries, jitter, caps, and safe replay boundaries. Easier to get wrong. | \n| Vertex AI pay-as-you-go | Shared-capacity behavior means burst shape matters a lot. | \n| Vertex AI Provisioned Throughput | Better when you need more consistent service and retries alone aren’t enough. | \n\nMy bias: if you’re doing direct HTTP in production, be honest that you’re taking on reliability work.\n\nThat’s fine.\n\nJust don’t pretend it’s the same as using an SDK with sane defaults.\n\nIf you’re testing Vertex AI auth manually:\n\n```\ngcloud auth print-access-token\n```\n\nIf you want to inspect whether your worker is replaying too aggressively, log attempt counts explicitly:\n\n```\ngrep \"gemini_attempt\" app.log | tail -100\n```\n\nAnd if you aren’t logging operation IDs yet, fix that first.\n\nA lot of debugging pain disappears once you can answer this question quickly:\n\nDid the model fail once, or did our workflow replay the same business event four times?\n\nOnce we stopped retrying the whole workflow, the incidents got much less dramatic.\n\nWe still saw transient model failures.\n\nThat part never fully goes away.\n\nBut the failures became contained:\n\nThat’s a very different operational story.\n\nA lot of teams end up here because per-token pricing makes them afraid to add the reliability layers they actually need.\n\nThey avoid extra retries.\n\nThey avoid fallback models.\n\nThey avoid always-on agents.\n\nThey avoid richer automation because every failure path has a billing consequence.\n\nThat’s exactly the problem Standard Compute is trying to remove.\n\nStandard Compute gives you an OpenAI-compatible API with unlimited AI compute at a flat monthly price, so you can run agents, retries, batching, and automations without token anxiety.\n\nIf you’re building on n8n, Make, Zapier, OpenClaw, or custom workers, that matters more than people admit.\n\nPredictable cost changes architecture decisions.\n\nIt’s a lot easier to build safe retry boundaries and fallback paths when every extra call doesn’t feel like a tiny financial penalty.\n\nIf your Google AI API `500` errors keep showing up in production, stop asking only:\n\nhow do we retry harder?\n\nAsk better questions:\n\nWhen an LLM stops responding, the winning move usually isn’t prompt magic.\n\nIt’s boring architecture:\n\nLess exciting than blaming Gemini.\n\nMuch more effective.", "url": "https://wpnews.pro/news/my-google-ai-api-500-errors-stopped-being-scary-when-i-stopped-retrying-the", "canonical_source": "https://dev.to/lars_winstand/my-google-ai-api-500-errors-stopped-being-scary-when-i-stopped-retrying-the-whole-workflow-ko1", "published_at": "2026-09-28 22:10:33+00:00", "updated_at": "2026-09-28 22:19:18.181614+00:00", "lang": "en", "topics": ["ai-agents", "mlops", "ai-infrastructure", "developer-tools"], "entities": ["Google", "Gemini", "Vertex AI", "n8n", "HubSpot", "Zendesk", "Make", "Zapier"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/my-google-ai-api-500-errors-stopped-being-scary-when-i-stopped-retrying-the", "markdown": "https://wpnews.pro/news/my-google-ai-api-500-errors-stopped-being-scary-when-i-stopped-retrying-the.md", "text": "https://wpnews.pro/news/my-google-ai-api-500-errors-stopped-being-scary-when-i-stopped-retrying-the.txt", "jsonld": "https://wpnews.pro/news/my-google-ai-api-500-errors-stopped-being-scary-when-i-stopped-retrying-the.jsonld"}}