cd /news/ai-agents/my-google-ai-api-500-errors-stopped-… · home › topics › ai-agents › article
[ARTICLE · art-141318] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

My Google AI API 500 errors stopped being scary when I stopped retrying the whole workflow

A developer found that intermittent 500, 502, 503, and 429 errors from Google's Gemini and Vertex AI APIs were being amplified by retrying entire automation workflows rather than just the model call, causing duplicate side effects such as repeated CRM writes and Slack alerts. The recommended fix is to move the retry boundary to wrap only the model request, isolate the LLM call in its own sub-workflow or worker, and pair retries with idempotency keys tied to the business event rather than the execution attempt.

by read6 min views1 publishedSep 28, 2026

I knew something was wrong when one flaky Gemini call turned into:

At first it looked like a normal "Google AI API is throwing random 500s" problem.

It wasn’t.

The real bug was that we were retrying the entire automation instead of retrying the model call.

That distinction matters a lot once your workflow has side effects.

If you’re seeing intermittent 500, 502, 503, or 429 errors from Gemini or Vertex AI, the fix usually is not “add more retries everywhere.” The fix is to move the retry boundary.

If you use the Gemini Python SDK, Google already retries transient failures by default.

That includes transient 429 and 5xx responses, with exponential backoff.

So if your production automation is still blowing up after “we added retries,” one of these is probably true:

That second one is the expensive mistake.

If your workflow does this:

...then one transient model failure becomes a duplicate generator.

The model error is annoying.

The replay damage is worse.

A lot of teams treat 500 as "Google is broken."

Sometimes that’s true.

But with Gemini and Vertex AI, 500 can also mean overload, dependency failures, shared-capacity pressure, or quota-related behavior that doesn’t show up as a clean 429.

That’s why these incidents feel spooky in production:

If you only watch for explicit rate limits, you miss the actual pattern.

Your "random 500s" may really be burst traffic, project-wide contention, or spend throttling wearing a different mask.

If a workflow has side effects, retrying the whole thing is the wrong default.

The retry boundary should sit around the model call, not around everything before and after it.

This is the rule I trust now:

If you don’t do that, you get the usual mess:

At that point it’s not really an LLM problem anymore.

It’s a workflow design problem.

The cleanest version of this is to isolate the LLM call into its own sub-workflow or worker.

Instead of this:

Do this instead:

That one change removes most of the blast radius.

If Gemini throws a transient 5xx, only the model step retries.

Your CRM write doesn’t happen twice.

Your Slack alert doesn’t fire twice.

Your upstream data fetch doesn’t get repeated for no reason.

n8n is actually pretty good at this if you use the primitives it gives you.

Useful pieces:

A decent shape looks like this:

Error Trigger That turns a noisy crash into something debuggable.

This is where a lot of teams accidentally create their own outage.

If you call Gemini through direct REST, an n8n HTTP Request node, Make, Zapier, or a custom worker, you need to implement retry policy yourself.

Minimum bar:

retry_on = [408, 429, 500, 502, 503, 504]
max_attempts = 4
base_delay_seconds = 1
max_delay_seconds = 60
use_jitter = True

And the retry should wrap only the model request.

Not the whole business process.

import random
import time
import requests

RETRY_ON = {408, 429, 500, 502, 503, 504}
MAX_ATTEMPTS = 4
BASE_DELAY = 1
MAX_DELAY = 60

def call_gemini_with_retry(url, headers, payload):
    attempt = 0

    while attempt < MAX_ATTEMPTS:
        attempt += 1
        response = requests.post(url, headers=headers, json=payload, timeout=60)

        if response.status_code < 400:
            return response.json()

        if response.status_code not in RETRY_ON:
            response.raise_for_status()

        if attempt == MAX_ATTEMPTS:
            response.raise_for_status()

        delay = min(BASE_DELAY * (2 ** (attempt - 1)), MAX_DELAY)
        jitter = random.uniform(0, delay * 0.25)
        time.sleep(delay + jitter)

That’s still not enough by itself.

You also need idempotency around whatever happens after the model returns.

This is the design I’d recommend to anyone running AI automations in production.

Give each model request a stable operation ID tied to the business event.

Not the execution attempt.

For example:

lead_enrichment:hubspot_contact_12345
support_triage:zendesk_ticket_98765
invoice_review:invoice_2026_00412

If the same job replays, your system should recognize it as the same operation.

Store:

If you don’t log the exact request shape, replay becomes guesswork.

Do not let a worker spin forever because one model is having a bad hour.

After max attempts, route to one of these:

Fallback routing is not cheating.

It’s production engineering.

A lot of “random instability” is really bursty traffic.

If your cron job wakes up and slams Gemini with a huge batch, shared-capacity systems can get weird fast.

Paced workers beat spiky workers.

Queues beat bursts.

This is another easy trap.

Gemini and Vertex AI limits are not always about a single request or a single API key.

They can be project-wide.

So if you have:

...all hitting the same Google project, failures can look random unless you correlate them with project-wide traffic.

That means you should track at least:

If you only inspect one failing execution, you’ll miss the real cause.

This is one of those boring implementation details that decides whether your week stays calm.

Option What happens when Gemini gets flaky
Gemini API via official SDK Safer defaults. Built-in transient retry behavior. Less custom work.
Gemini API via direct REST or n8n HTTP Request You own retries, jitter, caps, and safe replay boundaries. Easier to get wrong.
Vertex AI pay-as-you-go Shared-capacity behavior means burst shape matters a lot.
Vertex AI Provisioned Throughput Better when you need more consistent service and retries alone aren’t enough.

My bias: if you’re doing direct HTTP in production, be honest that you’re taking on reliability work.

That’s fine.

Just don’t pretend it’s the same as using an SDK with sane defaults.

If you’re testing Vertex AI auth manually:

gcloud auth print-access-token

If you want to inspect whether your worker is replaying too aggressively, log attempt counts explicitly:

grep "gemini_attempt" app.log | tail -100

And if you aren’t logging operation IDs yet, fix that first.

A lot of debugging pain disappears once you can answer this question quickly:

Did the model fail once, or did our workflow replay the same business event four times?

Once we stopped retrying the whole workflow, the incidents got much less dramatic.

We still saw transient model failures.

That part never fully goes away.

But the failures became contained:

That’s a very different operational story.

A lot of teams end up here because per-token pricing makes them afraid to add the reliability layers they actually need.

They avoid extra retries.

They avoid fallback models.

They avoid always-on agents.

They avoid richer automation because every failure path has a billing consequence.

That’s exactly the problem Standard Compute is trying to remove.

Standard Compute gives you an OpenAI-compatible API with unlimited AI compute at a flat monthly price, so you can run agents, retries, batching, and automations without token anxiety.

If you’re building on n8n, Make, Zapier, OpenClaw, or custom workers, that matters more than people admit.

Predictable cost changes architecture decisions.

It’s a lot easier to build safe retry boundaries and fallback paths when every extra call doesn’t feel like a tiny financial penalty.

If your Google AI API 500 errors keep showing up in production, stop asking only:

how do we retry harder?

Ask better questions:

When an LLM stops responding, the winning move usually isn’t prompt magic.

It’s boring architecture:

Less exciting than blaming Gemini.

Much more effective.

── more in #ai-agents 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/my-google-ai-api-500…] indexed:0 read:6min 2026-09-28 · —