I knew something was wrong when one flaky Gemini call turned into:
At first it looked like a normal "Google AI API is throwing random 500s" problem.
It wasn’t.
The real bug was that we were retrying the entire automation instead of retrying the model call.
That distinction matters a lot once your workflow has side effects.
If you’re seeing intermittent 500, 502, 503, or 429 errors from Gemini or Vertex AI, the fix usually is not “add more retries everywhere.” The fix is to move the retry boundary.
If you use the Gemini Python SDK, Google already retries transient failures by default.
That includes transient 429 and 5xx responses, with exponential backoff.
So if your production automation is still blowing up after “we added retries,” one of these is probably true:
That second one is the expensive mistake.
If your workflow does this:
...then one transient model failure becomes a duplicate generator.
The model error is annoying.
The replay damage is worse.
A lot of teams treat 500 as "Google is broken."
Sometimes that’s true.
But with Gemini and Vertex AI, 500 can also mean overload, dependency failures, shared-capacity pressure, or quota-related behavior that doesn’t show up as a clean 429.
That’s why these incidents feel spooky in production:
If you only watch for explicit rate limits, you miss the actual pattern.
Your "random 500s" may really be burst traffic, project-wide contention, or spend throttling wearing a different mask.
If a workflow has side effects, retrying the whole thing is the wrong default.
The retry boundary should sit around the model call, not around everything before and after it.
This is the rule I trust now:
If you don’t do that, you get the usual mess:
At that point it’s not really an LLM problem anymore.
It’s a workflow design problem.
The cleanest version of this is to isolate the LLM call into its own sub-workflow or worker.
Instead of this:
Do this instead:
That one change removes most of the blast radius.
If Gemini throws a transient 5xx, only the model step retries.
Your CRM write doesn’t happen twice.
Your Slack alert doesn’t fire twice.
Your upstream data fetch doesn’t get repeated for no reason.
n8n is actually pretty good at this if you use the primitives it gives you.
Useful pieces:
A decent shape looks like this:
Error Trigger
That turns a noisy crash into something debuggable.
This is where a lot of teams accidentally create their own outage.
If you call Gemini through direct REST, an n8n HTTP Request node, Make, Zapier, or a custom worker, you need to implement retry policy yourself.
Minimum bar:
retry_on = [408, 429, 500, 502, 503, 504]
max_attempts = 4
base_delay_seconds = 1
max_delay_seconds = 60
use_jitter = True
And the retry should wrap only the model request.
Not the whole business process.
import random
import time
import requests
RETRY_ON = {408, 429, 500, 502, 503, 504}
MAX_ATTEMPTS = 4
BASE_DELAY = 1
MAX_DELAY = 60
def call_gemini_with_retry(url, headers, payload):
attempt = 0
while attempt < MAX_ATTEMPTS:
attempt += 1
response = requests.post(url, headers=headers, json=payload, timeout=60)
if response.status_code < 400:
return response.json()
if response.status_code not in RETRY_ON:
response.raise_for_status()
if attempt == MAX_ATTEMPTS:
response.raise_for_status()
delay = min(BASE_DELAY * (2 ** (attempt - 1)), MAX_DELAY)
jitter = random.uniform(0, delay * 0.25)
time.sleep(delay + jitter)
That’s still not enough by itself.
You also need idempotency around whatever happens after the model returns.
This is the design I’d recommend to anyone running AI automations in production.
Give each model request a stable operation ID tied to the business event.
Not the execution attempt.
For example:
lead_enrichment:hubspot_contact_12345
support_triage:zendesk_ticket_98765
invoice_review:invoice_2026_00412
If the same job replays, your system should recognize it as the same operation.
Store:
If you don’t log the exact request shape, replay becomes guesswork.
Do not let a worker spin forever because one model is having a bad hour.
After max attempts, route to one of these:
Fallback routing is not cheating.
It’s production engineering.
A lot of “random instability” is really bursty traffic.
If your cron job wakes up and slams Gemini with a huge batch, shared-capacity systems can get weird fast.
Paced workers beat spiky workers.
Queues beat bursts.
This is another easy trap.
Gemini and Vertex AI limits are not always about a single request or a single API key.
They can be project-wide.
So if you have:
...all hitting the same Google project, failures can look random unless you correlate them with project-wide traffic.
That means you should track at least:
If you only inspect one failing execution, you’ll miss the real cause.
This is one of those boring implementation details that decides whether your week stays calm.
| Option | What happens when Gemini gets flaky |
|---|---|
| Gemini API via official SDK | Safer defaults. Built-in transient retry behavior. Less custom work. |
| Gemini API via direct REST or n8n HTTP Request | You own retries, jitter, caps, and safe replay boundaries. Easier to get wrong. |
| Vertex AI pay-as-you-go | Shared-capacity behavior means burst shape matters a lot. |
| Vertex AI Provisioned Throughput | Better when you need more consistent service and retries alone aren’t enough. |
My bias: if you’re doing direct HTTP in production, be honest that you’re taking on reliability work.
That’s fine.
Just don’t pretend it’s the same as using an SDK with sane defaults.
If you’re testing Vertex AI auth manually:
gcloud auth print-access-token
If you want to inspect whether your worker is replaying too aggressively, log attempt counts explicitly:
grep "gemini_attempt" app.log | tail -100
And if you aren’t logging operation IDs yet, fix that first.
A lot of debugging pain disappears once you can answer this question quickly:
Did the model fail once, or did our workflow replay the same business event four times?
Once we stopped retrying the whole workflow, the incidents got much less dramatic.
We still saw transient model failures.
That part never fully goes away.
But the failures became contained:
That’s a very different operational story.
A lot of teams end up here because per-token pricing makes them afraid to add the reliability layers they actually need.
They avoid extra retries.
They avoid fallback models.
They avoid always-on agents.
They avoid richer automation because every failure path has a billing consequence.
That’s exactly the problem Standard Compute is trying to remove.
Standard Compute gives you an OpenAI-compatible API with unlimited AI compute at a flat monthly price, so you can run agents, retries, batching, and automations without token anxiety.
If you’re building on n8n, Make, Zapier, OpenClaw, or custom workers, that matters more than people admit.
Predictable cost changes architecture decisions.
It’s a lot easier to build safe retry boundaries and fallback paths when every extra call doesn’t feel like a tiny financial penalty.
If your Google AI API 500 errors keep showing up in production, stop asking only:
how do we retry harder?
Ask better questions:
When an LLM stops responding, the winning move usually isn’t prompt magic.
It’s boring architecture:
Less exciting than blaming Gemini.
Much more effective.