{"slug": "the-day-all-our-ai-agents-stopped-working", "title": "The Day All Our AI Agents Stopped Working 💀", "summary": "A developer at Kanvas recounts how all of the company's AI agents failed simultaneously on September 24th when a usage spike triggered 429 RESOURCE_EXHAUSTED errors from their third-party AI provider, despite the company's own infrastructure, APIs and business logic remaining healthy. The team responded by building an internal routing layer with a fallback strategy that lets a request move across multiple models and multiple providers (Model A → Model B → Provider B → Model C) instead of dying with a single model. The developer's conclusion: highly available AI agents cannot be built on a single model or provider.", "body_md": "[Spanish Version Here](https://dev.to/frederickpeal/el-dia-que-todos-nuestros-agentes-dejaron-de-funcionar-3b2f)\n\n**September 24th.**\n\nKanvas is right in the middle of a presentation.\n\nThe tests from the day before? All successful.\n\nOur PO was happy with the results.\n\nEverything was ready.\n\nI had my protein shake that morning. I was prepared for the presentation.\n\nThen, out of nowhere, I get a call:\n\n— **What’s going on with our agents? None of them are working.**\n\nOkay...\n\nI check.\n\nHe’s right.\n\n**Not a single one is working. 💀**\n\nWhat the hell is going on?\n\nThe codebase hasn’t changed in the last 24 hours.\n\nOur core is still operational.\n\nThe APIs are working.\n\nThe business logic is working.\n\nEverything looks fine.\n\nExcept for our agents.\n\nSo I check our monitoring system and find the problem:\n\n`429 RESOURCE_EXHAUSTED`\n\n💀\n\nOkay.\n\nAfter that unnecessarily dramatic introduction, let me explain what actually happened.\n\nOur AI models experienced a usage spike right in the middle of a presentation.\n\nAnd that forced us to ask ourselves a question we hadn't seriously considered before:\n\n**How do we make AI agents highly available?**\n\nKanvas Agent has become one of our flagship products.\n\nAnd with Kanvas, we’ve always had something very important:\n\n**Control.**\n\nOur core is under our control.\n\nBusiness rules, resources, APIs, infrastructure — all of it goes through systems we manage.\n\nIf something fails, we can investigate it.\n\nWe can scale it.\n\nWe can deploy another instance.\n\nWe can change the infrastructure.\n\nWe have redundancy, autoscaling, health checks, blue/green deployments, and years of knowledge about building highly available systems.\n\nBut now we're living through this new wave of AI.\n\nAnd AI introduced a new problem.\n\nYou can have your infrastructure working perfectly.\n\nYour API can be healthy.\n\nYour database can be healthy.\n\nYour servers can be running without a problem.\n\nEverything can be green.\n\nBut if the provider powering your agents stops responding...\n\n**Your agents are dead.**\n\nWhen part of your infrastructure depends on a third-party AI provider, there's something you simply don't control.\n\nAt some point, all you can do is trust them.\n\nAnd this time...\n\n**our faith wasn't enough. 😂**\n\nFortunately, we managed to save the presentation thanks to our PO, Estrella.\n\nBut the incident left us with a problem we needed to solve.\n\nOur Lead came up with a simple idea:\n\n**If a request to an AI model fails, the request shouldn't die with that model.**\n\nThere should always be another route.\n\nIf Model A isn't available, try Model B.\n\nIf the entire provider is having problems, move to Provider B.\n\nAnd that's when we started implementing something we internally call **Routing**.\n\nMore specifically, one of the main strategies behind our router is **fallback**.\n\nOur agentic framework can now receive multiple models and multiple providers.\n\nA request might start with our primary model.\n\nIf that model fails, the request doesn't die.\n\nThe router tries the next available model.\n\nAnd if the problem affects the entire provider, we can move to another provider completely.\n\nSomething like:\n\n**Model A ❌ → Model B ❌ → Provider B → Model C ✅**\n\nSame mission.\n\nDifferent model.\n\nEvery production problem leaves you with a lesson.\n\nOurs was pretty simple:\n\n**You can't build highly available AI agents while depending on a single model or a single provider.**\n\nToday, we can assign multiple models from the same provider and also configure multiple providers with different models.\n\nAnd that opened another interesting door", "url": "https://wpnews.pro/news/the-day-all-our-ai-agents-stopped-working", "canonical_source": "https://dev.to/frederickpeal/the-day-all-our-ai-agents-stopped-working-1iia", "published_at": "2026-09-30 05:23:24+00:00", "updated_at": "2026-09-30 05:46:45.014829+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "mlops", "ai-tools"], "entities": ["Kanvas", "Kanvas Agent", "Estrella"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-day-all-our-ai-agents-stopped-working", "markdown": "https://wpnews.pro/news/the-day-all-our-ai-agents-stopped-working.md", "text": "https://wpnews.pro/news/the-day-all-our-ai-agents-stopped-working.txt", "jsonld": "https://wpnews.pro/news/the-day-all-our-ai-agents-stopped-working.jsonld"}}