Dmitry Prozorovsky is a freelance automation developer based in Tashkent, Uzbekistan. He builds n8n-based automation and AI agent systems for small businesses, handling everything from order tracking and pricing to customer support.
Over the past year and a half, one of those projects has grown into something much bigger.
Dmitry has been building and maintaining a multi-agent AI system for an e-commerce client, running inside a VK community of around 14,000 people.
Then the server hosting it went down for 17 hours.
The infrastructure belonged to the client, so Dmitry couldn’t fix the hosting problem himself. But there was another problem he could solve. Every message sent during the outage had gone unanswered, and those conversations couldn’t simply disappear once the server came back online.
Dmitry built a recovery workflow around the UptimeRobot API**.** It waits until the monitoring history shows that the server has genuinely recovered, processes the messages that never received a reply, then switches itself off again.
We talked with Dmitry about building AI systems that people actually depend on, what he learned from a recovery process that initially caused problems of its own, and why production engineering means planning for failures outside your control.
Building AI for a real e-commerce community #
Dmitry has been working commercially in automation development for around a year and a half. He didn’t originally set out to specialize in production AI systems. Instead, one client relationship kept growing.
The client sells motorcycle gear, with a VK community of around 14,000 people. What started as a simpler assistant developed into a multi-agent system that now handles real customer interactions every day.
Under the hood are more than 10 interconnected n8n workflows, with a tool-calling AI agent at the center.
The system can check order statuses, calculate prices using the client’s API, search for products, verify receipts and answer common questions. Customers can even send a photo instead of a product link when they want to find something. When the bot can’t handle a request itself, it can escalate the conversation to a person.
Dmitry also built a small FastAPI and Vue dashboard for monitoring and managing the bot.
Running that system in production changed where many of the difficult problems appeared.
“That’s where most of the real lessons came from. Not the AI part, but keeping something like that alive and correct in production.”
One of those lessons arrived when the server disappeared for most of a day.
What happens to 17 hours of unanswered messages? #
The server hosting the bot went down at roughly 7 AM and didn’t return until around midnight local time.
The problem was on the hosting provider’s side. Dmitry didn’t control the server infrastructure, which meant he couldn’t simply step in and repair it himself.
Meanwhile, people kept messaging the client’s VK community.
With the bot offline, those messages weren’t being processed or answered. Even when the server eventually returned, getting the bot running again wouldn’t deal with everything that had happened while it was gone.
“The actual problem I could solve wasn’t ‘get the server back faster.’ It was ‘once it’s back, make sure nothing that happened while it was down just gets silently lost.'”
Dmitry started building a recovery process around the part of the system he could control. First, he needed to know when the server was genuinely ready for it.
Waiting for a genuine recovery #
The server already had an UptimeRobot monitor, so Dmitry used the UptimeRobot API as the external signal for his recovery workflow.
There was an important reason not to put that responsibility on the server itself.
“A health check running on the same box has an obvious blind spot. If the whole box goes down, the health check goes down with it. It can’t tell you it’s down if it’s dead.”
UptimeRobot monitors the server from outside that infrastructure. Dmitry’s workflow could use its monitoring data to see the transition from down to up without depending on the failed server to report its own condition.
But a single successful check wasn’t enough. A server can briefly come back online and then disappear again, so triggering the backfill immediately could mean starting work before the infrastructure was stable.
Dmitry’s workflow reads the monitor’s recent up and down history through the UptimeRobot API and waits for the recovery signal he wants before starting the backfill.
Dmitry’s n8n workflow checks UptimeRobot for a recent recovery before activating the backfill workflow.
That extra caution came partly from an earlier version of the system that didn’t behave as planned.
When the first fix creates another problem #
Dmitry initially tried keeping the backfill process running continuously in the background.
It worked a little too enthusiastically.
The client had been handling some customer conversations manually. The always-on backfill started reopening and interfering with those conversations.
Dmitry redesigned it as a temporary process that activates after a confirmed recovery and shuts itself down again when it’s finished.
“That’s what pushed me toward the current design: something that only turns on when there’s a confirmed recovery, and turns itself off again afterward.”
Catching up without answering twice #
Once triggered, the backfill still needs to work out what actually requires attention.
It checks messages against the bot’s own state tracking and only processes conversations that never received a reply. Anything already handled manually or by the bot is left alone.
The backfill workflow pulls VK conversations and compares them against the bot’s records to identify missed messages.
Order statuses and prices need slightly different treatment. Those lookups happen when the bot actually answers, instead of relying on a snapshot taken when the original message arrived. The customer therefore gets current information even if their message has been waiting for hours.
When the missed messages have been processed, the recovery workflow’s job is finished.
Designing around what you can’t control #
For Dmitry, the outage came down to knowing where his control ends. “Separate clearly what’s your responsibility from what isn’t, but design as if the parts that aren’t your responsibility will fail anyway.”
Dmitry couldn’t prevent the hosting provider from going down. What he could decide was what happened after it came back.
His advice for engineers building similar tools follows the same principle.
“Build recovery tools to be self-limiting: something that watches for a real signal and then switches itself back off, rather than a patch that just runs forever and can end up causing its own bugs.”
It’s a small workflow next to the larger AI system Dmitry has been building. But it’s also the kind of work that only becomes important once software leaves the demo stage and people start depending on it.
The bot working is one problem. Making sure it knows how to come back from failure is another.
See what Dmitry is building #
Want to see more of Dmitry’s automation and AI projects? Follow him on LinkedIn, where he shares what he’s building and what he learns along the way.