# Building a Production WhatsApp AI Agent: Architecture That Actually Works

> Source: <https://dev.to/alessandrobinda114/building-a-production-whatsapp-ai-agent-architecture-that-actually-works-2gfd>
> Published: 2026-08-09 21:09:19+00:00

Everyone demos a WhatsApp chatbot. Few run one in production with real customers sending real messages 24/7.

After 18 months of running SARA — an open-source WhatsApp AI agent serving businesses across 20 industries — here's what we learned about architecture that survives contact with reality.

The numbers are simple:

But WhatsApp is NOT just another chat channel. It has unique constraints that break naive implementations.

```
WhatsApp (WAHA) → Bridge (:3008) → SARA API (:3006) → AI Provider Chain → Tool Dispatcher
                                                              ↓
                                                    Groq → Cerebras → SambaNova → Mistral
```

Single-provider AI is a production risk. We use a 4-provider chain:

```
Primary: Groq (fastest, free tier)
    ↓ fail
Fallback 1: Cerebras
    ↓ fail
Fallback 2: SambaNova
    ↓ fail
Fallback 3: Mistral (paid, always works)
```

Each provider gets 2 retries with exponential backoff before failover. Result: **99.7% uptime** over 6 months with $0 inference cost (free tiers).

SARA doesn't just answer questions. She executes actions:

`create_reservation`

— books a table with date normalization ("domani alle 8" → 2026-08-10T20:00)`check_inventory`

— queries stock levels`generate_invoice`

— creates a PDF from database records`schedule_appointment`

— manages calendar slotsThe dispatcher maps 30+ tools to handlers with an autonomy gate:

```
User message → Intent classification → Risk assessment → Tool execution
                                              ↓
                                    Low risk: execute immediately
                                    Medium: execute + notify owner
                                    High: ask for confirmation first
```

You do NOT want your AI agent booking a catering order for 500 people without human approval.

Messages contain names, phone numbers, addresses. Our pipeline:

WhatsApp doesn't have "sessions" — it's just a stream of messages. We manage context with:

SARA runs on a single VPS (4 vCPU, 8GB RAM):

| Component | Resource |
|---|---|
| WAHA (WhatsApp Web) | ~500MB RAM |
| Bridge service | ~50MB |
| SARA API | ~200MB |
| PostgreSQL + pgvector | ~2GB |
| Total | ~3GB |

No GPU needed — inference is offloaded to cloud providers (Groq, etc.).

**WhatsApp session contention** — running two instances with the same number = instant logout for both. We learned this the hard way.

**Date parsing across languages** — "dopodomani" (Italian for "day after tomorrow") + timezone handling + business hours awareness. This alone took weeks.

**Message ordering** — WhatsApp doesn't guarantee delivery order. Our bridge queues and re-orders by timestamp.

SARA is AGPL-3.0 on GitHub: [github.com/Alessandro114/sara](https://github.com/Alessandro114/sara)

Self-host it, extend it, build your own vertical agent on top. Cloud-only features (multi-tenant, white-label, analytics) stay in the commercial version.

The 20 industry-specific agent definitions are also open source: [scala-agent-definitions](https://github.com/Alessandro114/scala-agent-definitions) (Apache-2.0).

*Running AI in production is 10% model quality and 90% engineering. Follow for more war stories.*
