{"slug": "ai-customer-support-that-degrades-gracefully-instead-of-failing", "title": "AI customer support that degrades gracefully instead of failing.", "summary": "Backend lead and architect Jawad Ul Hadi built Omni.io, a multi-tenant RAG customer-support engine on NestJS, PostgreSQL with row-level security and pgvector, BullMQ, Gemini and React, designed so that every question still gets an answer when the AI layer fails. The system uses a three-tier fallback ladder — cited AI answer, then verbatim excerpts, then FAQ or human hand-off — plus a stored decision trace for each answer, and isolates tenants with Postgres row-level security under a NOBYPASSRLS role tested against a real database.", "body_md": "**AI customer support that degrades gracefully instead of failing.** Multi-tenant RAG on NestJS, PostgreSQL row-level security + pgvector, BullMQ, Gemini and React.\n\nJawad Ul Hadi · Backend Lead & Architect · October 2026 [Let's Connect](https://gravatar.com/juhbukhari)\n\n[Repo](https://github.com/JawadulHadi/omni-io)\n[Case study: *Designing for AI Failure*](https://juh-bukhari.vercel.app/case-study)\n\nMost RAG systems have one mode: working. If the embedding API is down, the model times out, or the model cites a passage it never saw, the customer gets either an error or a confident fabrication.\n\nOmni.io follows three rules:\n\n| Rule | How it is enforced | \n|---|---|\n| **Every question gets an answer** | A three-tier ladder in one service method that never throws: cited AI answer â†’ verbatim excerpts â†’ FAQ or human hand-off | \n| **Every answer explains itself** | A step-by-step decision trace, stored in the audit log and shown to support staff | \n| **Tenants can't see each other's data, even through a bug** | Postgres row-level security under a `NOBYPASSRLS` role, tested against a real database | \n\n```\nflowchart LR\n  subgraph People\n    OP[\"Support team<br/>(owner Â· admin Â· editor Â· viewer)\"]\n    VIS[\"Customer's website visitor<br/>(anonymous)\"]\n    AIA[\"AI assistant<br/>(Claude, Cursor, etc.)\"]\n  end\n\n  subgraph Omni[\"Omni.io\"]\n    SYS[\"Multi-tenant AI support engine\"]\n  end\n\n  subgraph External\n    GEM[\"Google Gemini<br/>chat + embeddings\"]\n    GOO[\"Google OAuth<br/>(optional)\"]\n  end\n\n  OP -- \"Admin console<br/>GraphQL + WebSocket\" --> SYS\n  VIS -- \"Embedded widget<br/>REST /w/:key\" --> SYS\n  AIA -- \"MCP Â· Streamable HTTP<br/>personal access token\" --> SYS\n  SYS -- \"embed Â· generate\" --> GEM\n  SYS -- \"PKCE sign-in\" --> GOO\n```\n\nThere are three ways in, each with its own trust level:\n\n| Channel | Caller | Authentication | What it can see | \n|---|---|---|---|\n| Console | Workspace members | 15-minute JWT + rotating refresh cookie | Everything their role allows | \n| Widget | Anonymous visitors | Public, rotatable widget key | Only documents and FAQs marked **public** | \n| MCP | AI assistants acting for a user | `omni_pat_â€¦` personal access token | Exactly what that user can see | \n\n```\nflowchart TB\n  subgraph Edge[\"Edge Â· Caddy (auto-HTTPS)\"]\n    CAD[\"Static console + widget.js<br/>reverse proxy\"]\n  end\n\n  subgraph App[\"Application tier\"]\n    API[\"NestJS API<br/>GraphQL Â· REST Â· MCP Â· WebSocket<br/>(stateless, horizontally scalable)\"]\n    WRK[\"Ingestion worker<br/>separate Node process<br/>chunk Â· embed Â· retention sweep\"]\n  end\n\n  subgraph Data[\"Data tier\"]\n    PG[(\"PostgreSQL 16 + pgvector 0.8<br/>RLS Â· HNSW Â· SECURITY DEFINER fns\")]\n    RD[(\"Redis 7<br/>BullMQ Â· rate limits Â· budgets Â· progress\")]\n  end\n\n  GEM[\"Gemini API\"]\n\n  CAD -- \"/graphql /auth/* /mcp /w/:key/* /health\" --> API\n  API -- \"enqueue job (ids only)\" --> RD\n  RD -- \"jobs\" --> WRK\n  WRK -. \"progress events\" .-> RD\n  RD -. \"QueueEvents â†’ GraphQL subscription\" .-> API\n  API -- \"omniio_app role\" --> PG\n  WRK -- \"omniio_app role\" --> PG\n  API -- \"query embedding + Tier 1 generation<br/>(timeout + circuit breaker)\" --> GEM\n  WRK -- \"batch embeddings\" --> GEM\n```\n\n**Why ingestion is queued but answering is synchronous.** The split follows failure semantics, not threading:\n\n- **Ingestion** is long and retryable, and nobody is waiting on it. BullMQ gives durability, retries with backoff, and backpressure. A separate process means a 20 MB PDF can't starve the API of CPU, memory or DB connections.\n- **Answering** has a customer waiting. It runs inline under a hard time budget with circuit breakers. The worst case is a fast Tier 2 answer, never a queue.\n\n```\nflowchart LR\n  subgraph Entry[\"Entry points\"]\n    GQL[\"GraphQL resolvers\"]\n    REST[\"REST controllers\"]\n    MCPC[\"MCP controller\"]\n  end\n\n  subgraph Features\n    AUTH[\"AuthModule<br/>login Â· register Â· refresh Â· Google\"]\n    TOK[\"ApiTokensModule<br/>MCP PATs\"]\n    WS[\"WorkspacesModule<br/>members Â· roles Â· invites Â· ladder settings\"]\n    DOC[\"DocumentsModule<br/>upload Â· paste Â· visibility Â· erasure\"]\n    ING[\"IngestionModule<br/>producer Â· progress relay\"]\n    ANS[\"AnswerModule<br/>resilience ladder\"]\n    FAQ[\"FaqModule<br/>Tier 3 floor\"]\n    WID[\"WidgetModule<br/>public key Â· theme Â· rotation\"]\n    AUD[\"AuditModule<br/>answer.completed listener\"]\n    MCP[\"McpModule<br/>4 tools\"]\n    HLT[\"HealthModule<br/>/health Â· systemInfo\"]\n  end\n\n  subgraph Infra[\"Shared infrastructure\"]\n    DB[\"DbService<br/>tenant() Â· withWorkspace() Â· global()\"]\n    AI[\"AiProvider<br/>Gemini | Fake\"]\n    RL[\"Redis rate limiter<br/>+ in-process fallback\"]\n    BS[\"BlobStorage<br/>none | local\"]\n  end\n\n  GQL --> AUTH & TOK & WS & DOC & ANS & FAQ & WID & AUD\n  REST --> AUTH & DOC & WID & HLT\n  MCPC --> MCP\n  MCP --> WS & DOC & FAQ & ANS\n  WID --> ANS\n  ANS -- \"emits answer.completed\" --> AUD\n  DOC --> ING\n  ANS --> AI\n  ANS & DOC & FAQ & WS & AUD & AUTH --> DB\n  WID & AUTH & ANS --> RL\n  DOC --> BS\n```\n\n| Module | Main operations | \n|---|---|\n| `AuthModule` | `POST /auth/login Â· /register Â· /refresh Â· /switch-workspace Â· /logout Â· /google` | \n| `ApiTokensModule` | `apiTokens` ,`createApiToken` ,`revokeApiToken` | \n| `WorkspacesModule` | `workspace` ,`myWorkspaces` ,`createInvitation` ,`acceptInvitation` ,`updateMemberRole` ,`updateLadderSettings` | \n| `DocumentsModule` | `POST /documents/upload` ,`documents` ,`createDocumentFromText` ,`deleteDocument` | \n| `IngestionModule` | BullMQ producer, `ingestionProgress` subscription | \n| `AnswerModule` | `askQuestion` | \n| `FaqModule` | FAQ CRUD | \n| `WidgetModule` | `GET /w/:key/config` ,`POST /w/:key/ask` ,`widgetConfig` ,`rotateWidgetKey` | \n| `AuditModule` | Internal listener; read-only `answers` | \n| `McpModule` | `POST /mcp` :`list_workspaces` ,`list_documents` ,`list_faqs` ,`ask_question` | \n\nEvery request passes the same gates in a fixed order. No feature code can skip them.\n\n``` php\nflowchart LR\n  R[\"HTTP / WS request\"] --> TC[\"TenantContext middleware<br/>opens AsyncLocalStorage scope\"]\n  TC --> CP[\"cookie-parser Â· JSON body 2 MB\"]\n  CP --> AG{\"AuthGuard<br/>@Public?\"}\n  AG -- \"public route\" --> RG\n  AG -- \"JWT valid<br/>(PAT only on /mcp)\" --> FILL[\"fill context:<br/>userId Â· workspaceId\"]\n  AG -- \"invalid\" --> X401[\"401\"]\n  FILL --> RG{\"RolesGuard<br/>@Roles(min) vs<br/>LIVE membership row\"}\n  RG -- \"insufficient\" --> X403[\"403\"]\n  RG -- \"ok\" --> VP[\"ValidationPipe<br/>whitelist Â· forbidNonWhitelisted\"]\n  VP --> H[\"Resolver / controller\"]\n  H --> DBS[\"DbService.tenant()<br/>BEGIN Â· set_config(app.workspace_id, local) Â· queries Â· COMMIT\"]\n  H -. \"any error\" .-> EF[\"AllExceptionsFilter<br/>sanitized, no stack traces\"]\n```\n\nRoles are checked against the **live** `workspace_members` row on every request, not against JWT claims. A demotion or removal takes effect on the next request, not when the token expires.\n\nOne method, `AnswerService.askQuestion()`, never throws. Every branch appends to the decision trace.\n\n``` php\nflowchart TD\n  Q([\"Question â‰¤ 1,000 chars\"]) --> RB{\"Retrieval breaker open?\"}\n  RB -- \"yes\" --> T3\n  RB -- \"no\" --> E[\"Embed query + match_chunks<br/>(RETRIEVAL_TIMEOUT_MS = 4 s)\"]\n  E -- \"error / timeout\" --> T3\n  E --> F{\"Any chunk â‰¥ workspace<br/>similarity floor?\"}\n  F -- \"no Â· no_relevant_context\" --> T3\n  F -- \"yes\" --> GB{\"Generation breaker open?\"}\n  GB -- \"yes Â· tier1_circuit_open\" --> T2\n  GB -- \"no\" --> BUD{\"Daily model-call<br/>budget left?\"}\n  BUD -- \"no Â· tier1_budget_exhausted\" --> T2\n  BUD -- \"yes\" --> G[\"Gemini Â· JSON schema output<br/>(TIER1_TIMEOUT_MS = 8 s)\"]\n  G -- \"error Â· tier1_model_error<br/>timeout Â· tier1_timeout\" --> T2\n  G --> V1{\"Valid JSON?\"}\n  V1 -- \"no Â· tier1_invalid_output\" --> T2\n  V1 -- \"yes\" --> V2{\"confidence â‰¥ threshold?\"}\n  V2 -- \"no Â· tier1_low_confidence\" --> T2\n  V2 -- \"yes\" --> V3{\"cites â‰¥ 1 passage, all<br/>ids âŠ† retrieved?\"}\n  V3 -- \"no Â· tier1_no_citations /<br/>tier1_invalid_citation\" --> T2\n  V3 -- \"yes\" --> V4{\"grounding â‰¥ TIER1_MIN_GROUNDING?<br/>(answer words found in citations)\"}\n  V4 -- \"no Â· tier1_ungrounded\" --> T2\n  V4 -- \"yes\" --> T1([\"Tier 1 Â· cited AI answer\"])\n  T2([\"Tier 2 Â· top â‰¤ 3 excerpts verbatim\"])\n  T3{\"FAQ whole-word<br/>keyword match?\"} -- \"yes\" --> F3([\"Tier 3 Â· FAQ answer\"])\n  T3 -- \"no / DB error\" --> H([\"Tier 3 Â· human hand-off message\"])\n\n  T1 & T2 & F3 & H --> EV[\"emit answer.completed\"]\n  EV -. \"audit write fails â†’ logged only\" .-> AUD[(\"answers table\")]\n```\n\n**What \"confident enough\" means.** Tier 1 needs four independent checks to pass:\n\n1. **Retrieval:** at least one passage clears the per-workspace similarity floor (default 0.6).\n2. **Structure:** the output is schema-valid JSON`{ answer, citedChunkIds, confidence }` .\n3. **Provenance:** every cited id was actually sent to the model.\n4. **Calibration + grounding:** self-reported confidence clears the threshold,*and* enough of the answer's content words appear in the cited text.\n\nSelf-reported confidence is poorly calibrated, so it is never the only gate.\n\n**Cost honesty.** Tokens spent on a Tier 1 attempt that was later rejected are still recorded, so the audit log reflects real spend.\n\nThere is one breaker per dependency, so an embedding outage and a chat-model outage trip independently.\n\n``` php\nstateDiagram-v2\n  [*] --> Closed\n  Closed --> Open: 5 consecutive failures\n  Open --> HalfOpen: 30 s cool-down elapsed\n  HalfOpen --> Closed: the single probe succeeds\n  HalfOpen --> Open: the probe fails\n  note right of HalfOpen\n    Exactly one probe request is let through,\n    concurrent requests are short-circuited.\n  end note\nsequenceDiagram\n  autonumber\n  actor U as Support agent\n  participant C as Console (urql)\n  participant A as NestJS API\n  participant P as Postgres (RLS)\n  participant G as Gemini\n  participant L as AuditListener\n\n  U->>C: types question in Playground\n  C->>A: mutation askQuestion(query)\n  A->>A: AuthGuard â†’ RolesGuard(viewer) â†’ AskLimiter\n  A->>P: tenant tx: read ladder settings\n  A->>G: embed(query) [4 s budget]\n  G-->>A: 768-dim vector\n  A->>P: tenant tx: match_chunks(ws, vec, k, public_only=false)\n  P-->>A: top-k chunks + similarity\n  Note over A: no DB transaction is held during the model call\n  A->>G: generate(JSON schema, delimited passages) [8 s budget]\n  G-->>A: { answer, citedChunkIds, confidence }\n  A->>A: validate JSON Â· citations âŠ† retrieved Â· confidence Â· grounding\n  A-->>C: AnswerOutcome { tier, answer, citations, trace }\n  A--)L: event answer.completed\n  L->>P: tenant tx: insert into answers (tokens, ids, trace, latency)\n  C-->>U: tier badge + citations + decision trace\nsequenceDiagram\n  autonumber\n  actor E as Editor\n  participant C as Console\n  participant A as API\n  participant X as PDF worker thread\n  participant Q as Redis / BullMQ\n  participant W as Ingestion worker\n  participant G as Gemini\n  participant P as Postgres\n\n  E->>C: drop file (PDF / TXT / MD â‰¤ 20 MB)\n  C->>A: POST /documents/upload\n  A->>A: ingest rate limit (per workspace / hour)\n  opt PDF\n    A->>X: parse (30 s limit Â· 512 MB heap Â· max 2 concurrent)\n    X-->>A: extracted text\n  end\n  A->>P: insert document (status = pending)\n  A->>Q: add job { documentId, workspaceId } â€” ids only\n  alt Redis down\n    A->>P: status = failed (Retry button in console)\n  end\n  Q->>W: deliver job\n  W->>P: withWorkspace: load text\n  W->>W: chunk 1,200 chars / 200 overlap, word-aligned\n  loop batches\n    W->>G: embed(batch)\n    W--)Q: progress %\n    Q--)A: QueueEvents\n    A--)C: subscription ingestionProgress\n  end\n  W->>P: upsert chunks by \"documentId:chunkIndex\" + delete stale tail\n  W->>P: status = ready\n  Note over W: failures retry with backoff, UnrecoverableError â†’ failed\n```\n\n**Idempotency.** Chunk ids are deterministic (`${documentId}:${chunkIndex}`), so a retried job overwrites rather than duplicates. Stale tail chunks from a shorter re-ingest are deleted in the same transaction.\n\n**Retention sweep.** Every six hours the worker calls `run_retention()`. It deletes answer-audit rows older than `ANSWER_RETENTION_DAYS` (default 90), plus expired invitations, refresh tokens and API tokens.\n\n```\nerDiagram\n  USERS ||--o{ WORKSPACE_MEMBERS : \"belongs to\"\n  WORKSPACES ||--o{ WORKSPACE_MEMBERS : has\n  WORKSPACES ||--o{ DOCUMENTS : owns\n  DOCUMENTS ||--o{ CHUNKS : \"split into\"\n  WORKSPACES ||--o{ FAQS : owns\n  WORKSPACES ||--o{ ANSWERS : audits\n  WORKSPACES ||--o| WIDGET_CONFIGS : themes\n  WORKSPACES ||--o{ WORKSPACE_INVITATIONS : issues\n  USERS ||--o{ REFRESH_TOKENS : \"sessions\"\n  USERS ||--o{ API_TOKENS : \"MCP PATs\"\n\n  WORKSPACES {\n    uuid id PK\n    text name\n    text plan\n    text widget_key \"single source of truth\"\n    float confidence_threshold\n    float similarity_floor\n  }\n  WORKSPACE_MEMBERS {\n    uuid workspace_id FK\n    uuid user_id FK\n    text role \"owner|admin|editor|viewer\"\n  }\n  USERS {\n    uuid id PK\n    text email\n    text password_hash \"scrypt, not readable by app role\"\n    text google_sub\n  }\n  DOCUMENTS {\n    uuid id PK\n    uuid workspace_id FK\n    text status \"pending|processing|ready|failed\"\n    text source_type\n    text visibility \"internal|public\"\n    text content\n    text storage_key\n  }\n  CHUNKS {\n    text id PK \"documentId:chunkIndex\"\n    uuid workspace_id FK\n    uuid document_id FK\n    vector embedding \"768 dims, HNSW\"\n    text embedding_model\n    text content\n  }\n  FAQS {\n    uuid id PK\n    uuid workspace_id FK\n    text question\n    text answer\n    text_array keywords \"normalized\"\n    text visibility \"internal|public\"\n  }\n  ANSWERS {\n    uuid id PK\n    uuid workspace_id FK\n    text tier\n    text channel \"console|widget|mcp\"\n    text model\n    int tokens_in\n    int tokens_out\n    text_array retrieved_chunk_ids\n    text_array cited_chunk_ids\n    float confidence\n    float top_similarity\n    text decision_note\n    jsonb decision_trace\n    int latency_ms\n  }\n  WIDGET_CONFIGS {\n    uuid workspace_id FK\n    jsonb theme\n    timestamptz rotated_at\n  }\n  WORKSPACE_INVITATIONS {\n    uuid id PK\n    uuid workspace_id FK\n    text token_hash \"SHA-256\"\n    text role\n    timestamptz expires_at \"7 days, single use\"\n  }\n  REFRESH_TOKENS {\n    text token_hash \"SHA-256\"\n    uuid family_id\n    timestamptz used_at\n    timestamptz revoked_at\n  }\n  API_TOKENS {\n    text token_hash \"SHA-256\"\n    text name\n    timestamptz last_used_at\n    timestamptz expires_at\n    timestamptz revoked_at\n  }\n```\n\n**Vector search.** `match_chunks(workspace_id, query_embedding, top_k, public_only)` runs server-side:\n\n- It filters by `workspace_id` inside the SQL, on top of RLS.\n- It uses an HNSW index with `hnsw.iterative_scan = relaxed_order` . IVFFlat would post-filter, so small tenants would get fewer than k rows.\n- If the iterative scan still comes back short, it falls back to an exact search for that tenant only, which is cheap for small tenants.\n\nMigrations are plain, forward-only SQL files (`0001_init` â†’ `0004_hardening`), applied by a ~60-line runner under an advisory lock.\n\nIsolation is enforced in Postgres, in layers, so that no single application bug can cross tenants.\n\n```\nflowchart TB\n  L1[\"â‘  Request context<br/>AuthGuard puts workspaceId in AsyncLocalStorage<br/>(no workspace â†’ DbService throws: fail closed)\"]\n  L2[\"â‘¡ Unit of work<br/>one short transaction Â· set_config('app.workspace_id', id, true)<br/>transaction-local, can't leak to the next pooled request\"]\n  L3[\"â‘¢ Database role<br/>omniio_app: not owner Â· not superuser Â· NOBYPASSRLS<br/>column grants hide users.password_hash\"]\n  L4[\"â‘£ RLS policy on every tenant table<br/>USING / WITH CHECK (workspace_id = nullif(current_setting(...), '')::uuid)\"]\n  L5[\"â‘¤ Narrow SECURITY DEFINER functions<br/>auth_find_user Â· user_workspaces Â· workspace_role Â· resolve_widget_key<br/>invitation_preview Â· accept_invitation Â· run_retention (pinned search_path)\"]\n  L6[\"â‘¥ Proof<br/>e2e tests as omniio_app: no-WHERE selects, cross-tenant insert/update/delete,<br/>vector search aimed at another tenant\"]\n  L1 --> L2 --> L3 --> L4\n  L4 -.-> L5\n  L4 -.-> L6\n```\n\n**Why `nullif(â€¦, '')`?** After a transaction-local `set_config`, a pooled connection reads the setting back as `''`, not `NULL`, and `''::uuid` raises an error on every later unscoped query.\n\n**Why no transaction across model calls?** A slow LLM would pin pool connections. Each unit of work holds a connection for milliseconds.\n\n```\nsequenceDiagram\n  autonumber\n  participant B as Browser (tab)\n  participant A as API /auth\n  participant P as Postgres\n\n  B->>A: POST /auth/login (email, password)\n  A->>P: auth_find_user() â€” SECURITY DEFINER\n  A->>A: scrypt verify\n  A->>P: insert refresh token (SHA-256, new family)\n  A-->>B: access JWT (15 min, memory only) + httpOnly SameSite=Strict cookie (path /auth)\n  Note over B: access token expires\n  B->>B: Web Lock: one refresh across all tabs\n  B->>A: POST /auth/refresh (cookie)\n  A->>P: mark old token used Â· insert next token (same family)\n  A-->>B: new access JWT + rotated cookie\n  alt an already-used token is presented (theft)\n    A->>P: revoke the whole family\n    A-->>B: 401 â€” every session in that family is signed out\n  end\n```\n\n| Concern | Decision | \n|---|---|\n| Access token | 15-minute JWT with `typ: \"access\"` , held in memory only | \n| Refresh token | Opaque, SHA-256 at rest, rotated on every use, reuse revokes the whole family | \n| Google sign-in | Authorization code + PKCE + state cookie. **Never** auto-linked to an existing password account by email, because sign-up doesn't verify email ownership | \n| Invitations | Single-use links valid for 7 days, bound to possession of the link. With `ALLOW_SIGNUP=false` they are the only way to create an account | \n| Roles | `owner > admin > editor > viewer` . Nobody can grant above their own role, only owners can modify owners, and there is always â‰¥ 1 owner | \n| Production boot guard | The API refuses to start with `COOKIE_SECURE=false` or the example`JWT_SECRET` | \n\n```\nsequenceDiagram\n  autonumber\n  actor V as Visitor\n  participant S as Customer site\n  participant J as widget.js (Shadow DOM)\n  participant A as API /w/:key\n  participant R as Redis\n  participant P as Postgres\n\n  S->>J: <script src=\".../widget.js\" data-omniio-key=KEY>\n  J->>A: GET /w/KEY/config\n  A->>P: resolve_widget_key(KEY)\n  alt key unknown or rotated\n    A-->>J: 404 â†’ widget stays hidden\n  else Postgres down\n    A-->>J: default theme\n  end\n  V->>J: asks a question\n  J->>A: POST /w/KEY/ask (CORS *, no credentials)\n  A->>R: rate limit per IP + per workspace\n  A->>A: ladder with public_only = true Â· no trace returned\n  A-->>J: { tier, answer, citations }\n```\n\n- **Isolation from the host page:** the widget renders in a Shadow DOM, ships as a separate ~72 kB gzipped IIFE bundle, and is ASCII-only so it works on pages without a UTF-8 charset.\n- **Data exposure:** it answers only from public documents and FAQs. Both are internal by default, so Tier 2 and Tier 3 can't leak internal content to anonymous visitors.\n- **Key rotation:** takes effect immediately and doesn't affect console sessions.\n\n```\nsequenceDiagram\n  autonumber\n  participant M as MCP client\n  participant A as POST /mcp\n  participant T as McpToolsService\n  participant P as Postgres\n\n  M->>A: JSON-RPC tools/call + Bearer omni_pat_â€¦\n  A->>A: AuthGuard (PAT accepted only on /mcp) â†’ user\n  A->>T: new stateless McpServer bound to user\n  T->>P: workspace_role(workspaceId, userId)\n  alt not a member\n    T-->>M: isError: \"You are not a member of that workspace.\"\n  else member\n    T->>P: withWorkspace(workspaceId) â€” same RLS scope as the console\n    T-->>M: JSON result\n  end\n```\n\n| Tool | Input | Effect | \n|---|---|---|\n| `list_workspaces` | â€” | Read-only: the caller's memberships and roles | \n| `list_documents` | `workspaceId` | Read-only: documents with ingestion status | \n| `list_faqs` | `workspaceId` | Read-only: Tier 3 FAQ entries | \n| `ask_question` | `workspaceId` ,`query` (â‰¤ 1,000 chars) | Runs the ladder. It calls Gemini, writes an audit row and is rate-limited per user | \n\nEvery tool declares a title, an input schema and all four MCP annotation hints (`readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint`), so clients can warn before a tool with side effects runs.\n\nThe transport is Streamable HTTP in stateless mode (HTTP+SSE is deprecated in the MCP spec), so there are no sessions to store or scale. OAuth 2.1 is on the roadmap.\n\nThe reference deployment is one server running Docker Compose. The same images run unchanged on container platforms.\n\n``` php\nflowchart TB\n  NET((\"Internet\")) -- \":80 / :443 (HTTP/3)\" --> WEB\n\n  subgraph Host[\"Single Linux host Â· docker compose (project: omniio)\"]\n    WEB[\"web Â· Caddy<br/>auto-HTTPS Â· HSTS Â· nosniff<br/>console + widget static files\"]\n    API[\"api Â· node dist/main.js :3000<br/>healthcheck /health\"]\n    WRK[\"worker Â· node dist/worker.js\"]\n    MIG[\"migrate Â· node dist/migrate.js<br/>(runs once, then exits)\"]\n    PG[(\"postgres Â· pgvector/pgvector:pg16<br/>volume pgdata\")]\n    RD[(\"redis:7-alpine Â· AOF<br/>volume redisdata\")]\n  end\n\n  WEB -- \"reverse_proxy\" --> API\n  MIG -- \"owner role (DATABASE_MIGRATOR_URL)\" --> PG\n  API -- \"omniio_app\" --> PG\n  WRK -- \"omniio_app\" --> PG\n  API --> RD\n  WRK --> RD\n  API & WRK -- \"HTTPS\" --> GEM[\"Gemini API\"]\n\n  PG -. \"healthy\" .-> MIG\n  MIG -. \"completed successfully\" .-> API & WRK\n  API -. \"healthy\" .-> WEB\n```\n\n**Start-up order:** Postgres becomes healthy â†’ `migrate` runs as the schema owner and exits 0 â†’ the API and worker start as `omniio_app` â†’ Caddy starts once the API is healthy.\n\n**Scaling path:**\n\n| Component | Scales | Note | \n|---|---|---|\n| API | Horizontally | Stateless apart from in-process circuit breakers | \n| Worker | Horizontally | BullMQ distributes jobs; `concurrency: 2` per process | \n| Postgres | Vertically + read replicas | Managed options need pgvector â‰¥ 0.8 for `hnsw.iterative_scan` | \n| Redis | Managed instance | Holds only queues and counters, no source data | \n| Console / widget | CDN | Same-site with the API, because the refresh cookie is `SameSite=Strict` | \n\n| Dependency fails | Detection | Customer gets | Recorded as | \n|---|---|---|---|\n| Gemini chat model errors | exception | Tier 2 excerpts | `tier1_model_error` | \n| Gemini chat model hangs | 8 s `AbortSignal` | Tier 2 excerpts | `tier1_timeout` | \n| Repeated model failures | breaker open | Tier 2, no model call | `tier1_circuit_open` | \n| Model returns bad JSON | schema parse | Tier 2 | `tier1_invalid_output` | \n| Model invents a citation | ids âŠ„ retrieved | Tier 2 | `tier1_invalid_citation` | \n| Answer not supported by its citations | grounding ratio | Tier 2 | `tier1_ungrounded` | \n| Daily budget spent | Redis counter | Tier 2 | `tier1_budget_exhausted` | \n| Embedding API down or slow | exception / 4 s budget / breaker | Tier 3 FAQ or hand-off | `retrieval_failed` /`retrieval_timeout` /`retrieval_circuit_open` | \n| Nothing relevant indexed | similarity floor | Tier 3 | `no_relevant_context` | \n| Audit write fails | listener catch | Same answer | logged only | \n| Redis down | client error | Same answers; per-process rate limits; uploads saved as `failed` + Retry | `/health` â†’`degraded` | \n| Postgres down | client error | Widget: hand-off message + default theme | `/health` â†’ 503 | \n| PDF bomb or slow parse | worker-thread limits | Upload rejected; API stays responsive | document `error` | \n\n| Layer | Stack | \n|---|---|\n| API | NestJS 12, GraphQL (Apollo 5, code-first), REST, MCP SDK, zod, class-validator | \n| Data | PostgreSQL 16, pgvector 0.8 (HNSW + iterative scan), RLS, plain SQL migrations | \n| Jobs | BullMQ 6 on Redis 7, separate worker process | \n| AI | Gemini `gemini-3.5-flash` +`gemini-embedding-2` (768 dims) behind an`AiProvider` interface, plus a deterministic fake with failure injection | \n| Console | React 19, Vite 8, React Router 7, Tailwind v4, shadcn/ui, urql, graphql-ws | \n| Widget | Standalone React IIFE bundle in a Shadow DOM | \n| Delivery | Docker (Node 24 LTS), Caddy, GitHub Actions (typecheck, unit, RLS e2e on real pgvector, image builds), Dependabot | \n\n| ADR | Decision | \n|---|---|\n| [0001](https://github.com/JawadulHadi/omni-io/blob/main/docs/adr/0001-tenant-isolation-in-postgres.md) | Tenant isolation in Postgres, not application code | \n| [0002](https://github.com/JawadulHadi/omni-io/blob/main/docs/adr/0002-short-tenant-transactions.md) | One short tenant transaction per unit of work | \n| [0003](https://github.com/JawadulHadi/omni-io/blob/main/docs/adr/0003-plain-sql-migrations.md) | Plain SQL migrations instead of an ORM | \n| [0004](https://github.com/JawadulHadi/omni-io/blob/main/docs/adr/0004-sync-answers-queued-ingestion.md) | Synchronous answers, queued ingestion | \n| [0005](https://github.com/JawadulHadi/omni-io/blob/main/docs/adr/0005-model-output-is-untrusted.md) | Model output is untrusted input | \n| [0006](https://github.com/JawadulHadi/omni-io/blob/main/docs/adr/0006-rotating-refresh-tokens.md) | Rotating refresh tokens with reuse detection | \n| [0007](https://github.com/JawadulHadi/omni-io/blob/main/docs/adr/0007-urql-and-a-separate-widget-bundle.md) | urql and a separate widget bundle | \n| [0008](https://github.com/JawadulHadi/omni-io/blob/main/docs/adr/0008-mcp-streamable-http.md) | MCP over Streamable HTTP, stateless, under the caller's scope | \n\n## *Source code, ADRs and full documentation:\n[Source](https://github.com/JawadulHadi/omni-io) Â· [Case study: *Designing for AI Failure*](https://juh-bukhari.vercel.app/case-study)\n\n*Designing for AI Failure*", "url": "https://wpnews.pro/news/ai-customer-support-that-degrades-gracefully-instead-of-failing", "canonical_source": "https://gist.github.com/JawadulHadi/200e92104392d9b91dafab467df4a1a7", "published_at": "2026-10-02 17:29:05+00:00", "updated_at": "2026-10-02 18:06:40.123094+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "large-language-models", "mlops", "agent-protocols"], "entities": ["Omni.io", "Jawad Ul Hadi", "NestJS", "PostgreSQL", "pgvector", "BullMQ", "Google Gemini", "React"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-customer-support-that-degrades-gracefully-instead-of-failing", "markdown": "https://wpnews.pro/news/ai-customer-support-that-degrades-gracefully-instead-of-failing.md", "text": "https://wpnews.pro/news/ai-customer-support-that-degrades-gracefully-instead-of-failing.txt", "jsonld": "https://wpnews.pro/news/ai-customer-support-that-degrades-gracefully-instead-of-failing.jsonld"}}