Your prompts are code. Manage them like data — but know what you lose when you do A developer detailed the ChimerAI stack's approach to treating prompts as database-backed data rather than code, using a PromptTemplate model with Mustache-style variables, per-category defaults, and permission-gated editing. The writeup also covers per-request cost tracking via trackApiUsage and a regex-based PII guardrail endpoint set, while flagging tradeoffs: losing git diffs and review history for prompts, and a character-based pre-request budget estimate (chars / 4 × 1.2) that is a soft gate rather than a billing-grade token count. The three things that separate an AI demo from an AI product aren't models. They're: being able to change a prompt without a deploy, knowing what each request cost you, and not leaking a user's credit card into logs. None of these are exciting. All three are in the ChimerAI stack, so here's what they do and — more usefully — the tradeoffs they hide. Instead of prompt strings scattered across .py files, there's a PromptTemplate model with Mustache-style variables: model PromptTemplate { id String @id @default cuid name String @unique category String // "system" | "user" | "rag" | "agent" | "chat" | "custom" content String @db.Text variables String // extracted from {{placeholders}} at save time language String @default "en" version Int @default 1 isDefault Boolean @default false isActive Boolean @default true tags String } The variables array is derived from the content, so the app can validate that a caller supplied every {{context}} / {{query}} before rendering — a missing variable becomes a 400 at the API boundary, not a prompt that silently ships a literal {{context}} to the model. POST /api/prompts { "name": "RAG System Prompt", "category": "rag", "content": "Use {{context}} to answer {{query}}", "isDefault": true } Exactly one default per category, multi-language EN/DE/FR/ES/IT , and the whole thing is behind an manage prompts permission so non-admins can't edit production prompts. The tradeoff nobody tells you about: the moment a prompt lives in the DB instead of git, you lose your diff, your code review, and your "what shipped on Tuesday" answer. Versioning here is an Int counter, not a history — it tells you this is v7 , not what v5 said . If prompts matter to you and for a RAG system they do , either export template changes to your audit log or keep the canonical copy in git and treat the DB as a runtime override. I lean toward the latter and haven't fully committed to either, which is honest. Every chat call runs through trackApiUsage , which records the endpoint, model, token counts, success/failure, and status code. The provider's own token counts are used where available; the streaming path accumulates them from chunks and reports once at the end: if provider id and user id and total prompt tokens or total completion tokens : await provider client.report usage provider id=provider id, user id=user id, model=request.model, prompt tokens=total prompt tokens, completion tokens=total completion tokens, endpoint="/api/chat/stream", This is what lets the credit check in the request path exist at all, and what turns "our AI bill doubled" from a mystery into a query. The model is chosen at runtime and can be swapped per provider without a deploy, so cost is tracked against the model that actually answered, not a config constant. Caveat: the pre-request budget check uses a character-based token estimate chars / 4 × 1.2 , not the real tokenizer. It's a soft gate to stop runaway spend, not a billing system. The recorded usage is the accurate number; the estimate is a heuristic that will occasionally reject a request slightly early. Don't invoice off the estimate. chimerai add guardrails gives you four endpoints. The PII detector is regex-based — which is both its strength fast, no network, deterministic and its ceiling it catches patterns, not meaning : "email": re.compile r'\b A-Za-z0-9. %+- +@ A-Za-z0-9.- +\. A-Z|a-z {2,}\b' , "phone": re.compile r'\b\d{3} -. ?\d{3} -. ?\d{4}\b' , "ssn": re.compile r'\b\d{3}-\d{2}-\d{4}\b' , "credit card": re.compile r'\b\d{4} - ?\d{4} - ?\d{4} - ?\d{4}\b' , Endpoints: | Method | Path | What | |---|---|---| | POST | /api/guardrails/pii/detect | find email/phone/SSN/CC/IP | | POST | /api/guardrails/pii/redact | mask the above | | POST | /api/guardrails/toxicity | keyword score 0–1 + risk level | | POST | /api/guardrails/injection | prompt-injection patterns | Toxicity and injection detection are keyword/pattern-based with a confidence score and a risk classification. That's enough to catch the obvious stuff and to add a cheap first line of defense before a request hits your paid model. Where it stops being sufficient: US phone/SSN/CC formats only a German project will want different patterns — worth noting since this kit is partly German-authored ; "toxicity" by keyword misses rephrased hostility entirely; regex injection detection is a tripwire, not a classifier. If compliance demands real content moderation, put a dedicated model in front and keep these as a fast pre-filter. detect vs redact as separate calls is deliberate — you often want to log what was found even when you forward a masked version. The RAG pipeline in the same stack reads its system prompt from category rag feature 1 , runs the user query through guardrails before embedding feature 3 , and reports token usage after the LLM call feature 2 . Three boring features, one grounded-and-metered request. None of this is the part that demos well. It's the part that lets you sleep once users find your app. Repo: github.com/armbur19-collab/chimerai-app