The three things that separate an AI demo from an AI product aren't models. They're: being able
to change a prompt without a deploy, knowing what each request cost you, and not leaking a
user's credit card into logs. None of these are exciting. All three are in the ChimerAI stack,
so here's what they do and — more usefully — the tradeoffs they hide.
Instead of prompt strings scattered across .py files, there's a PromptTemplate model with
Mustache-style variables:
model PromptTemplate {
id String @id @default(cuid())
name String @unique
category String // "system" | "user" | "rag" | "agent" | "chat" | "custom"
content String @db.Text
variables String[] // extracted from {{placeholders}} at save time
language String @default("en")
version Int @default(1)
isDefault Boolean @default(false)
isActive Boolean @default(true)
tags String[]
}
The variables array is derived from the content, so the app can validate that a caller
supplied every {{context}} / {{query}} before rendering — a missing variable becomes a
400 at the API boundary, not a prompt that silently ships a literal {{context}} to the model.
POST /api/prompts
{ "name": "RAG System Prompt",
"category": "rag",
"content": "Use {{context}} to answer {{query}}",
"isDefault": true }
Exactly one default per category, multi-language (EN/DE/FR/ES/IT), and the whole thing is
behind an manage_prompts permission so non-admins can't edit production prompts.
The tradeoff nobody tells you about: the moment a prompt lives in the DB instead of git, you
lose your diff, your code review, and your "what shipped on Tuesday" answer. Versioning here is
an Int counter, not a history — it tells you this is v7, not what v5 said. If prompts
matter to you (and for a RAG system they do), either export template changes to your audit log
or keep the canonical copy in git and treat the DB as a runtime override. I lean toward the
latter and haven't fully committed to either, which is honest.
Every chat call runs through trackApiUsage, which records the endpoint, model, token counts,
success/failure, and status code. The provider's own token counts are used where available; the
streaming path accumulates them from chunks and reports once at the end:
if provider_id and user_id and (total_prompt_tokens or total_completion_tokens):
await provider_client.report_usage(
provider_id=provider_id, user_id=user_id, model=request.model,
prompt_tokens=total_prompt_tokens, completion_tokens=total_completion_tokens,
endpoint="/api/chat/stream",
)
This is what lets the credit check in the request path exist at all, and what turns "our AI
bill doubled" from a mystery into a query. The model is chosen at runtime and can be swapped
per provider without a deploy, so cost is tracked against the model that actually answered, not
a config constant.
Caveat: the pre-request budget check uses a character-based token estimate
(chars / 4 × 1.2), not the real tokenizer. It's a soft gate to stop runaway spend, not a
billing system. The recorded usage is the accurate number; the estimate is a heuristic that
will occasionally reject a request slightly early. Don't invoice off the estimate.
chimerai add guardrails gives you four endpoints. The PII detector is regex-based — which is
both its strength (fast, no network, deterministic) and its ceiling (it catches patterns, not
meaning):
"email": re.compile(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b'),
"phone": re.compile(r'\b\d{3}[-.]?\d{3}[-.]?\d{4}\b'),
"ssn": re.compile(r'\b\d{3}-\d{2}-\d{4}\b'),
"credit_card": re.compile(r'\b\d{4}[- ]?\d{4}[- ]?\d{4}[- ]?\d{4}\b'),
Endpoints:
| Method | Path | What |
|---|---|---|
| POST | /api/guardrails/pii/detect |
find email/phone/SSN/CC/IP |
| POST | /api/guardrails/pii/redact |
mask the above |
| POST | /api/guardrails/toxicity |
keyword score 0–1 + risk level |
| POST | /api/guardrails/injection |
prompt-injection patterns |
Toxicity and injection detection are keyword/pattern-based with a confidence score and a risk
classification. That's enough to catch the obvious stuff and to add a cheap first line of
defense before a request hits your paid model.
Where it stops being sufficient: US phone/SSN/CC formats only (a German project will want
different patterns — worth noting since this kit is partly German-authored); "toxicity" by
keyword misses rephrased hostility entirely; regex injection detection is a tripwire, not a
classifier. If compliance demands real content moderation, put a dedicated model in front and
keep these as a fast pre-filter. detect vs redact as separate calls is deliberate — you
often want to log what was found even when you forward a masked version.
The RAG pipeline in the same stack reads its system prompt from category rag (feature 1),
runs the user query through guardrails before embedding (feature 3), and reports token usage
after the LLM call (feature 2). Three boring features, one grounded-and-metered request.
None of this is the part that demos well. It's the part that lets you sleep once users find
your app.
Repo: github.com/armbur19-collab/chimerai-app