# Your prompts are code. Manage them like data — but know what you lose when you do

> Source: <https://dev.to/armin_burger_ab136b2f8bb1/your-prompts-are-code-manage-them-like-data-but-know-what-you-lose-when-you-do-4i3j>
> Published: 2026-09-30 11:43:16+00:00

The three things that separate an AI demo from an AI product aren't models. They're: being able

to change a prompt without a deploy, knowing what each request cost you, and not leaking a

user's credit card into logs. None of these are exciting. All three are in the ChimerAI stack,

so here's what they do and — more usefully — the tradeoffs they hide.

Instead of prompt strings scattered across `.py` files, there's a `PromptTemplate` model with

Mustache-style variables:

```
model PromptTemplate {
  id        String   @id @default(cuid())
  name      String   @unique
  category  String   // "system" | "user" | "rag" | "agent" | "chat" | "custom"
  content   String   @db.Text
  variables String[] // extracted from {{placeholders}} at save time
  language  String   @default("en")
  version   Int      @default(1)
  isDefault Boolean  @default(false)
  isActive  Boolean  @default(true)
  tags      String[]
}
```

The `variables` array is derived from the content, so the app can validate that a caller

supplied every `{{context}}` / `{{query}}` before rendering — a missing variable becomes a

400 at the API boundary, not a prompt that silently ships a literal `{{context}}` to the model.

```
POST /api/prompts
{ "name": "RAG System Prompt",
  "category": "rag",
  "content": "Use {{context}} to answer {{query}}",
  "isDefault": true }
```

Exactly one default per category, multi-language (EN/DE/FR/ES/IT), and the whole thing is

behind an `manage_prompts` permission so non-admins can't edit production prompts.

**The tradeoff nobody tells you about:** the moment a prompt lives in the DB instead of git, you

lose your diff, your code review, and your "what shipped on Tuesday" answer. Versioning here is

an `Int` counter, not a history — it tells you *this is v7*, not *what v5 said*. If prompts

matter to you (and for a RAG system they do), either export template changes to your audit log

or keep the canonical copy in git and treat the DB as a runtime override. I lean toward the

latter and haven't fully committed to either, which is honest.

Every chat call runs through `trackApiUsage`, which records the endpoint, model, token counts,

success/failure, and status code. The provider's own token counts are used where available; the

streaming path accumulates them from chunks and reports once at the end:

```
if provider_id and user_id and (total_prompt_tokens or total_completion_tokens):
    await provider_client.report_usage(
        provider_id=provider_id, user_id=user_id, model=request.model,
        prompt_tokens=total_prompt_tokens, completion_tokens=total_completion_tokens,
        endpoint="/api/chat/stream",
    )
```

This is what lets the credit check in the request path exist at all, and what turns "our AI

bill doubled" from a mystery into a query. The model is chosen at runtime and can be swapped

per provider without a deploy, so cost is tracked against the model that actually answered, not

a config constant.

**Caveat:** the *pre-request* budget check uses a character-based token estimate

(`chars / 4 × 1.2`), not the real tokenizer. It's a soft gate to stop runaway spend, not a

billing system. The recorded usage is the accurate number; the estimate is a heuristic that

will occasionally reject a request slightly early. Don't invoice off the estimate.

`chimerai add guardrails` gives you four endpoints. The PII detector is regex-based — which is

both its strength (fast, no network, deterministic) and its ceiling (it catches patterns, not

meaning):

```
"email":       re.compile(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b'),
"phone":       re.compile(r'\b\d{3}[-.]?\d{3}[-.]?\d{4}\b'),
"ssn":         re.compile(r'\b\d{3}-\d{2}-\d{4}\b'),
"credit_card": re.compile(r'\b\d{4}[- ]?\d{4}[- ]?\d{4}[- ]?\d{4}\b'),
```

Endpoints:

| Method | Path | What | 
|---|---|---|
| POST | `/api/guardrails/pii/detect` | find email/phone/SSN/CC/IP | 
| POST | `/api/guardrails/pii/redact` | mask the above | 
| POST | `/api/guardrails/toxicity` | keyword score 0–1 + risk level | 
| POST | `/api/guardrails/injection` | prompt-injection patterns | 

Toxicity and injection detection are keyword/pattern-based with a confidence score and a risk

classification. That's enough to catch the obvious stuff and to add a cheap first line of

defense before a request hits your paid model.

**Where it stops being sufficient:** US phone/SSN/CC formats only (a German project will want

different patterns — worth noting since this kit is partly German-authored); "toxicity" by

keyword misses rephrased hostility entirely; regex injection detection is a tripwire, not a

classifier. If compliance demands real content moderation, put a dedicated model in front and

keep these as a fast pre-filter. `detect` vs `redact` as separate calls is deliberate — you

often want to log what was found even when you forward a masked version.

The RAG pipeline in the same stack reads its system prompt from category `rag` (feature 1),

runs the user query through guardrails before embedding (feature 3), and reports token usage

after the LLM call (feature 2). Three boring features, one grounded-and-metered request.

None of this is the part that demos well. It's the part that lets you sleep once users find

your app.

Repo: `github.com/armbur19-collab/chimerai-app`
