{"slug": "your-prompts-are-code-manage-them-like-data-but-know-what-you-lose-when-you-do", "title": "Your prompts are code. Manage them like data — but know what you lose when you do", "summary": "A developer detailed the ChimerAI stack's approach to treating prompts as database-backed data rather than code, using a PromptTemplate model with Mustache-style variables, per-category defaults, and permission-gated editing. The writeup also covers per-request cost tracking via trackApiUsage and a regex-based PII guardrail endpoint set, while flagging tradeoffs: losing git diffs and review history for prompts, and a character-based pre-request budget estimate (chars / 4 × 1.2) that is a soft gate rather than a billing-grade token count.", "body_md": "The three things that separate an AI demo from an AI product aren't models. They're: being able\n\nto change a prompt without a deploy, knowing what each request cost you, and not leaking a\n\nuser's credit card into logs. None of these are exciting. All three are in the ChimerAI stack,\n\nso here's what they do and — more usefully — the tradeoffs they hide.\n\nInstead of prompt strings scattered across `.py` files, there's a `PromptTemplate` model with\n\nMustache-style variables:\n\n```\nmodel PromptTemplate {\n  id        String   @id @default(cuid())\n  name      String   @unique\n  category  String   // \"system\" | \"user\" | \"rag\" | \"agent\" | \"chat\" | \"custom\"\n  content   String   @db.Text\n  variables String[] // extracted from {{placeholders}} at save time\n  language  String   @default(\"en\")\n  version   Int      @default(1)\n  isDefault Boolean  @default(false)\n  isActive  Boolean  @default(true)\n  tags      String[]\n}\n```\n\nThe `variables` array is derived from the content, so the app can validate that a caller\n\nsupplied every `{{context}}` / `{{query}}` before rendering — a missing variable becomes a\n\n400 at the API boundary, not a prompt that silently ships a literal `{{context}}` to the model.\n\n```\nPOST /api/prompts\n{ \"name\": \"RAG System Prompt\",\n  \"category\": \"rag\",\n  \"content\": \"Use {{context}} to answer {{query}}\",\n  \"isDefault\": true }\n```\n\nExactly one default per category, multi-language (EN/DE/FR/ES/IT), and the whole thing is\n\nbehind an `manage_prompts` permission so non-admins can't edit production prompts.\n\n**The tradeoff nobody tells you about:** the moment a prompt lives in the DB instead of git, you\n\nlose your diff, your code review, and your \"what shipped on Tuesday\" answer. Versioning here is\n\nan `Int` counter, not a history — it tells you *this is v7*, not *what v5 said*. If prompts\n\nmatter to you (and for a RAG system they do), either export template changes to your audit log\n\nor keep the canonical copy in git and treat the DB as a runtime override. I lean toward the\n\nlatter and haven't fully committed to either, which is honest.\n\nEvery chat call runs through `trackApiUsage`, which records the endpoint, model, token counts,\n\nsuccess/failure, and status code. The provider's own token counts are used where available; the\n\nstreaming path accumulates them from chunks and reports once at the end:\n\n```\nif provider_id and user_id and (total_prompt_tokens or total_completion_tokens):\n    await provider_client.report_usage(\n        provider_id=provider_id, user_id=user_id, model=request.model,\n        prompt_tokens=total_prompt_tokens, completion_tokens=total_completion_tokens,\n        endpoint=\"/api/chat/stream\",\n    )\n```\n\nThis is what lets the credit check in the request path exist at all, and what turns \"our AI\n\nbill doubled\" from a mystery into a query. The model is chosen at runtime and can be swapped\n\nper provider without a deploy, so cost is tracked against the model that actually answered, not\n\na config constant.\n\n**Caveat:** the *pre-request* budget check uses a character-based token estimate\n\n(`chars / 4 × 1.2`), not the real tokenizer. It's a soft gate to stop runaway spend, not a\n\nbilling system. The recorded usage is the accurate number; the estimate is a heuristic that\n\nwill occasionally reject a request slightly early. Don't invoice off the estimate.\n\n`chimerai add guardrails` gives you four endpoints. The PII detector is regex-based — which is\n\nboth its strength (fast, no network, deterministic) and its ceiling (it catches patterns, not\n\nmeaning):\n\n```\n\"email\":       re.compile(r'\\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\\.[A-Z|a-z]{2,}\\b'),\n\"phone\":       re.compile(r'\\b\\d{3}[-.]?\\d{3}[-.]?\\d{4}\\b'),\n\"ssn\":         re.compile(r'\\b\\d{3}-\\d{2}-\\d{4}\\b'),\n\"credit_card\": re.compile(r'\\b\\d{4}[- ]?\\d{4}[- ]?\\d{4}[- ]?\\d{4}\\b'),\n```\n\nEndpoints:\n\n| Method | Path | What | \n|---|---|---|\n| POST | `/api/guardrails/pii/detect` | find email/phone/SSN/CC/IP | \n| POST | `/api/guardrails/pii/redact` | mask the above | \n| POST | `/api/guardrails/toxicity` | keyword score 0–1 + risk level | \n| POST | `/api/guardrails/injection` | prompt-injection patterns | \n\nToxicity and injection detection are keyword/pattern-based with a confidence score and a risk\n\nclassification. That's enough to catch the obvious stuff and to add a cheap first line of\n\ndefense before a request hits your paid model.\n\n**Where it stops being sufficient:** US phone/SSN/CC formats only (a German project will want\n\ndifferent patterns — worth noting since this kit is partly German-authored); \"toxicity\" by\n\nkeyword misses rephrased hostility entirely; regex injection detection is a tripwire, not a\n\nclassifier. If compliance demands real content moderation, put a dedicated model in front and\n\nkeep these as a fast pre-filter. `detect` vs `redact` as separate calls is deliberate — you\n\noften want to log what was found even when you forward a masked version.\n\nThe RAG pipeline in the same stack reads its system prompt from category `rag` (feature 1),\n\nruns the user query through guardrails before embedding (feature 3), and reports token usage\n\nafter the LLM call (feature 2). Three boring features, one grounded-and-metered request.\n\nNone of this is the part that demos well. It's the part that lets you sleep once users find\n\nyour app.\n\nRepo: `github.com/armbur19-collab/chimerai-app`", "url": "https://wpnews.pro/news/your-prompts-are-code-manage-them-like-data-but-know-what-you-lose-when-you-do", "canonical_source": "https://dev.to/armin_burger_ab136b2f8bb1/your-prompts-are-code-manage-them-like-data-but-know-what-you-lose-when-you-do-4i3j", "published_at": "2026-09-30 11:43:16+00:00", "updated_at": "2026-09-30 11:47:35.234583+00:00", "lang": "en", "topics": ["ai-products", "ai-tools", "large-language-models", "mlops", "ai-agents"], "entities": ["ChimerAI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/your-prompts-are-code-manage-them-like-data-but-know-what-you-lose-when-you-do", "markdown": "https://wpnews.pro/news/your-prompts-are-code-manage-them-like-data-but-know-what-you-lose-when-you-do.md", "text": "https://wpnews.pro/news/your-prompts-are-code-manage-them-like-data-but-know-what-you-lose-when-you-do.txt", "jsonld": "https://wpnews.pro/news/your-prompts-are-code-manage-them-like-data-but-know-what-you-lose-when-you-do.jsonld"}}