{"slug": "slashing-llm-bills-by-58-building-an-opentelemetry-semantic-cache-proxy-in", "title": "Slashing LLM Bills by 58%: Building an OpenTelemetry Semantic Cache Proxy in FastAPI", "summary": "A developer built an OpenAI-compatible FastAPI reverse proxy that combines Redis Vector Search semantic caching with OpenTelemetry instrumentation, reporting a 58% reduction in LLM inference bills. The proxy matches prompts by cosine similarity at a 0.92 threshold instead of exact SHA256 hashing, which the developer says yields cache hit rates under 4%, and exports per-tenant span and token metrics to Grafana, Jaeger, or Prometheus.", "body_md": "Production LLM deployments come with a dirty secret: between 30% and 60% of LLM queries in enterprise SaaS applications are duplicates or slight paraphrases. Support chat queries, recurring agent tasks, automated enrichment jobs, and document Q&A pipelines repeatedly hit upstream inference providers (OpenAI, Anthropic, Mistral) with functionally identical prompts.\n\nThe result? Insane monthly inference bills, unpredictable latency spikes, and engineering teams flying completely blind without granular, per-tenant observability.\n\nIn this technical breakdown, we'll build a production-grade, drop-in OpenAI-compatible reverse proxy using asynchronous Python (FastAPI), Redis Vector Search for semantic caching, and OpenTelemetry (OTel) for unified span and token metrics.\n\nWhen engineering leaders realize LLM spending is spiraling out of control, they usually evaluate two paths:\n\n`SHA256(prompt)`) in Redis yields miserable cache hit rates (under 4%) because even an extra whitespace or punctuation mark invalidates the cache.\n\n``` bash\nTraditional Flow:\nClient ──> Commercial Gateway ($$$ per token) ──> OpenAI API ($$$ per run)\n\nOptimized Self-Hosted Flow:\nClient ──> FastAPI Proxy ──> Redis Vector Search (Cosine >= 0.92) ──> Cache HIT (0ms upstream cost)\n                    │ (Cache MISS)\n                    └───> Upstream Provider (OTel Metric Exported)\n```\n\nTo solve this sustainably, we need an in-house proxy that acts as an exact drop-in replacement (`base_url=\"http://proxy-host/v1\"`), verifies semantic similarity via vector embeddings, and tracks granular OTel spans directly into Grafana, Jaeger, or Prometheus.\n\nThe proxy pipeline executes in five sequential phases:\n\n`Authorization`, `X-Tenant-ID`) and validate against token budget quotas.` text-embedding-3-small` or an on-premise SentenceTransformer).`similarity >= threshold` (e.g., 0.92), return cached completions instantly.\nLet's construct the core engine. Below is the semantic cache router and vector manager implemented in FastAPI and Redis.\n\n``` python\nimport numpy as np\nfrom redis.asyncio import Redis\nfrom redis.commands.search.field import VectorField, TextField\nfrom redis.commands.search.indexDefinition import IndexDefinition, IndexType\nfrom redis.commands.search.query import Query\n\nINDEX_NAME = \"idx:semantic_cache\"\nVECTOR_DIM = 1536  # text-embedding-3-small dimension\n\nasync def init_redis_indices(redis_client: Redis):\n    try:\n        await redis_client.ft(INDEX_NAME).info()\n    except Exception:\n        schema = (\n            TextField(\"prompt_text\"),\n            TextField(\"completion_text\"),\n            VectorField(\n                \"prompt_vector\",\n                \"HNSW\",\n                {\n                    \"TYPE\": \"FLOAT32\",\n                    \"DIM\": VECTOR_DIM,\n                    \"DISTANCE_METRIC\": \"COSINE\",\n                }\n            )\n        )\n        definition = IndexDefinition(prefix=[\"cache:prompt:\"], index_type=IndexType.HASH)\n        await redis_client.ft(INDEX_NAME).create_index(schema, definition=definition)\npython\nasync def check_semantic_cache(redis_client: Redis, query_vector: list[float], threshold: float = 0.92):\n    query_bytes = np.array(query_vector, dtype=np.float32).tobytes()\n\n    # KNN search retrieving nearest prompt vector\n    query = (\n        Query(\"*=>[KNN 1 @prompt_vector $vec AS score]\")\n        .sort_by(\"score\")\n        .return_fields(\"completion_text\", \"score\")\n        .dialect(2)\n    )\n\n    results = await redis_client.ft(INDEX_NAME).search(query, query_params={\"vec\": query_bytes})\n\n    if results.docs:\n        doc = results.docs[0]\n        # Cosine distance: 0 = identical, 2 = opposite. Similarity = 1 - distance\n        distance = float(doc.score)\n        similarity = 1.0 - distance\n\n        if similarity >= threshold:\n            return doc.completion_text, similarity\n\n    return None, 0.0\npython\nfrom fastapi import FastAPI, Request, HTTPException\nfrom opentelemetry import trace, metrics\nimport httpx\n\napp = FastAPI()\ntracer = trace.get_tracer(\"llm-proxy\")\nmeter = metrics.get_meter(\"llm-proxy\")\n\ntoken_counter = meter.create_counter(\n    name=\"llm_tokens_consumed_total\",\n    description=\"Total tokens consumed split by tenant and cache hit status\"\n)\n\n@app.post(\"/v1/chat/completions\")\nasync def chat_completions_proxy(request: Request):\n    tenant_id = request.headers.get(\"X-Tenant-ID\", \"default_tenant\")\n    payload = await request.json()\n    messages = payload.get(\"messages\", [])\n    last_prompt = messages[-1][\"content\"] if messages else \"\"\n\n    with tracer.start_as_current_span(\"llm_completion_router\") as span:\n        span.set_attribute(\"llm.tenant_id\", tenant_id)\n\n        # 1. Compute embedding (simplified dummy hook)\n        prompt_vec = await compute_embedding(last_prompt)\n\n        # 2. Check Semantic Cache\n        cached_response, similarity = await check_semantic_cache(app.state.redis, prompt_vec, threshold=0.92)\n\n        if cached_response:\n            span.set_attribute(\"llm.cache_hit\", True)\n            span.set_attribute(\"llm.cosine_similarity\", similarity)\n            token_counter.add(0, {\"tenant\": tenant_id, \"cache_hit\": \"true\"})\n\n            return {\n                \"id\": \"cached-completion\",\n                \"object\": \"chat.completion\",\n                \"choices\": [{\n                    \"index\": 0,\n                    \"message\": {\"role\": \"assistant\", \"content\": cached_response},\n                    \"finish_reason\": \"stop\"\n                }],\n                \"usage\": {\"prompt_tokens\": 0, \"completion_tokens\": 0, \"total_tokens\": 0}\n            }\n\n        # 3. Cache Miss - Forward to OpenAI\n        span.set_attribute(\"llm.cache_hit\", False)\n        async with httpx.AsyncClient() as client:\n            upstream_resp = await client.post(\n                \"https://api.openai.com/v1/chat/completions\",\n                headers={\"Authorization\": request.headers.get(\"Authorization\")},\n                json=payload,\n                timeout=30.0\n            )\n\n        data = upstream_resp.json()\n\n        # Extract token metrics for OTel\n        usage = data.get(\"usage\", {})\n        total_tokens = usage.get(\"total_tokens\", 0)\n        token_counter.add(total_tokens, {\"tenant\": tenant_id, \"cache_hit\": \"false\"})\n\n        # Store in Redis vector cache asynchronously\n        reply_text = data[\"choices\"][0][\"message\"][\"content\"]\n        await store_cache(app.state.redis, last_prompt, prompt_vec, reply_text)\n\n        return data\n```\n\nWhen pushing a semantic proxy to production, three critical real-world edge cases must be handled:\n\n`INCRBY` operations keyed by tenant and current minute (`tenant:{id}:budget:{YYYYMMDDHHmm}`). If the tenant exceeds their token ceiling, trip the circuit and return an HTTP `429 Too Many Requests` instantly.`0.88` similarity threshold, whereas legal or financial extraction tasks require `0.97` or strictly deterministic cache hits. Make the threshold configurable via HTTP headers (`X-Semantic-Threshold: 0.95`).\nYou now have the architectural blueprint to eliminate redundant API calls, enforce strict per-tenant token guardrails, and export enterprise-grade OpenTelemetry metrics without paying software seat taxes.\n\nYou can implement this architecture from scratch using the code snippets above. However, if you want a complete, battle-tested, production-ready solution with Docker Compose stacks, automated migrations, Jaeger/Grafana dashboards, and test fixtures ready for zero-downtime deployment, grab the turnkey package:\n\n`EARLYBIRD`** for 20% off)*\nTake control of your inference bill, secure your tenancy boundaries, and give your infrastructure the observability it deserves.", "url": "https://wpnews.pro/news/slashing-llm-bills-by-58-building-an-opentelemetry-semantic-cache-proxy-in", "canonical_source": "https://dev.to/reigen/slashing-llm-bills-by-58-building-an-opentelemetry-semantic-cache-proxy-in-fastapi-59j4", "published_at": "2026-10-11 16:49:56+00:00", "updated_at": "2026-10-11 17:00:12.160681+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "mlops", "ai-agents", "developer-tools"], "entities": ["FastAPI", "Redis", "OpenTelemetry", "OpenAI", "Anthropic", "Mistral", "Grafana", "Prometheus"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/slashing-llm-bills-by-58-building-an-opentelemetry-semantic-cache-proxy-in", "markdown": "https://wpnews.pro/news/slashing-llm-bills-by-58-building-an-opentelemetry-semantic-cache-proxy-in.md", "text": "https://wpnews.pro/news/slashing-llm-bills-by-58-building-an-opentelemetry-semantic-cache-proxy-in.txt", "jsonld": "https://wpnews.pro/news/slashing-llm-bills-by-58-building-an-opentelemetry-semantic-cache-proxy-in.jsonld"}}