Slashing LLM Bills by 58%: Building an OpenTelemetry Semantic Cache Proxy in FastAPI A developer built an OpenAI-compatible FastAPI reverse proxy that combines Redis Vector Search semantic caching with OpenTelemetry instrumentation, reporting a 58% reduction in LLM inference bills. The proxy matches prompts by cosine similarity at a 0.92 threshold instead of exact SHA256 hashing, which the developer says yields cache hit rates under 4%, and exports per-tenant span and token metrics to Grafana, Jaeger, or Prometheus. Production LLM deployments come with a dirty secret: between 30% and 60% of LLM queries in enterprise SaaS applications are duplicates or slight paraphrases. Support chat queries, recurring agent tasks, automated enrichment jobs, and document Q&A pipelines repeatedly hit upstream inference providers OpenAI, Anthropic, Mistral with functionally identical prompts. The result? Insane monthly inference bills, unpredictable latency spikes, and engineering teams flying completely blind without granular, per-tenant observability. In this technical breakdown, we'll build a production-grade, drop-in OpenAI-compatible reverse proxy using asynchronous Python FastAPI , Redis Vector Search for semantic caching, and OpenTelemetry OTel for unified span and token metrics. When engineering leaders realize LLM spending is spiraling out of control, they usually evaluate two paths: SHA256 prompt in Redis yields miserable cache hit rates under 4% because even an extra whitespace or punctuation mark invalidates the cache. bash Traditional Flow: Client ── Commercial Gateway $$$ per token ── OpenAI API $$$ per run Optimized Self-Hosted Flow: Client ── FastAPI Proxy ── Redis Vector Search Cosine = 0.92 ── Cache HIT 0ms upstream cost │ Cache MISS └─── Upstream Provider OTel Metric Exported To solve this sustainably, we need an in-house proxy that acts as an exact drop-in replacement base url="http://proxy-host/v1" , verifies semantic similarity via vector embeddings, and tracks granular OTel spans directly into Grafana, Jaeger, or Prometheus. The proxy pipeline executes in five sequential phases: Authorization , X-Tenant-ID and validate against token budget quotas. text-embedding-3-small or an on-premise SentenceTransformer . similarity = threshold e.g., 0.92 , return cached completions instantly. Let's construct the core engine. Below is the semantic cache router and vector manager implemented in FastAPI and Redis. python import numpy as np from redis.asyncio import Redis from redis.commands.search.field import VectorField, TextField from redis.commands.search.indexDefinition import IndexDefinition, IndexType from redis.commands.search.query import Query INDEX NAME = "idx:semantic cache" VECTOR DIM = 1536 text-embedding-3-small dimension async def init redis indices redis client: Redis : try: await redis client.ft INDEX NAME .info except Exception: schema = TextField "prompt text" , TextField "completion text" , VectorField "prompt vector", "HNSW", { "TYPE": "FLOAT32", "DIM": VECTOR DIM, "DISTANCE METRIC": "COSINE", } definition = IndexDefinition prefix= "cache:prompt:" , index type=IndexType.HASH await redis client.ft INDEX NAME .create index schema, definition=definition python async def check semantic cache redis client: Redis, query vector: list float , threshold: float = 0.92 : query bytes = np.array query vector, dtype=np.float32 .tobytes KNN search retrieving nearest prompt vector query = Query " = KNN 1 @prompt vector $vec AS score " .sort by "score" .return fields "completion text", "score" .dialect 2 results = await redis client.ft INDEX NAME .search query, query params={"vec": query bytes} if results.docs: doc = results.docs 0 Cosine distance: 0 = identical, 2 = opposite. Similarity = 1 - distance distance = float doc.score similarity = 1.0 - distance if similarity = threshold: return doc.completion text, similarity return None, 0.0 python from fastapi import FastAPI, Request, HTTPException from opentelemetry import trace, metrics import httpx app = FastAPI tracer = trace.get tracer "llm-proxy" meter = metrics.get meter "llm-proxy" token counter = meter.create counter name="llm tokens consumed total", description="Total tokens consumed split by tenant and cache hit status" @app.post "/v1/chat/completions" async def chat completions proxy request: Request : tenant id = request.headers.get "X-Tenant-ID", "default tenant" payload = await request.json messages = payload.get "messages", last prompt = messages -1 "content" if messages else "" with tracer.start as current span "llm completion router" as span: span.set attribute "llm.tenant id", tenant id 1. Compute embedding simplified dummy hook prompt vec = await compute embedding last prompt 2. Check Semantic Cache cached response, similarity = await check semantic cache app.state.redis, prompt vec, threshold=0.92 if cached response: span.set attribute "llm.cache hit", True span.set attribute "llm.cosine similarity", similarity token counter.add 0, {"tenant": tenant id, "cache hit": "true"} return { "id": "cached-completion", "object": "chat.completion", "choices": { "index": 0, "message": {"role": "assistant", "content": cached response}, "finish reason": "stop" } , "usage": {"prompt tokens": 0, "completion tokens": 0, "total tokens": 0} } 3. Cache Miss - Forward to OpenAI span.set attribute "llm.cache hit", False async with httpx.AsyncClient as client: upstream resp = await client.post "https://api.openai.com/v1/chat/completions", headers={"Authorization": request.headers.get "Authorization" }, json=payload, timeout=30.0 data = upstream resp.json Extract token metrics for OTel usage = data.get "usage", {} total tokens = usage.get "total tokens", 0 token counter.add total tokens, {"tenant": tenant id, "cache hit": "false"} Store in Redis vector cache asynchronously reply text = data "choices" 0 "message" "content" await store cache app.state.redis, last prompt, prompt vec, reply text return data When pushing a semantic proxy to production, three critical real-world edge cases must be handled: INCRBY operations keyed by tenant and current minute tenant:{id}:budget:{YYYYMMDDHHmm} . If the tenant exceeds their token ceiling, trip the circuit and return an HTTP 429 Too Many Requests instantly. 0.88 similarity threshold, whereas legal or financial extraction tasks require 0.97 or strictly deterministic cache hits. Make the threshold configurable via HTTP headers X-Semantic-Threshold: 0.95 . You now have the architectural blueprint to eliminate redundant API calls, enforce strict per-tenant token guardrails, and export enterprise-grade OpenTelemetry metrics without paying software seat taxes. You can implement this architecture from scratch using the code snippets above. However, if you want a complete, battle-tested, production-ready solution with Docker Compose stacks, automated migrations, Jaeger/Grafana dashboards, and test fixtures ready for zero-downtime deployment, grab the turnkey package: EARLYBIRD for 20% off Take control of your inference bill, secure your tenancy boundaries, and give your infrastructure the observability it deserves.