5 Proven Techniques for Token Compression and Prompt Optimization A guide outlines five token compression and prompt optimization techniques for reducing LLM API costs, including replacing verbose instructions with structured constraints, limiting few-shot examples to three to five, dynamic context trimming via sentence embeddings, and prompt caching of repeated prefixes. The guide cites a worked example in which a 36-token instruction is compressed to roughly 14 tokens, and states that trimming a 10,000-token knowledge base to the 800 relevant tokens can cut context costs by over 90%. It attributes the diminishing-returns finding on few-shot examples beyond three to five to research from Anthropic and academic benchmarks. 5 Proven Techniques for Token Compression and Prompt Optimization Reduce costs, improve response quality, and build leaner AI applications with these prompt engineering strategies. Every token counts. Whether you're building production applications with large language models LLMs or running experiments in a notebook, bloated prompts silently drain budgets and degrade response quality. Token compression is the practice of transmitting more intent with fewer tokens, and prompt optimization is how you structure that intent so models respond accurately and efficiently. This guide covers five techniques you can apply right away to reduce token consumption without sacrificing output quality, along with the reasoning behind each approach and practical code examples. 1. Replacing Verbose Instructions with Structured Constraints Long, conversational system prompts feel natural to write but cost significantly more than tightly structured equivalents. The fix is moving from narrative instructions to declarative constraints, using schema-like formatting that models parse efficiently. Instead of writing: Please make sure that when you respond, you always use bullet points and keep answers under 100 words. Do not include any preamble or sign-off at the end of your reply. Compress it to: Format: bullet points | Max: 100 words | Omit: preamble, sign-off That single line replaces 36 tokens with roughly 14. Across thousands of API calls, the savings compound quickly. Use pipe-delimited key-value pairs, YAML-style constraints, or JSON schema snippets depending on the model family you're working with. 2. Using Few-Shot Examples Strategically, Not Exhaustively Few-shot prompting — providing example input-output pairs before your actual request — dramatically improves output format consistency. The mistake most practitioners make is adding too many examples. Research from Anthropic and academic benchmarks consistently shows diminishing returns beyond three to five examples for most classification and generation tasks. Here's a lean three-shot prompt for sentiment labeling: system = """Label sentiment. Reply with one word: Positive, Negative, or Neutral. Examples: Input: "Shipped on time and well packaged." - Positive Input: "Completely broken out of the box." - Negative Input: "It arrived." - Neutral""" Three examples establish the pattern. Adding ten more rarely improves accuracy and often introduces contradictions that confuse the model. Audit your existing few-shot prompts and benchmark quality at one, three, and five examples before committing to a larger set. 3. Applying Dynamic Context Trimming for Long Documents When you pass long documents into a prompt — transcripts, legal text, knowledge base articles — you're almost always paying for tokens the model doesn't need. Dynamic context trimming retrieves only the relevant passage rather than the entire document. Here's a minimal implementation using cosine similarity with sentence embeddings https://www.sbert.net/ : python from sentence transformers import SentenceTransformer from sklearn.metrics.pairwise import cosine similarity import numpy as np model = SentenceTransformer "all-MiniLM-L6-v2" def trim context query, passages, top k=3 : q emb = model.encode query p embs = model.encode passages scores = cosine similarity q emb, p embs 0 top idx = np.argsort scores -top k: ::-1 return passages i for i in top idx Pass the filtered list of passages instead of the raw document. For a 10,000-token knowledge base where only 800 tokens are relevant, this technique alone can cut context costs by over 90%. 4. Caching Repeated Prompt Prefixes with Prompt Caching Many applications repeat identical system prompts across every user request: the same persona definition, the same tool descriptions, the same policy constraints. Sending those tokens fresh each time is unnecessary. Several inference providers — including Anthropic with its prompt caching https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching feature and OpenAI with automatic prefix caching — now store and reuse static prompt prefixes server-side, billing cached tokens at a fraction of standard input pricing. Structure your prompts so the stable content comes first and the dynamic content comes last: System prompt - static, 800 tokens <- Cached after first call Retrieved context - semi-static, 400 tokens <- Potentially cached User message - dynamic, 50 tokens <- Always fresh Before implementing, check your provider's caching documentation. Anthropic's prompt caching kicks in when the cached prefix exceeds a minimum token threshold and the cache is hit within a defined time window. 5. Compressing Chain-of-Thought Reasoning with Scratchpad Separation Chain-of-thought CoT prompting https://www.kdnuggets.com/2023/07/power-chain-thought-prompting-large-language-models.html improves model reasoning on complex tasks, but the reasoning trace itself — sometimes hundreds of tokens — often appears verbatim in your API response even when you only need the final answer. That inflates output token costs fast. The fix is to separate the reasoning scratchpad from the final answer using structured output markers: prompt = """Solve the problem step by step inside tags. Then provide only your final answer inside tags. Problem: A warehouse ships 240 units over 6 days at an uneven rate. Day 1-3 average: 30/day. What is the Day 4-6 average?""" Your application then parses and discards the