{"slug": "compress-before-you-prompt-how-a-74k-star-token-first-architecture-is-making-ai", "title": "Compress Before You Prompt: How a 74K-Star Token-First Architecture Is Making AI Coding Agents Smarter, Cheaper, and Actually Honest", "summary": "A developer-built project that has accumulated over 74,000 GitHub stars in under a year introduces a \"token-first\" architecture that compresses code, documents, and conversation history before they reach an AI coding agent's context window. The approach reports 60-80% token cost reductions on real-world code understanding tasks with higher accuracy than uncompressed baselines, and sharply lower hallucination rates because agents work from compressed, high-fidelity representations of actual code rather than inventing function signatures and imports.", "body_md": "*Originally published on [tamiz.pro](https://tamiz.pro/insights/token-first-architecture-ai-coding-agents-compression).*\n\nThe context window is the new bottleneck. Every AI coding agent you've used — whether it's a local CLI assistant, a cloud IDE copilot, or an autonomous refactor bot — is fundamentally constrained by the same problem: models have finite context, but codebases are infinite in complexity. The result is a brutal engineering tradeoff. Stuff more context in, and you pay exponentially more per token while quality degrades from attention dilution. Stuff less in, and your agent hallucinates APIs, invents dependencies, and confidently writes code that doesn't compile against your actual codebase.\n\nA project that has accumulated over 74,000 GitHub stars in under a year offers a radically different approach. Instead of treating the context window as a fill-to-capacity resource, it treats token budget as a *scarce asset to be managed* — compressing code, documents, and conversation history before they ever reach the model. This \"token-first\" architecture inverts the traditional prompt engineering paradigm: rather than asking \"what should I put in the prompt?\", it asks \"what is the minimum information the model needs, expressed in the most information-dense form possible?\"\n\nThe results are striking. Benchmarks show 60-80% token cost reduction on real-world code understanding tasks, with *higher* accuracy than uncompressed baselines. More importantly, hallucination rates drop dramatically — the agent stops inventing function signatures, fake imports, and nonexistent methods because it's working with surgically compressed, high-fidelity representations of actual code.\n\nThis article dissects the architecture, explains the compression pipeline in detail, and walks through the engineering decisions that make this approach viable at scale.\n\nModern LLMs advertise context windows of 128K to 200K tokens. This sounds generous until you realize what actually goes into a coding agent's context:\n\n| Context Component | Typical Token Cost | Notes | \n|---|---|---|\n| System prompt + tool definitions | 2,000–5,000 | Fixed overhead | \n| Conversation history | 5,000–20,000 | Grows linearly with turns | \n| Repository structure (file tree) | 3,000–10,000 | For medium projects | \n| Referenced source files | 10,000–50,000 | The biggest variable | \n| Documentation / README | 2,000–8,000 | Often low signal | \n| Linting/test output | 1,000–5,000 | Noisy | \n\nA single \"refactor this module\" request can consume 80,000+ tokens of context, leaving almost no room for the model's actual reasoning. Worse, attention mechanisms don't handle uniform token importance well — a 50,000-token context with only 3,000 relevant tokens produces measurably worse outputs than a focused 5,000-token context.\n\nThe traditional approach has been *retrieval augmentation*: use a vector database to find \"relevant\" chunks and stuff them in. This helps, but it's blunt. Vector similarity doesn't understand code semantics — a function named `calculate` might match a completely unrelated `calculate` in another module. And once chunks are in the context window, they're immutable. The model sees them all with equal weight.\n\nThe token-first architecture introduces a **compression layer** between the raw codebase and the model. Instead of feeding source files directly into the prompt, every piece of code passes through a multi-stage pipeline that reduces token count while preserving semantic fidelity.\n\nHere's the high-level data flow:\n\n```\n┌─────────────────────────────────────────────────────────────────┐\n│                    AI Coding Agent Runtime                        │\n├─────────────────────────────────────────────────────────────────┤\n│                                                                   │\n│  ┌──────────┐    ┌──────────────┐    ┌───────────────────────┐  │\n│  │  Raw     │───▶│  Compression │───▶│  Token Budget         │  │\n│  │  Codebase│    │  Pipeline    │    │  Allocator             │  │\n│  └──────────┘    └──────────────┘    └───────────┬───────────┘  │\n│       │                    │                      │              │\n│       │                    ▼                      ▼              │\n│       │            ┌──────────────┐    ┌───────────────────────┐ │\n│       │            │  AST Parser  │    │  Context Assembly     │ │\n│       │            │  + Indexer   │    │  + Prompt Builder     │ │\n│       │            └──────────────┘    └───────────┬───────────┘ │\n│       │                                             │             │\n│       │                                             ▼             │\n│       │                                    ┌───────────────┐     │\n│       │                                    │  LLM Backend  │     │\n│       │                                    │  (Any Model)  │     │\n│       │                                    └───────────────┘     │\n│       ▼                                                         │\n│  ┌──────────┐                                                   │\n│  │  Git /   │                                                   │\n│  │  FS      │                                                   │\n│  └──────────┘                                                   │\n└─────────────────────────────────────────────────────────────────┘\n```\n\nThe key architectural insight is that compression happens *before* any model call. The pipeline is deterministic, fast (typically under 50ms for a 2,000-line file), and produces representations that are dramatically more information-dense than raw source code.\n\nThe compression pipeline consists of four stages, each targeting a different class of redundancy:\n\nRaw source code contains massive redundancy from a semantic perspective. Consider this TypeScript function:\n\n```\n// Original: 187 tokens\nexport async function processUserData(\n  userId: string,\n  options: ProcessOptions = {\n    includeHistory: false,\n    maxRecords: 100,\n    timeout: 5000\n  }: ProcessOptions\n): Promise<UserDataResponse> {\n  if (!userId) {\n    throw new Error('User ID is required');\n  }\n\n  const user = await userService.findById(userId);\n  if (!user) {\n    throw new Error(`User not found: ${userId}`);\n  }\n\n  const records = await recordService.getRecent(userId, options.maxRecords);\n\n  return {\n    id: user.id,\n    name: user.name,\n    email: user.email,\n    records: records.map(r => ({\n      id: r.id,\n      timestamp: r.timestamp,\n      value: r.value\n    })),\n    processedAt: new Date().toISOString()\n  };\n}\n```\n\nThe AST-based compressor transforms this into a **semantic skeleton**:\n\n```\n// Compressed: 42 tokens\nfn processUserData(userId: string, options?: ProcessOptions) -> UserDataResponse\n  deps: userService.findById, recordService.getRecent\n  throws: Error(userId required), Error(user not found)\n  returns: {id, name, email, records[{id,timestamp,value}], processedAt}\n  defaults: includeHistory=false, maxRecords=100, timeout=5000\n```\n\nThis preserves every piece of information the model actually needs — the function signature, its dependencies, error conditions, return shape, and default values — while eliminating formatting, boilerplate, and implementation details that are irrelevant for *understanding* the code's interface.\n\nRather than including full source files for every imported module, the architecture maintains a **compressed dependency graph**. Each module is represented as a node with its exported interfaces, and edges represent import relationships.\n\n```\n{\n  \"module\": \"src/services/userService.ts\",\n  \"exports\": [\n    {\n      \"name\": \"findById\",\n      \"signature\": \"(id: string) => Promise<User | null>\",\n      \"sideEffects\": [\"db.query\"]\n    },\n    {\n      \"name\": \"createUser\",\n      \"signature\": \"(data: CreateUserInput) => Promise<User>\",\n      \"sideEffects\": [\"db.insert\", \"eventBus.emit\"]\n    }\n  ],\n  \"imports\": [\"src/models/User.ts\", \"src/db/connection.ts\"],\n  \"tokenCountOriginal\": 890,\n  \"tokenCountCompressed\": 95\n}\n```\n\nWhen the model needs to understand how `processUserData` works, it sees the compressed signatures of `userService` and `recordService` — not their full implementations. If it needs deeper detail, it can request expansion of specific functions.\n\nMulti-turn coding conversations accumulate enormous context. The architecture applies **progressive summarization** to earlier turns:\n\n```\n// Turn 1-3 (raw): ~3,200 tokens\nUser: Can you refactor the auth module to use JWT instead of sessions?\nAgent: I'll start by examining the current auth implementation...\n[agent reads 3 files, proposes changes, applies patches]\n\n// After compression: ~340 tokens\n[Turns 1-3 Summary] Refactored auth from session-based to JWT.\nModified: auth/middleware.ts (JWT validation), auth/routes.ts (token extraction),\nauth/models.ts (added TokenPayload interface). Removed: sessionStore.ts.\nKey decision: Using HS256 with 24h expiry, refresh tokens in httpOnly cookies.\n```\n\nThe compression preserves decisions, file modifications, and architectural choices — the things a model needs to maintain consistency across turns — while discarding exploratory reasoning, intermediate states, and verbose explanations.\n\nThis is the most architecturally significant innovation. Rather than letting the model generate freely, the system allocates a **token budget** across different components of the response:\n\n```\n{\n  \"totalBudget\": 4096,\n  \"allocation\": {\n    \"reasoning\": 512,\n    \"code\": 2560,\n    \"explanation\": 768,\n    \"metadata\": 256\n  }\n}\n```\n\nThe model is instructed to stay within these bounds, and the runtime enforces truncation or regeneration if any section exceeds its allocation. This prevents the common failure mode where a model spends 3,000 tokens explaining context before producing 500 tokens of actual code.\n\nThe core of the compression engine is a language-aware AST parser that transforms source code into a compact semantic representation. Let's look at how this works in practice.\n\n``` python\n# Simplified representation of the compression pipeline\nfrom dataclasses import dataclass\nfrom typing import Literal\n\n@dataclass\nclass CompressedFunction:\n    name: str\n    params: list[tuple[str, str]]  # (name, type)\n    return_type: str\n    dependencies: list[str]        # external calls\n    side_effects: list[str]       # mutations, I/O\n    errors: list[str]             # thrown errors\n    complexity: Literal[\"trivial\", \"moderate\", \"complex\"]\n\n    def to_prompt_tokens(self) -> str:\n        \"\"\"Serialize to minimal token representation.\"\"\"\n        params_str = \", \".join(f\"{n}: {t}\" for n, t in self.params)\n        parts = [f\"fn {self.name}({params_str}) -> {self.return_type}\"]\n        if self.dependencies:\n            parts.append(f\"  calls: {', '.join(self.dependencies)}\")\n        if self.side_effects:\n            parts.append(f\"  mutates: {', '.join(self.side_effects)}\")\n        if self.errors:\n            parts.append(f\"  throws: {', '.join(self.errors)}\")\n        parts.append(f\"  complexity: {self.complexity}\")\n        return \"\\n\".join(parts)\n\ndef compress_file(source: str, language: str, budget: int) -> str:\n    \"\"\"\n    Compress a source file to fit within token budget.\n\n    Strategy: lossless for public interfaces, lossy for internals.\n    Priority order: exported > used-by-current-task > private > unused\n    \"\"\"\n    ast = parse_ast(source, language)\n\n    # Phase 1: Extract all symbols with metadata\n    symbols = extract_symbols(ast)\n\n    # Phase 2: Score each symbol by relevance\n    for sym in symbols:\n        sym.score = compute_relevance(sym, current_task_context)\n\n    # Phase 3: Greedy selection within budget\n    symbols.sort(key=lambda s: s.score, reverse=True)\n    selected = []\n    used_tokens = 0\n    for sym in symbols:\n        cost = estimate_tokens(sym.to_prompt_tokens())\n        if used_tokens + cost <= budget:\n            selected.append(sym)\n            used_tokens += cost\n\n    # Phase 4: If under budget, expand top-scoring symbols\n    remaining = budget - used_tokens\n    for sym in selected:\n        if remaining <= 0:\n            break\n        expanded_cost = estimate_tokens(sym.to_full_source())\n        if expanded_cost <= remaining:\n            sym.expanded = True\n            remaining -= expanded_cost\n\n    return serialize(selected)\n```\n\nThe key insight is **adaptive compression**: different parts of the codebase get different compression ratios based on their relevance to the current task. A function the model is about to modify gets near-full fidelity. A dependency it merely calls gets a signature-only representation. Unrelated code in the same file might be omitted entirely.\n\nDifferent languages have different redundancy patterns. The compressor uses language-specific rules:\n\n| Language | Primary Redundancy | Compression Strategy | Typical Ratio | \n|---|---|---|---|\n| TypeScript/JavaScript | Type annotations, JSDoc, verbose object literals | Strip types for internal, keep for exports | 4-6x | \n| Python | Docstrings, type hints, decorator boilerplate | Preserve signatures, compress bodies | 3-5x | \n| Go | Error checking patterns, context propagation | Collapse error guards to `throws:` lists | 5-8x | \n| Rust | Trait bounds, lifetimes, boilerplate impl blocks | Abstract trait impls to capability lists | 6-10x | \n| Java | Getters/setters, annotations, imports | Collapse CRUD, strip annotations | 8-12x | \n\nTraditional RAG systems use fixed-size chunks (typically 512-1024 tokens) with overlap. This is fundamentally wrong for code, where a 50-line function is an atomic unit and splitting it across chunks destroys meaning.\n\nThe token-first architecture uses **semantic chunking** based on AST boundaries:\n\n``` php\ndef semantic_chunk(ast: AST, target_tokens: int) -> list[Chunk]:\n    \"\"\"\n    Split code into semantically coherent chunks.\n    Never splits across function/class boundaries.\n    Groups related symbols into single chunks.\n    \"\"\"\n    nodes = ast.body  # top-level declarations\n    chunks = []\n    current_chunk = Chunk()\n\n    for node in nodes:\n        node_tokens = estimate_tokens(node)\n\n        # Check if adding this node would exceed budget\n        if current_chunk.token_count + node_tokens > target_tokens:\n            if current_chunk.nodes:\n                chunks.append(current_chunk)\n                current_chunk = Chunk()\n\n        # Special handling for large classes\n        if isinstance(node, ClassNode) and node_tokens > target_tokens:\n            chunks.append(chunk_large_class(node, target_tokens))\n            continue\n\n        current_chunk.add(node)\n\n    if current_chunk.nodes:\n        chunks.append(current_chunk)\n\n    # Merge adjacent small chunks (avoid fragmentation)\n    chunks = merge_small_chunks(chunks, min_tokens=target_tokens // 3)\n\n    return chunks\n```\n\nEach chunk receives a relevance score based on multiple signals:\n\n``` php\ndef compute_relevance(chunk: Chunk, query: str, context: TaskContext) -> float:\n    \"\"\"\n    Multi-signal relevance scoring for context selection.\n    Returns float in [0.0, 1.0].\n    \"\"\"\n    signals = {}\n\n    # Signal 1: Direct symbol match (highest weight)\n    if any(sym in query for sym in chunk.exported_symbols):\n        signals['direct_match'] = 0.9\n\n    # Signal 2: Import graph proximity\n    distance = shortest_import_path(chunk.module, context.target_module)\n    signals['import_proximity'] = max(0, 1.0 - distance * 0.3)\n\n    # Signal 3: Embedding similarity (semantic)\n    signals['semantic'] = cosine_similarity(\n        embed(chunk.compressed_repr),\n        embed(query)\n    )\n\n    # Signal 4: Recent modification (temporal relevance)\n    if chunk.last_modified_hours < 24:\n        signals['recency'] = 0.3\n\n    # Signal 5: Git blame overlap with modified files\n    if chunk.file in context.modified_files:\n        signals['modified'] = 0.7\n\n    # Weighted combination\n    weights = {\n        'direct_match': 0.30,\n        'import_proximity': 0.25,\n        'semantic': 0.20,\n        'recency': 0.10,\n        'modified': 0.15\n    }\n\n    score = sum(signals.get(k, 0) * w for k, w in weights.items())\n    return min(1.0, score)\n```\n\nThis multi-signal approach is dramatically more accurate than pure vector similarity for code. A function that's semantically dissimilar to the query but sits in the same module as the target file will still score high due to import proximity and modification signals.\n\nHere's where the architecture gets philosophically interesting. The primary complaint about AI coding agents is dishonesty — they hallucinate APIs, invent imports, and confidently write code that references nonexistent functions. The token-first architecture addresses this through **grounded compression**.\n\nTraditional agents hallucinate because of context dilution. When a model sees 80,000 tokens of context, it cannot maintain precise recall of every function signature. Under pressure to produce output, it fills gaps with plausible-sounding hallucinations:\n\n``` python\n# Model hallucination example (traditional agent)\nfrom myapp.utils import parse_config  # DOES NOT EXIST\nresult = process_data(config)          # process_data doesn't accept config param\n```\n\nThe compressed context is **exhaustive for interfaces**. Every function, class, and exported symbol in the dependency graph appears in the compressed representation with its exact signature. There is no gap for the model to hallucinate into.\n\n```\n# What the model actually sees (compressed but complete):\n\nAvailable functions in scope:\n  fn parseConfig(path: string) -> AppConfig  [src/config/parser.ts]\n  fn loadEnv(file?: string) -> Record<string,string>  [src/config/env.ts]\n  fn process_data(input: InputData, opts?: ProcessOpts) -> Output  [src/core/engine.ts]\n\n  NOTE: These are ALL exported functions in the dependency graph.\n  If you need a function not listed here, it does not exist.\n```\n\nThe architecture explicitly tells the model: *this is the complete set of available functions*. There is no implicit knowledge, no retrieval gap. If the model needs something that isn't listed, it must ask or admit it doesn't exist.\n\nThe system prompt includes an explicit honesty contract:\n\n```\nHONESTY CONTRACT:\n1. You are given the COMPLETE list of available functions, types, and modules.\n2. If you need something not in this list, state: \"Not available in current context\"\n3. Never invent function names, parameters, or import paths.\n4. If uncertain whether something exists, request clarification.\n5. Your compressed context is authoritative — it reflects the actual codebase state.\n```\n\nThis combination of exhaustive compressed interfaces plus explicit honesty instructions dramatically reduces hallucination. In benchmarks, the rate of invented function calls drops from ~12% (traditional RAG) to ~2% (token-first compression).\n\nThe token budget allocator is the central control mechanism that balances cost, quality, and completeness. It operates as a real-time optimization problem:\n\n``` python\nclass TokenBudgetAllocator:\n    def __init__(self, model_max_context: int, cost_per_token: float, budget_limit: float):\n        self.max_context = model_max_context\n        self.cost_per_token = cost_per_token\n        self.budget_limit = budget_limit\n\n        # Fixed allocations\n        self.system_prompt_tokens = 2048\n        self.safety_margin = 512  # room for model reasoning\n\n        # Dynamic allocation (the interesting part)\n        self.available = model_max_context - self.system_prompt_tokens - self.safety_margin\n\n    def allocate(self, components: list[ContextComponent]) -> dict[str, int]:\n        \"\"\"\n        Distribute available tokens across context components\n        using utility-maximizing allocation.\n        \"\"\"\n        # Score each component by utility density (info per token)\n        for comp in components:\n            comp.utility_density = comp.expected_usefulness / comp.token_cost\n            comp.compressed_cost = comp.token_cost / comp.compression_ratio\n\n        # Greedy allocation by utility density\n        components.sort(key=lambda c: c.utility_density, reverse=True)\n\n        allocation = {}\n        remaining = self.available\n\n        for comp in components:\n            # Allocate compressed version first\n            comp_allocation = min(comp.compressed_cost, remaining)\n            allocation[comp.name] = int(comp_allocation)\n            remaining -= comp_allocation\n\n            if remaining <= 0:\n                break\n\n        # Validate against cost budget\n        total_cost = sum(\n            allocation[c.name] * c.cost_per_token\n            for c in components\n        )\n        if total_cost > self.budget_limit:\n            # Scale down proportionally, preserving highest-utility components\n            scale = self.budget_limit / total_cost\n            for comp in components:\n                allocation[comp.name] = int(allocation[comp.name] * scale)\n\n        return allocation\n```\n\nThe allocator uses a tiered priority system:\n\n| Tier | Component | Allocation Priority | Compression Level | \n|---|---|---|---|\n| 0 | Current task description | Always full | None (raw) | \n| 1 | Modified files (this session) | High | Light (2-3x) | \n| 2 | Direct dependencies | Medium-High | Moderate (4-6x) | \n| 3 | Indirect dependencies | Medium | Heavy (6-10x) | \n| 4 | Conversation summary | Medium | Progressive (varies) | \n| 5 | Project config / conventions | Low | Maximum (10-15x) | \n| 6 | Documentation / README | Low | Maximum or omit | \n\nLet's walk through a concrete implementation of the compression pipeline for a TypeScript project.\n\n``` js\n// src/compression/ast-parser.ts\nimport { parse, Node, FunctionDeclaration, ClassDeclaration, ExportStatement } from 'typescript';\n\nexport interface CompressedSymbol {\n  name: string;\n  kind: 'function' | 'class' | 'interface' | 'const' | 'enum';\n  signature: string;\n  dependencies: string[];\n  isExported: boolean;\n  tokenCostOriginal: number;\n  tokenCostCompressed: number;\n}\n\nexport function extractSymbols(source: string, fileName: string): CompressedSymbol[] {\n  const sf = parse(source, { fileName });\n  const symbols: CompressedSymbol[] = [];\n\n  sf.forEachChild(node => {\n    if (isExportStatement(node)) {\n      node.forEachChild(inner => handleExported(inner, symbols));\n    } else if (isFunctionDeclaration(node) || isClassDeclaration(node)) {\n      handleExported(node, symbols);\n    } else if (isVariableStatement(node)) {\n      // Check if it's an exported const\n      if (hasModifier(node, 'export')) {\n        handleExported(node, symbols);\n      }\n    }\n  });\n\n  return symbols;\n}\n\nfunction handleExported(\n  node: Node,\n  symbols: CompressedSymbol[]\n): void {\n  if (isFunctionDeclaration(node)) {\n    const fn = node as FunctionDeclaration;\n    const params = fn.parameters?.map(p => \n      `${p.name.getText()}: ${p.type?.getText() || 'any'}`\n    ) || [];\n    const returnType = fn.type?.getText() || 'void';\n\n    // Extract called functions\n    const deps = extractCalledFunctions(fn.body);\n\n    symbols.push({\n      name: fn.name.getText(),\n      kind: 'function',\n      signature: `${fn.name.getText()}(${params.join(', ')}) -> ${returnType}`,\n      dependencies: deps,\n      isExported: true,\n      tokenCostOriginal: countTokens(fn.getText()),\n      tokenCostCompressed: countTokens(formatCompressedSymbol(symbols[symbols.length - 1]))\n    });\n  }\n  // ... handle classes, interfaces, etc.\n}\n\nfunction extractCalledFunctions(body: Node): string[] {\n  const called = new Set<string>();\n\n  function visit(node: Node) {\n    if (isCallExpression(node)) {\n      const expr = node.expression;\n      if (isIdentifier(expr)) {\n        called.add(expr.text);\n      } else if (isPropertyAccessExpression(expr)) {\n        called.add(expr.getText());\n      }\n    }\n    node.forEachChild(visit);\n  }\n\n  visit(body);\n  return [...called];\n}\njs\n// src/compression/engine.ts\nimport { extractSymbols, CompressedSymbol } from './ast-parser';\nimport { estimateTokenCount } from './tokenizer';\n\nexport interface CompressionResult {\n  compressed: string;\n  originalTokens: number;\n  compressedTokens: number;\n  ratio: number;\n  symbolsIncluded: number;\n  symbolsOmitted: number;\n}\n\nexport class CompressionEngine {\n  private compressionProfile: CompressionProfile;\n\n  constructor(profile: CompressionProfile) {\n    this.compressionProfile = profile;\n  }\n\n  compress(\n    source: string,\n    options: {\n      budget: number;\n      focusSymbols?: string[];\n      includeDependencies?: boolean;\n    }\n  ): CompressionResult {\n    const symbols = extractSymbols(source, options.fileName || 'unknown');\n    const originalTokens = estimateTokenCount(source);\n\n    // Score symbols by relevance\n    const scored = symbols.map(sym => ({\n      symbol: sym,\n      score: this.scoreSymbol(sym, options.focusSymbols || [])\n    }));\n\n    // Sort by score, then greedily allocate budget\n    scored.sort((a, b) => b.score - a.score);\n\n    const included: CompressedSymbol[] = [];\n    let usedTokens = 0;\n\n    for (const { symbol } of scored) {\n      const compressed = this.formatSymbol(symbol);\n      const cost = estimateTokenCount(compressed);\n\n      if (usedTokens + cost <= options.budget) {\n        included.push(symbol);\n        usedTokens += cost;\n      }\n    }\n\n    // Build output\n    const compressed = included.map(s => this.formatSymbol(s)).join('\\n');\n    const compressedTokens = estimateTokenCount(compressed);\n\n    return {\n      compressed,\n      originalTokens,\n      compressedTokens,\n      ratio: originalTokens / compressedTokens,\n      symbolsIncluded: included.length,\n      symbolsOmitted: symbols.length - included.length\n    };\n  }\n\n  private scoreSymbol(sym: CompressedSymbol, focus: string[]): number {\n    let score = 0;\n\n    if (sym.isExported) score += 0.5;\n    if (focus.includes(sym.name)) score += 1.0;\n    if (sym.kind === 'function') score += 0.3;\n    if (sym.kind === 'interface') score += 0.4;\n\n    // Penalize huge symbols (likely complex internals)\n    const complexityPenalty = Math.min(\n      0.5,\n      sym.tokenCostOriginal / 1000\n    );\n    score -= complexityPenalty;\n\n    return score;\n  }\n\n  private formatSymbol(sym: CompressedSymbol): string {\n    const prefix = sym.isExported ? 'export ' : '';\n    let result = `${prefix}${sym.kind} ${sym.signature}`;\n\n    if (sym.dependencies.length > 0) {\n      result += `\\n  // calls: ${sym.dependencies.join(', ')}`;\n    }\n\n    return result;\n  }\n}\njs\n// src/agent/context-builder.ts\nimport { CompressionEngine } from '../compression/engine';\nimport { TokenBudgetAllocator } from '../budget/allocator';\n\nexport class ContextBuilder {\n  private compressor: CompressionEngine;\n  private allocator: TokenBudgetAllocator;\n  private dependencyGraph: DependencyGraph;\n\n  async buildContext(\n    task: string,\n    modifiedFiles: string[],\n    model: string\n  ): Promise<AssembledContext> {\n    const modelConfig = getModelConfig(model);\n\n    // Step 1: Identify relevant modules via dependency graph\n    const relevantModules = this.dependencyGraph.getTransitiveDependencies(\n      modifiedFiles,\n      { depth: 2 }\n    );\n\n    // Step 2: Allocate token budget across components\n    const budget = this.allocator.allocate({\n      task,\n      modifiedFiles,\n      relevantModules,\n      totalBudget: modelConfig.maxContext - 2048 // reserve for system prompt\n    });\n\n    // Step 3: Compress each component within its allocation\n    const compressed = await Promise.all(\n      relevantModules.map(mod => {\n        const source = this.readFile(mod.path);\n        return this.compressor.compress(source, {\n          budget: budget[mod.path] || 500,\n          focusSymbols: this.extractFocusSymbols(task, mod.path)\n        });\n      })\n    );\n\n    // Step 4: Assemble final context\n    return {\n      systemPrompt: this.buildSystemPrompt(compressed),\n      totalTokensUsed: compressed.reduce((sum, c) => sum + c.compressedTokens, 0),\n      compressionRatio: this.computeOverallRatio(compressed),\n      honestyGuarantees: this.generateHonestyContract(compressed)\n    };\n  }\n}\n```\n\nThe token-first architecture has been benchmarked across multiple coding tasks. Here are representative results:\n\n| Task | Traditional Agent (tokens) | Token-First (tokens) | Savings | Quality Delta | \n|---|---|---|---|---|\n| Understand 50-file module | 78,400 | 9,200 | 88% | +3.2% accuracy | \n| Refactor auth system | 62,100 | 14,800 | 76% | +5.1% accuracy | \n| Add new API endpoint | 45,300 | 11,400 | 75% | +2.8% accuracy | \n| Debug failing test | 38,700 | 8,900 | 77% | +7.4% accuracy | \n| Cross-module rename | 91,200 | 18,600 | 79% | +4.6% accuracy | \n\n| Metric | Traditional RAG | Token-First | \n|---|---|---|\n| Invented function calls | 12.3% | 1.8% | \n| Wrong import paths | 8.7% | 0.9% | \n| Incorrect parameter types | 15.1% | 3.2% | \n| Nonexistent method calls | 9.4% | 1.1% | \n| Overall hallucination rate | 11.4% | 1.8% | \n\nThe compression pipeline adds minimal latency:\n\nThis is negligible compared to the 2-30 second LLM inference time it enables by reducing input size.\n\nNo architecture is universally superior. The token-first approach has specific failure modes:\n\nWhen debugging a subtle off-by-one error or a race condition, the model needs to see the *exact* implementation, not a compressed summary. The architecture must detect these scenarios and expand compression selectively.\n\nIf the task requires creating a pattern that doesn't exist in the codebase (e.g., \"implement a circuit breaker pattern\"), compressed context provides no reference material. The system needs to either fetch external knowledge or operate in a \"generation mode\" with minimal context.\n\nFor tasks like \"optimize this for throughput,\" the model needs to reason about specific implementation details — loop structures, allocation patterns, cache behavior. Heavy compression loses these details.\n\nThe dependency graph grows superlinearly in monorepos with 10,000+ modules. The graph traversal itself becomes expensive, and the compressed representation of \"all dependencies\" may exceed the context window even after compression.\n\n```\n// Adaptive compression: detect when full fidelity is needed\nfunction detectCompressionLevel(task: string, context: TaskContext): CompressionLevel {\n  if (/(debug|race|off-by-one|deadlock|memory leak)/i.test(task)) {\n    return 'minimal'; // Near-full source\n  }\n  if (/(refactor|rename|add endpoint|new feature)/i.test(task)) {\n    return 'aggressive'; // Heavy compression OK\n  }\n  if (/(optimize|performance|throughput|latency)/i.test(task)) {\n    return 'moderate'; // Keep implementation details\n  }\n  return 'standard'; // Default balanced compression\n}\n```\n\nIf you want to implement a token-first architecture for your own coding agent, here's the recommended build order:\n\n``` js\n// A minimal token-first agent in ~100 lines\nimport { parse } from 'typescript';\nimport OpenAI from 'openai';\n\nconst client = new OpenAI();\n\nasync function tokenFirstAgent(task: string, codebase: Map<string, string>) {\n  // Step 1: Compress all files to signatures\n  const compressed = new Map<string, string>();\n\n  for (const [path, source] of codebase) {\n    const sf = parse(source, { fileName: path });\n    const symbols: string[] = [];\n\n    sf.forEachChild(node => {\n      const text = node.getText();\n      // Keep only exported declarations' signatures\n      if (text.startsWith('export ')) {\n        // Truncate to first 200 chars (signature area)\n        const signature = text.split('\\n').slice(0, 5).join('\\n');\n        symbols.push(signature);\n      }\n    });\n\n    compressed.set(path, symbols.join('\\n'));\n  }\n\n  // Step 2: Build compressed context\n  const context = [...compressed.entries()]\n    .map(([path, sigs]) => `// ${path}\\n${sigs}`)\n    .join('\\n\\n');\n\n  // Step 3: Call model with compressed context\n  const response = await client.chat.completions.create({\n    model: 'gpt-4',\n    messages: [\n      {\n        role: 'system',\n        content: `\nYou have access to the following codebase interfaces:\n\n${context}\n\nHONESTY RULES:\n- These are the ONLY available functions/classes in scope\n- If you need something not listed, say it's not available\n- Never invent function names or import paths\n- Request clarification if uncertain\n        `.  rim()\n      },\n      { role: 'user', content: task }\n    ]\n  });\n\n  return response.choices[0].message.content;\n}\n```\n\nLarger context windows are *necessary but not sufficient*. Attention mechanisms have diminishing returns beyond ~30-40K tokens of actual context — the model can't maintain precise recall across that much information. Token-first compression achieves better accuracy *within* a smaller window than uncompressed context in a larger window. The combination of compression + larger windows is ideal, but compression provides independent value even with fixed window sizes.\n\nThe architecture works best with languages that have mature parsers (TypeScript, Python, Go, Rust, Java). For languages without standard parsers, you can use tree-sitter as a universal parser backend. The compression quality will be lower (you lose type information), but structural compression still provides 2-3x reduction.\n\nYes, and arguably more so. Open-source models typically have smaller context windows (8K-32K), making token budget management even more critical. The compressed representations also tend to be more compatible with smaller models because they reduce the cognitive load — the model doesn't need to parse and understand verbose source code, just work with clean interface descriptions.\n\nBuild an evaluation harness that compares outputs from compressed vs. uncompressed contexts on a held-out set of tasks. Key metrics: (1) Does the output compile? (2) Does it pass existing tests? (3) Does it reference only real functions? (4) Is the implementation semantically correct? If compression causes quality degradation on specific task types, tune the compression level for those scenarios using the adaptive system described above.\n\nThe token-first architecture represents a fundamental shift in how we think about AI coding agents. Instead of treating context as a passive container to fill, it treats it as an active resource to manage — compressing, prioritizing, and guaranteeing honesty at the architectural level. The 74K-star traction reflects something real: developers are frustrated with agents that hallucinate, waste tokens, and produce unreliable output. A compression-first approach addresses all three problems simultaneously, and the engineering is accessible enough that you can build a working version in under a week.\n\nFor more technical deep-dives on AI architecture and developer tooling, explore [Tamiz's Insights](https://tamiz.pro/insights).", "url": "https://wpnews.pro/news/compress-before-you-prompt-how-a-74k-star-token-first-architecture-is-making-ai", "canonical_source": "https://dev.to/tamizuddin/compress-before-you-prompt-how-a-74k-star-token-first-architecture-is-making-ai-coding-agents-d87", "published_at": "2026-10-02 18:02:25+00:00", "updated_at": "2026-10-02 18:07:38.000603+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "ai-infrastructure", "developer-tools"], "entities": ["GitHub", "tamiz.pro"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/compress-before-you-prompt-how-a-74k-star-token-first-architecture-is-making-ai", "markdown": "https://wpnews.pro/news/compress-before-you-prompt-how-a-74k-star-token-first-architecture-is-making-ai.md", "text": "https://wpnews.pro/news/compress-before-you-prompt-how-a-74k-star-token-first-architecture-is-making-ai.txt", "jsonld": "https://wpnews.pro/news/compress-before-you-prompt-how-a-74k-star-token-first-architecture-is-making-ai.jsonld"}}