cd /news/ai-agents/compress-before-you-prompt-how-a-74k… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-144047] src=dev.to β†— pub= topic=ai-agents verified=true sentiment=↑ positive

Compress Before You Prompt: How a 74K-Star Token-First Architecture Is Making AI Coding Agents Smarter, Cheaper, and Actually Honest

A developer-built project that has accumulated over 74,000 GitHub stars in under a year introduces a "token-first" architecture that compresses code, documents, and conversation history before they reach an AI coding agent's context window. The approach reports 60-80% token cost reductions on real-world code understanding tasks with higher accuracy than uncompressed baselines, and sharply lower hallucination rates because agents work from compressed, high-fidelity representations of actual code rather than inventing function signatures and imports.

by read21 min views1 publishedOct 2, 2026

Originally published on tamiz.pro.

The context window is the new bottleneck. Every AI coding agent you've used β€” whether it's a local CLI assistant, a cloud IDE copilot, or an autonomous refactor bot β€” is fundamentally constrained by the same problem: models have finite context, but codebases are infinite in complexity. The result is a brutal engineering tradeoff. Stuff more context in, and you pay exponentially more per token while quality degrades from attention dilution. Stuff less in, and your agent hallucinates APIs, invents dependencies, and confidently writes code that doesn't compile against your actual codebase.

A project that has accumulated over 74,000 GitHub stars in under a year offers a radically different approach. Instead of treating the context window as a fill-to-capacity resource, it treats token budget as a scarce asset to be managed β€” compressing code, documents, and conversation history before they ever reach the model. This "token-first" architecture inverts the traditional prompt engineering paradigm: rather than asking "what should I put in the prompt?", it asks "what is the minimum information the model needs, expressed in the most information-dense form possible?"

The results are striking. Benchmarks show 60-80% token cost reduction on real-world code understanding tasks, with higher accuracy than uncompressed baselines. More importantly, hallucination rates drop dramatically β€” the agent stops inventing function signatures, fake imports, and nonexistent methods because it's working with surgically compressed, high-fidelity representations of actual code.

This article dissects the architecture, explains the compression pipeline in detail, and walks through the engineering decisions that make this approach viable at scale.

Modern LLMs advertise context windows of 128K to 200K tokens. This sounds generous until you realize what actually goes into a coding agent's context:

Context Component Typical Token Cost Notes
System prompt + tool definitions 2,000–5,000 Fixed overhead
Conversation history 5,000–20,000 Grows linearly with turns
Repository structure (file tree) 3,000–10,000 For medium projects
Referenced source files 10,000–50,000 The biggest variable
Documentation / README 2,000–8,000 Often low signal
Linting/test output 1,000–5,000 Noisy

A single "refactor this module" request can consume 80,000+ tokens of context, leaving almost no room for the model's actual reasoning. Worse, attention mechanisms don't handle uniform token importance well β€” a 50,000-token context with only 3,000 relevant tokens produces measurably worse outputs than a focused 5,000-token context.

The traditional approach has been retrieval augmentation: use a vector database to find "relevant" chunks and stuff them in. This helps, but it's blunt. Vector similarity doesn't understand code semantics β€” a function named calculate might match a completely unrelated calculate in another module. And once chunks are in the context window, they're immutable. The model sees them all with equal weight.

The token-first architecture introduces a compression layer between the raw codebase and the model. Instead of feeding source files directly into the prompt, every piece of code passes through a multi-stage pipeline that reduces token count while preserving semantic fidelity.

Here's the high-level data flow:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    AI Coding Agent Runtime                        β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                                                                   β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚  β”‚  Raw     │───▢│  Compression │───▢│  Token Budget         β”‚  β”‚
β”‚  β”‚  Codebaseβ”‚    β”‚  Pipeline    β”‚    β”‚  Allocator             β”‚  β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚       β”‚                    β”‚                      β”‚              β”‚
β”‚       β”‚                    β–Ό                      β–Ό              β”‚
β”‚       β”‚            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚       β”‚            β”‚  AST Parser  β”‚    β”‚  Context Assembly     β”‚ β”‚
β”‚       β”‚            β”‚  + Indexer   β”‚    β”‚  + Prompt Builder     β”‚ β”‚
β”‚       β”‚            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β”‚       β”‚                                             β”‚             β”‚
β”‚       β”‚                                             β–Ό             β”‚
β”‚       β”‚                                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”‚
β”‚       β”‚                                    β”‚  LLM Backend  β”‚     β”‚
β”‚       β”‚                                    β”‚  (Any Model)  β”‚     β”‚
β”‚       β”‚                                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β”‚
β”‚       β–Ό                                                         β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                                                   β”‚
β”‚  β”‚  Git /   β”‚                                                   β”‚
β”‚  β”‚  FS      β”‚                                                   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                                                   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

The key architectural insight is that compression happens before any model call. The pipeline is deterministic, fast (typically under 50ms for a 2,000-line file), and produces representations that are dramatically more information-dense than raw source code.

The compression pipeline consists of four stages, each targeting a different class of redundancy:

Raw source code contains massive redundancy from a semantic perspective. Consider this TypeScript function:

// Original: 187 tokens
export async function processUserData(
  userId: string,
  options: ProcessOptions = {
    includeHistory: false,
    maxRecords: 100,
    timeout: 5000
  }: ProcessOptions
): Promise<UserDataResponse> {
  if (!userId) {
    throw new Error('User ID is required');
  }

  const user = await userService.findById(userId);
  if (!user) {
    throw new Error(`User not found: ${userId}`);
  }

  const records = await recordService.getRecent(userId, options.maxRecords);

  return {
    id: user.id,
    name: user.name,
    email: user.email,
    records: records.map(r => ({
      id: r.id,
      timestamp: r.timestamp,
      value: r.value
    })),
    processedAt: new Date().toISOString()
  };
}

The AST-based compressor transforms this into a semantic skeleton:

// Compressed: 42 tokens
fn processUserData(userId: string, options?: ProcessOptions) -> UserDataResponse
  deps: userService.findById, recordService.getRecent
  throws: Error(userId required), Error(user not found)
  returns: {id, name, email, records[{id,timestamp,value}], processedAt}
  defaults: includeHistory=false, maxRecords=100, timeout=5000

This preserves every piece of information the model actually needs β€” the function signature, its dependencies, error conditions, return shape, and default values β€” while eliminating formatting, boilerplate, and implementation details that are irrelevant for understanding the code's interface.

Rather than including full source files for every imported module, the architecture maintains a compressed dependency graph. Each module is represented as a node with its exported interfaces, and edges represent import relationships.

{
  "module": "src/services/userService.ts",
  "exports": [
    {
      "name": "findById",
      "signature": "(id: string) => Promise<User | null>",
      "sideEffects": ["db.query"]
    },
    {
      "name": "createUser",
      "signature": "(data: CreateUserInput) => Promise<User>",
      "sideEffects": ["db.insert", "eventBus.emit"]
    }
  ],
  "imports": ["src/models/User.ts", "src/db/connection.ts"],
  "tokenCountOriginal": 890,
  "tokenCountCompressed": 95
}

When the model needs to understand how processUserData works, it sees the compressed signatures of userService and recordService β€” not their full implementations. If it needs deeper detail, it can request expansion of specific functions.

Multi-turn coding conversations accumulate enormous context. The architecture applies progressive summarization to earlier turns:

// Turn 1-3 (raw): ~3,200 tokens
User: Can you refactor the auth module to use JWT instead of sessions?
Agent: I'll start by examining the current auth implementation...
[agent reads 3 files, proposes changes, applies patches]

// After compression: ~340 tokens
[Turns 1-3 Summary] Refactored auth from session-based to JWT.
Modified: auth/middleware.ts (JWT validation), auth/routes.ts (token extraction),
auth/models.ts (added TokenPayload interface). Removed: sessionStore.ts.
Key decision: Using HS256 with 24h expiry, refresh tokens in httpOnly cookies.

The compression preserves decisions, file modifications, and architectural choices β€” the things a model needs to maintain consistency across turns β€” while discarding exploratory reasoning, intermediate states, and verbose explanations.

This is the most architecturally significant innovation. Rather than letting the model generate freely, the system allocates a token budget across different components of the response:

{
  "totalBudget": 4096,
  "allocation": {
    "reasoning": 512,
    "code": 2560,
    "explanation": 768,
    "metadata": 256
  }
}

The model is instructed to stay within these bounds, and the runtime enforces truncation or regeneration if any section exceeds its allocation. This prevents the common failure mode where a model spends 3,000 tokens explaining context before producing 500 tokens of actual code.

The core of the compression engine is a language-aware AST parser that transforms source code into a compact semantic representation. Let's look at how this works in practice.

from dataclasses import dataclass
from typing import Literal

@dataclass
class CompressedFunction:
    name: str
    params: list[tuple[str, str]]  # (name, type)
    return_type: str
    dependencies: list[str]        # external calls
    side_effects: list[str]       # mutations, I/O
    errors: list[str]             # thrown errors
    complexity: Literal["trivial", "moderate", "complex"]

    def to_prompt_tokens(self) -> str:
        """Serialize to minimal token representation."""
        params_str = ", ".join(f"{n}: {t}" for n, t in self.params)
        parts = [f"fn {self.name}({params_str}) -> {self.return_type}"]
        if self.dependencies:
            parts.append(f"  calls: {', '.join(self.dependencies)}")
        if self.side_effects:
            parts.append(f"  mutates: {', '.join(self.side_effects)}")
        if self.errors:
            parts.append(f"  throws: {', '.join(self.errors)}")
        parts.append(f"  complexity: {self.complexity}")
        return "\n".join(parts)

def compress_file(source: str, language: str, budget: int) -> str:
    """
    Compress a source file to fit within token budget.

    Strategy: lossless for public interfaces, lossy for internals.
    Priority order: exported > used-by-current-task > private > unused
    """
    ast = parse_ast(source, language)

    symbols = extract_symbols(ast)

    for sym in symbols:
        sym.score = compute_relevance(sym, current_task_context)

    symbols.sort(key=lambda s: s.score, reverse=True)
    selected = []
    used_tokens = 0
    for sym in symbols:
        cost = estimate_tokens(sym.to_prompt_tokens())
        if used_tokens + cost <= budget:
            selected.append(sym)
            used_tokens += cost

    remaining = budget - used_tokens
    for sym in selected:
        if remaining <= 0:
            break
        expanded_cost = estimate_tokens(sym.to_full_source())
        if expanded_cost <= remaining:
            sym.expanded = True
            remaining -= expanded_cost

    return serialize(selected)

The key insight is adaptive compression: different parts of the codebase get different compression ratios based on their relevance to the current task. A function the model is about to modify gets near-full fidelity. A dependency it merely calls gets a signature-only representation. Unrelated code in the same file might be omitted entirely.

Different languages have different redundancy patterns. The compressor uses language-specific rules:

Language Primary Redundancy Compression Strategy Typical Ratio
TypeScript/JavaScript Type annotations, JSDoc, verbose object literals Strip types for internal, keep for exports 4-6x
Python Docstrings, type hints, decorator boilerplate Preserve signatures, compress bodies 3-5x
Go Error checking patterns, context propagation Collapse error guards to throws: lists 5-8x
Rust Trait bounds, lifetimes, boilerplate impl blocks Abstract trait impls to capability lists 6-10x
Java Getters/setters, annotations, imports Collapse CRUD, strip annotations 8-12x

Traditional RAG systems use fixed-size chunks (typically 512-1024 tokens) with overlap. This is fundamentally wrong for code, where a 50-line function is an atomic unit and splitting it across chunks destroys meaning.

The token-first architecture uses semantic chunking based on AST boundaries:

def semantic_chunk(ast: AST, target_tokens: int) -> list[Chunk]:
    """
    Split code into semantically coherent chunks.
    Never splits across function/class boundaries.
    Groups related symbols into single chunks.
    """
    nodes = ast.body  # top-level declarations
    chunks = []
    current_chunk = Chunk()

    for node in nodes:
        node_tokens = estimate_tokens(node)

        if current_chunk.token_count + node_tokens > target_tokens:
            if current_chunk.nodes:
                chunks.append(current_chunk)
                current_chunk = Chunk()

        if isinstance(node, ClassNode) and node_tokens > target_tokens:
            chunks.append(chunk_large_class(node, target_tokens))
            continue

        current_chunk.add(node)

    if current_chunk.nodes:
        chunks.append(current_chunk)

    chunks = merge_small_chunks(chunks, min_tokens=target_tokens // 3)

    return chunks

Each chunk receives a relevance score based on multiple signals:

def compute_relevance(chunk: Chunk, query: str, context: TaskContext) -> float:
    """
    Multi-signal relevance scoring for context selection.
    Returns float in [0.0, 1.0].
    """
    signals = {}

    if any(sym in query for sym in chunk.exported_symbols):
        signals['direct_match'] = 0.9

    distance = shortest_import_path(chunk.module, context.target_module)
    signals['import_proximity'] = max(0, 1.0 - distance * 0.3)

    signals['semantic'] = cosine_similarity(
        embed(chunk.compressed_repr),
        embed(query)
    )

    if chunk.last_modified_hours < 24:
        signals['recency'] = 0.3

    if chunk.file in context.modified_files:
        signals['modified'] = 0.7

    weights = {
        'direct_match': 0.30,
        'import_proximity': 0.25,
        'semantic': 0.20,
        'recency': 0.10,
        'modified': 0.15
    }

    score = sum(signals.get(k, 0) * w for k, w in weights.items())
    return min(1.0, score)

This multi-signal approach is dramatically more accurate than pure vector similarity for code. A function that's semantically dissimilar to the query but sits in the same module as the target file will still score high due to import proximity and modification signals.

Here's where the architecture gets philosophically interesting. The primary complaint about AI coding agents is dishonesty β€” they hallucinate APIs, invent imports, and confidently write code that references nonexistent functions. The token-first architecture addresses this through grounded compression.

Traditional agents hallucinate because of context dilution. When a model sees 80,000 tokens of context, it cannot maintain precise recall of every function signature. Under pressure to produce output, it fills gaps with plausible-sounding hallucinations:

from myapp.utils import parse_config  # DOES NOT EXIST
result = process_data(config)          # process_data doesn't accept config param

The compressed context is exhaustive for interfaces. Every function, class, and exported symbol in the dependency graph appears in the compressed representation with its exact signature. There is no gap for the model to hallucinate into.


Available functions in scope:
  fn parseConfig(path: string) -> AppConfig  [src/config/parser.ts]
  fn loadEnv(file?: string) -> Record<string,string>  [src/config/env.ts]
  fn process_data(input: InputData, opts?: ProcessOpts) -> Output  [src/core/engine.ts]

  NOTE: These are ALL exported functions in the dependency graph.
  If you need a function not listed here, it does not exist.

The architecture explicitly tells the model: this is the complete set of available functions. There is no implicit knowledge, no retrieval gap. If the model needs something that isn't listed, it must ask or admit it doesn't exist.

The system prompt includes an explicit honesty contract:

HONESTY CONTRACT:
1. You are given the COMPLETE list of available functions, types, and modules.
2. If you need something not in this list, state: "Not available in current context"
3. Never invent function names, parameters, or import paths.
4. If uncertain whether something exists, request clarification.
5. Your compressed context is authoritative β€” it reflects the actual codebase state.

This combination of exhaustive compressed interfaces plus explicit honesty instructions dramatically reduces hallucination. In benchmarks, the rate of invented function calls drops from ~12% (traditional RAG) to ~2% (token-first compression).

The token budget allocator is the central control mechanism that balances cost, quality, and completeness. It operates as a real-time optimization problem:

class TokenBudgetAllocator:
    def __init__(self, model_max_context: int, cost_per_token: float, budget_limit: float):
        self.max_context = model_max_context
        self.cost_per_token = cost_per_token
        self.budget_limit = budget_limit

        self.system_prompt_tokens = 2048
        self.safety_margin = 512  # room for model reasoning

        self.available = model_max_context - self.system_prompt_tokens - self.safety_margin

    def allocate(self, components: list[ContextComponent]) -> dict[str, int]:
        """
        Distribute available tokens across context components
        using utility-maximizing allocation.
        """
        for comp in components:
            comp.utility_density = comp.expected_usefulness / comp.token_cost
            comp.compressed_cost = comp.token_cost / comp.compression_ratio

        components.sort(key=lambda c: c.utility_density, reverse=True)

        allocation = {}
        remaining = self.available

        for comp in components:
            comp_allocation = min(comp.compressed_cost, remaining)
            allocation[comp.name] = int(comp_allocation)
            remaining -= comp_allocation

            if remaining <= 0:
                break

        total_cost = sum(
            allocation[c.name] * c.cost_per_token
            for c in components
        )
        if total_cost > self.budget_limit:
            scale = self.budget_limit / total_cost
            for comp in components:
                allocation[comp.name] = int(allocation[comp.name] * scale)

        return allocation

The allocator uses a tiered priority system:

Tier Component Allocation Priority Compression Level
0 Current task description Always full None (raw)
1 Modified files (this session) High Light (2-3x)
2 Direct dependencies Medium-High Moderate (4-6x)
3 Indirect dependencies Medium Heavy (6-10x)
4 Conversation summary Medium Progressive (varies)
5 Project config / conventions Low Maximum (10-15x)
6 Documentation / README Low Maximum or omit

Let's walk through a concrete implementation of the compression pipeline for a TypeScript project.

// src/compression/ast-parser.ts
import { parse, Node, FunctionDeclaration, ClassDeclaration, ExportStatement } from 'typescript';

export interface CompressedSymbol {
  name: string;
  kind: 'function' | 'class' | 'interface' | 'const' | 'enum';
  signature: string;
  dependencies: string[];
  isExported: boolean;
  tokenCostOriginal: number;
  tokenCostCompressed: number;
}

export function extractSymbols(source: string, fileName: string): CompressedSymbol[] {
  const sf = parse(source, { fileName });
  const symbols: CompressedSymbol[] = [];

  sf.forEachChild(node => {
    if (isExportStatement(node)) {
      node.forEachChild(inner => handleExported(inner, symbols));
    } else if (isFunctionDeclaration(node) || isClassDeclaration(node)) {
      handleExported(node, symbols);
    } else if (isVariableStatement(node)) {
      // Check if it's an exported const
      if (hasModifier(node, 'export')) {
        handleExported(node, symbols);
      }
    }
  });

  return symbols;
}

function handleExported(
  node: Node,
  symbols: CompressedSymbol[]
): void {
  if (isFunctionDeclaration(node)) {
    const fn = node as FunctionDeclaration;
    const params = fn.parameters?.map(p => 
      `${p.name.getText()}: ${p.type?.getText() || 'any'}`
    ) || [];
    const returnType = fn.type?.getText() || 'void';

    // Extract called functions
    const deps = extractCalledFunctions(fn.body);

    symbols.push({
      name: fn.name.getText(),
      kind: 'function',
      signature: `${fn.name.getText()}(${params.join(', ')}) -> ${returnType}`,
      dependencies: deps,
      isExported: true,
      tokenCostOriginal: countTokens(fn.getText()),
      tokenCostCompressed: countTokens(formatCompressedSymbol(symbols[symbols.length - 1]))
    });
  }
  // ... handle classes, interfaces, etc.
}

function extractCalledFunctions(body: Node): string[] {
  const called = new Set<string>();

  function visit(node: Node) {
    if (isCallExpression(node)) {
      const expr = node.expression;
      if (isIdentifier(expr)) {
        called.add(expr.text);
      } else if (isPropertyAccessExpression(expr)) {
        called.add(expr.getText());
      }
    }
    node.forEachChild(visit);
  }

  visit(body);
  return [...called];
}
js
// src/compression/engine.ts
import { extractSymbols, CompressedSymbol } from './ast-parser';
import { estimateTokenCount } from './tokenizer';

export interface CompressionResult {
  compressed: string;
  originalTokens: number;
  compressedTokens: number;
  ratio: number;
  symbolsIncluded: number;
  symbolsOmitted: number;
}

export class CompressionEngine {
  private compressionProfile: CompressionProfile;

  constructor(profile: CompressionProfile) {
    this.compressionProfile = profile;
  }

  compress(
    source: string,
    options: {
      budget: number;
      focusSymbols?: string[];
      includeDependencies?: boolean;
    }
  ): CompressionResult {
    const symbols = extractSymbols(source, options.fileName || 'unknown');
    const originalTokens = estimateTokenCount(source);

    // Score symbols by relevance
    const scored = symbols.map(sym => ({
      symbol: sym,
      score: this.scoreSymbol(sym, options.focusSymbols || [])
    }));

    // Sort by score, then greedily allocate budget
    scored.sort((a, b) => b.score - a.score);

    const included: CompressedSymbol[] = [];
    let usedTokens = 0;

    for (const { symbol } of scored) {
      const compressed = this.formatSymbol(symbol);
      const cost = estimateTokenCount(compressed);

      if (usedTokens + cost <= options.budget) {
        included.push(symbol);
        usedTokens += cost;
      }
    }

    // Build output
    const compressed = included.map(s => this.formatSymbol(s)).join('\n');
    const compressedTokens = estimateTokenCount(compressed);

    return {
      compressed,
      originalTokens,
      compressedTokens,
      ratio: originalTokens / compressedTokens,
      symbolsIncluded: included.length,
      symbolsOmitted: symbols.length - included.length
    };
  }

  private scoreSymbol(sym: CompressedSymbol, focus: string[]): number {
    let score = 0;

    if (sym.isExported) score += 0.5;
    if (focus.includes(sym.name)) score += 1.0;
    if (sym.kind === 'function') score += 0.3;
    if (sym.kind === 'interface') score += 0.4;

    // Penalize huge symbols (likely complex internals)
    const complexityPenalty = Math.min(
      0.5,
      sym.tokenCostOriginal / 1000
    );
    score -= complexityPenalty;

    return score;
  }

  private formatSymbol(sym: CompressedSymbol): string {
    const prefix = sym.isExported ? 'export ' : '';
    let result = `${prefix}${sym.kind} ${sym.signature}`;

    if (sym.dependencies.length > 0) {
      result += `\n  // calls: ${sym.dependencies.join(', ')}`;
    }

    return result;
  }
}
js
// src/agent/context-builder.ts
import { CompressionEngine } from '../compression/engine';
import { TokenBudgetAllocator } from '../budget/allocator';

export class ContextBuilder {
  private compressor: CompressionEngine;
  private allocator: TokenBudgetAllocator;
  private dependencyGraph: DependencyGraph;

  async buildContext(
    task: string,
    modifiedFiles: string[],
    model: string
  ): Promise<AssembledContext> {
    const modelConfig = getModelConfig(model);

    // Step 1: Identify relevant modules via dependency graph
    const relevantModules = this.dependencyGraph.getTransitiveDependencies(
      modifiedFiles,
      { depth: 2 }
    );

    // Step 2: Allocate token budget across components
    const budget = this.allocator.allocate({
      task,
      modifiedFiles,
      relevantModules,
      totalBudget: modelConfig.maxContext - 2048 // reserve for system prompt
    });

    // Step 3: Compress each component within its allocation
    const compressed = await Promise.all(
      relevantModules.map(mod => {
        const source = this.readFile(mod.path);
        return this.compressor.compress(source, {
          budget: budget[mod.path] || 500,
          focusSymbols: this.extractFocusSymbols(task, mod.path)
        });
      })
    );

    // Step 4: Assemble final context
    return {
      systemPrompt: this.buildSystemPrompt(compressed),
      totalTokensUsed: compressed.reduce((sum, c) => sum + c.compressedTokens, 0),
      compressionRatio: this.computeOverallRatio(compressed),
      honestyGuarantees: this.generateHonestyContract(compressed)
    };
  }
}

The token-first architecture has been benchmarked across multiple coding tasks. Here are representative results:

Task Traditional Agent (tokens) Token-First (tokens) Savings Quality Delta
Understand 50-file module 78,400 9,200 88% +3.2% accuracy
Refactor auth system 62,100 14,800 76% +5.1% accuracy
Add new API endpoint 45,300 11,400 75% +2.8% accuracy
Debug failing test 38,700 8,900 77% +7.4% accuracy
Cross-module rename 91,200 18,600 79% +4.6% accuracy
Metric Traditional RAG Token-First
Invented function calls 12.3% 1.8%
Wrong import paths 8.7% 0.9%
Incorrect parameter types 15.1% 3.2%
Nonexistent method calls 9.4% 1.1%
Overall hallucination rate 11.4% 1.8%

The compression pipeline adds minimal latency:

This is negligible compared to the 2-30 second LLM inference time it enables by reducing input size.

No architecture is universally superior. The token-first approach has specific failure modes:

When debugging a subtle off-by-one error or a race condition, the model needs to see the exact implementation, not a compressed summary. The architecture must detect these scenarios and expand compression selectively.

If the task requires creating a pattern that doesn't exist in the codebase (e.g., "implement a circuit breaker pattern"), compressed context provides no reference material. The system needs to either fetch external knowledge or operate in a "generation mode" with minimal context.

For tasks like "optimize this for throughput," the model needs to reason about specific implementation details β€” loop structures, allocation patterns, cache behavior. Heavy compression loses these details.

The dependency graph grows superlinearly in monorepos with 10,000+ modules. The graph traversal itself becomes expensive, and the compressed representation of "all dependencies" may exceed the context window even after compression.

// Adaptive compression: detect when full fidelity is needed
function detectCompressionLevel(task: string, context: TaskContext): CompressionLevel {
  if (/(debug|race|off-by-one|deadlock|memory leak)/i.test(task)) {
    return 'minimal'; // Near-full source
  }
  if (/(refactor|rename|add endpoint|new feature)/i.test(task)) {
    return 'aggressive'; // Heavy compression OK
  }
  if (/(optimize|performance|throughput|latency)/i.test(task)) {
    return 'moderate'; // Keep implementation details
  }
  return 'standard'; // Default balanced compression
}

If you want to implement a token-first architecture for your own coding agent, here's the recommended build order:

// A minimal token-first agent in ~100 lines
import { parse } from 'typescript';
import OpenAI from 'openai';

const client = new OpenAI();

async function tokenFirstAgent(task: string, codebase: Map<string, string>) {
  // Step 1: Compress all files to signatures
  const compressed = new Map<string, string>();

  for (const [path, source] of codebase) {
    const sf = parse(source, { fileName: path });
    const symbols: string[] = [];

    sf.forEachChild(node => {
      const text = node.getText();
      // Keep only exported declarations' signatures
      if (text.startsWith('export ')) {
        // Truncate to first 200 chars (signature area)
        const signature = text.split('\n').slice(0, 5).join('\n');
        symbols.push(signature);
      }
    });

    compressed.set(path, symbols.join('\n'));
  }

  // Step 2: Build compressed context
  const context = [...compressed.entries()]
    .map(([path, sigs]) => `// ${path}\n${sigs}`)
    .join('\n\n');

  // Step 3: Call model with compressed context
  const response = await client.chat.completions.create({
    model: 'gpt-4',
    messages: [
      {
        role: 'system',
        content: `
You have access to the following codebase interfaces:

${context}

HONESTY RULES:
- These are the ONLY available functions/classes in scope
- If you need something not listed, say it's not available
- Never invent function names or import paths
- Request clarification if uncertain
        `.  rim()
      },
      { role: 'user', content: task }
    ]
  });

  return response.choices[0].message.content;
}

Larger context windows are necessary but not sufficient. Attention mechanisms have diminishing returns beyond ~30-40K tokens of actual context β€” the model can't maintain precise recall across that much information. Token-first compression achieves better accuracy within a smaller window than uncompressed context in a larger window. The combination of compression + larger windows is ideal, but compression provides independent value even with fixed window sizes.

The architecture works best with languages that have mature parsers (TypeScript, Python, Go, Rust, Java). For languages without standard parsers, you can use tree-sitter as a universal parser backend. The compression quality will be lower (you lose type information), but structural compression still provides 2-3x reduction.

Yes, and arguably more so. Open-source models typically have smaller context windows (8K-32K), making token budget management even more critical. The compressed representations also tend to be more compatible with smaller models because they reduce the cognitive load β€” the model doesn't need to parse and understand verbose source code, just work with clean interface descriptions.

Build an evaluation harness that compares outputs from compressed vs. uncompressed contexts on a held-out set of tasks. Key metrics: (1) Does the output compile? (2) Does it pass existing tests? (3) Does it reference only real functions? (4) Is the implementation semantically correct? If compression causes quality degradation on specific task types, tune the compression level for those scenarios using the adaptive system described above.

The token-first architecture represents a fundamental shift in how we think about AI coding agents. Instead of treating context as a passive container to fill, it treats it as an active resource to manage β€” compressing, prioritizing, and guaranteeing honesty at the architectural level. The 74K-star traction reflects something real: developers are frustrated with agents that hallucinate, waste tokens, and produce unreliable output. A compression-first approach addresses all three problems simultaneously, and the engineering is accessible enough that you can build a working version in under a week.

For more technical deep-dives on AI architecture and developer tooling, explore Tamiz's Insights.

── more in #ai-agents 4 stories Β· sorted by recency
── more on @github 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/compress-before-you-…] indexed:0 read:21min 2026-10-02 Β· β€”