# Compress Before You Prompt: How a 74K-Star Token-First Architecture Is Making AI Coding Agents Smarter, Cheaper, and Actually Honest

> Source: <https://dev.to/tamizuddin/compress-before-you-prompt-how-a-74k-star-token-first-architecture-is-making-ai-coding-agents-d87>
> Published: 2026-10-02 18:02:25+00:00

*Originally published on [tamiz.pro](https://tamiz.pro/insights/token-first-architecture-ai-coding-agents-compression).*

The context window is the new bottleneck. Every AI coding agent you've used — whether it's a local CLI assistant, a cloud IDE copilot, or an autonomous refactor bot — is fundamentally constrained by the same problem: models have finite context, but codebases are infinite in complexity. The result is a brutal engineering tradeoff. Stuff more context in, and you pay exponentially more per token while quality degrades from attention dilution. Stuff less in, and your agent hallucinates APIs, invents dependencies, and confidently writes code that doesn't compile against your actual codebase.

A project that has accumulated over 74,000 GitHub stars in under a year offers a radically different approach. Instead of treating the context window as a fill-to-capacity resource, it treats token budget as a *scarce asset to be managed* — compressing code, documents, and conversation history before they ever reach the model. This "token-first" architecture inverts the traditional prompt engineering paradigm: rather than asking "what should I put in the prompt?", it asks "what is the minimum information the model needs, expressed in the most information-dense form possible?"

The results are striking. Benchmarks show 60-80% token cost reduction on real-world code understanding tasks, with *higher* accuracy than uncompressed baselines. More importantly, hallucination rates drop dramatically — the agent stops inventing function signatures, fake imports, and nonexistent methods because it's working with surgically compressed, high-fidelity representations of actual code.

This article dissects the architecture, explains the compression pipeline in detail, and walks through the engineering decisions that make this approach viable at scale.

Modern LLMs advertise context windows of 128K to 200K tokens. This sounds generous until you realize what actually goes into a coding agent's context:

| Context Component | Typical Token Cost | Notes | 
|---|---|---|
| System prompt + tool definitions | 2,000–5,000 | Fixed overhead | 
| Conversation history | 5,000–20,000 | Grows linearly with turns | 
| Repository structure (file tree) | 3,000–10,000 | For medium projects | 
| Referenced source files | 10,000–50,000 | The biggest variable | 
| Documentation / README | 2,000–8,000 | Often low signal | 
| Linting/test output | 1,000–5,000 | Noisy | 

A single "refactor this module" request can consume 80,000+ tokens of context, leaving almost no room for the model's actual reasoning. Worse, attention mechanisms don't handle uniform token importance well — a 50,000-token context with only 3,000 relevant tokens produces measurably worse outputs than a focused 5,000-token context.

The traditional approach has been *retrieval augmentation*: use a vector database to find "relevant" chunks and stuff them in. This helps, but it's blunt. Vector similarity doesn't understand code semantics — a function named `calculate` might match a completely unrelated `calculate` in another module. And once chunks are in the context window, they're immutable. The model sees them all with equal weight.

The token-first architecture introduces a **compression layer** between the raw codebase and the model. Instead of feeding source files directly into the prompt, every piece of code passes through a multi-stage pipeline that reduces token count while preserving semantic fidelity.

Here's the high-level data flow:

```
┌─────────────────────────────────────────────────────────────────┐
│                    AI Coding Agent Runtime                        │
├─────────────────────────────────────────────────────────────────┤
│                                                                   │
│  ┌──────────┐    ┌──────────────┐    ┌───────────────────────┐  │
│  │  Raw     │───▶│  Compression │───▶│  Token Budget         │  │
│  │  Codebase│    │  Pipeline    │    │  Allocator             │  │
│  └──────────┘    └──────────────┘    └───────────┬───────────┘  │
│       │                    │                      │              │
│       │                    ▼                      ▼              │
│       │            ┌──────────────┐    ┌───────────────────────┐ │
│       │            │  AST Parser  │    │  Context Assembly     │ │
│       │            │  + Indexer   │    │  + Prompt Builder     │ │
│       │            └──────────────┘    └───────────┬───────────┘ │
│       │                                             │             │
│       │                                             ▼             │
│       │                                    ┌───────────────┐     │
│       │                                    │  LLM Backend  │     │
│       │                                    │  (Any Model)  │     │
│       │                                    └───────────────┘     │
│       ▼                                                         │
│  ┌──────────┐                                                   │
│  │  Git /   │                                                   │
│  │  FS      │                                                   │
│  └──────────┘                                                   │
└─────────────────────────────────────────────────────────────────┘
```

The key architectural insight is that compression happens *before* any model call. The pipeline is deterministic, fast (typically under 50ms for a 2,000-line file), and produces representations that are dramatically more information-dense than raw source code.

The compression pipeline consists of four stages, each targeting a different class of redundancy:

Raw source code contains massive redundancy from a semantic perspective. Consider this TypeScript function:

```
// Original: 187 tokens
export async function processUserData(
  userId: string,
  options: ProcessOptions = {
    includeHistory: false,
    maxRecords: 100,
    timeout: 5000
  }: ProcessOptions
): Promise<UserDataResponse> {
  if (!userId) {
    throw new Error('User ID is required');
  }

  const user = await userService.findById(userId);
  if (!user) {
    throw new Error(`User not found: ${userId}`);
  }

  const records = await recordService.getRecent(userId, options.maxRecords);

  return {
    id: user.id,
    name: user.name,
    email: user.email,
    records: records.map(r => ({
      id: r.id,
      timestamp: r.timestamp,
      value: r.value
    })),
    processedAt: new Date().toISOString()
  };
}
```

The AST-based compressor transforms this into a **semantic skeleton**:

```
// Compressed: 42 tokens
fn processUserData(userId: string, options?: ProcessOptions) -> UserDataResponse
  deps: userService.findById, recordService.getRecent
  throws: Error(userId required), Error(user not found)
  returns: {id, name, email, records[{id,timestamp,value}], processedAt}
  defaults: includeHistory=false, maxRecords=100, timeout=5000
```

This preserves every piece of information the model actually needs — the function signature, its dependencies, error conditions, return shape, and default values — while eliminating formatting, boilerplate, and implementation details that are irrelevant for *understanding* the code's interface.

Rather than including full source files for every imported module, the architecture maintains a **compressed dependency graph**. Each module is represented as a node with its exported interfaces, and edges represent import relationships.

```
{
  "module": "src/services/userService.ts",
  "exports": [
    {
      "name": "findById",
      "signature": "(id: string) => Promise<User | null>",
      "sideEffects": ["db.query"]
    },
    {
      "name": "createUser",
      "signature": "(data: CreateUserInput) => Promise<User>",
      "sideEffects": ["db.insert", "eventBus.emit"]
    }
  ],
  "imports": ["src/models/User.ts", "src/db/connection.ts"],
  "tokenCountOriginal": 890,
  "tokenCountCompressed": 95
}
```

When the model needs to understand how `processUserData` works, it sees the compressed signatures of `userService` and `recordService` — not their full implementations. If it needs deeper detail, it can request expansion of specific functions.

Multi-turn coding conversations accumulate enormous context. The architecture applies **progressive summarization** to earlier turns:

```
// Turn 1-3 (raw): ~3,200 tokens
User: Can you refactor the auth module to use JWT instead of sessions?
Agent: I'll start by examining the current auth implementation...
[agent reads 3 files, proposes changes, applies patches]

// After compression: ~340 tokens
[Turns 1-3 Summary] Refactored auth from session-based to JWT.
Modified: auth/middleware.ts (JWT validation), auth/routes.ts (token extraction),
auth/models.ts (added TokenPayload interface). Removed: sessionStore.ts.
Key decision: Using HS256 with 24h expiry, refresh tokens in httpOnly cookies.
```

The compression preserves decisions, file modifications, and architectural choices — the things a model needs to maintain consistency across turns — while discarding exploratory reasoning, intermediate states, and verbose explanations.

This is the most architecturally significant innovation. Rather than letting the model generate freely, the system allocates a **token budget** across different components of the response:

```
{
  "totalBudget": 4096,
  "allocation": {
    "reasoning": 512,
    "code": 2560,
    "explanation": 768,
    "metadata": 256
  }
}
```

The model is instructed to stay within these bounds, and the runtime enforces truncation or regeneration if any section exceeds its allocation. This prevents the common failure mode where a model spends 3,000 tokens explaining context before producing 500 tokens of actual code.

The core of the compression engine is a language-aware AST parser that transforms source code into a compact semantic representation. Let's look at how this works in practice.

``` python
# Simplified representation of the compression pipeline
from dataclasses import dataclass
from typing import Literal

@dataclass
class CompressedFunction:
    name: str
    params: list[tuple[str, str]]  # (name, type)
    return_type: str
    dependencies: list[str]        # external calls
    side_effects: list[str]       # mutations, I/O
    errors: list[str]             # thrown errors
    complexity: Literal["trivial", "moderate", "complex"]

    def to_prompt_tokens(self) -> str:
        """Serialize to minimal token representation."""
        params_str = ", ".join(f"{n}: {t}" for n, t in self.params)
        parts = [f"fn {self.name}({params_str}) -> {self.return_type}"]
        if self.dependencies:
            parts.append(f"  calls: {', '.join(self.dependencies)}")
        if self.side_effects:
            parts.append(f"  mutates: {', '.join(self.side_effects)}")
        if self.errors:
            parts.append(f"  throws: {', '.join(self.errors)}")
        parts.append(f"  complexity: {self.complexity}")
        return "\n".join(parts)

def compress_file(source: str, language: str, budget: int) -> str:
    """
    Compress a source file to fit within token budget.

    Strategy: lossless for public interfaces, lossy for internals.
    Priority order: exported > used-by-current-task > private > unused
    """
    ast = parse_ast(source, language)

    # Phase 1: Extract all symbols with metadata
    symbols = extract_symbols(ast)

    # Phase 2: Score each symbol by relevance
    for sym in symbols:
        sym.score = compute_relevance(sym, current_task_context)

    # Phase 3: Greedy selection within budget
    symbols.sort(key=lambda s: s.score, reverse=True)
    selected = []
    used_tokens = 0
    for sym in symbols:
        cost = estimate_tokens(sym.to_prompt_tokens())
        if used_tokens + cost <= budget:
            selected.append(sym)
            used_tokens += cost

    # Phase 4: If under budget, expand top-scoring symbols
    remaining = budget - used_tokens
    for sym in selected:
        if remaining <= 0:
            break
        expanded_cost = estimate_tokens(sym.to_full_source())
        if expanded_cost <= remaining:
            sym.expanded = True
            remaining -= expanded_cost

    return serialize(selected)
```

The key insight is **adaptive compression**: different parts of the codebase get different compression ratios based on their relevance to the current task. A function the model is about to modify gets near-full fidelity. A dependency it merely calls gets a signature-only representation. Unrelated code in the same file might be omitted entirely.

Different languages have different redundancy patterns. The compressor uses language-specific rules:

| Language | Primary Redundancy | Compression Strategy | Typical Ratio | 
|---|---|---|---|
| TypeScript/JavaScript | Type annotations, JSDoc, verbose object literals | Strip types for internal, keep for exports | 4-6x | 
| Python | Docstrings, type hints, decorator boilerplate | Preserve signatures, compress bodies | 3-5x | 
| Go | Error checking patterns, context propagation | Collapse error guards to `throws:` lists | 5-8x | 
| Rust | Trait bounds, lifetimes, boilerplate impl blocks | Abstract trait impls to capability lists | 6-10x | 
| Java | Getters/setters, annotations, imports | Collapse CRUD, strip annotations | 8-12x | 

Traditional RAG systems use fixed-size chunks (typically 512-1024 tokens) with overlap. This is fundamentally wrong for code, where a 50-line function is an atomic unit and splitting it across chunks destroys meaning.

The token-first architecture uses **semantic chunking** based on AST boundaries:

``` php
def semantic_chunk(ast: AST, target_tokens: int) -> list[Chunk]:
    """
    Split code into semantically coherent chunks.
    Never splits across function/class boundaries.
    Groups related symbols into single chunks.
    """
    nodes = ast.body  # top-level declarations
    chunks = []
    current_chunk = Chunk()

    for node in nodes:
        node_tokens = estimate_tokens(node)

        # Check if adding this node would exceed budget
        if current_chunk.token_count + node_tokens > target_tokens:
            if current_chunk.nodes:
                chunks.append(current_chunk)
                current_chunk = Chunk()

        # Special handling for large classes
        if isinstance(node, ClassNode) and node_tokens > target_tokens:
            chunks.append(chunk_large_class(node, target_tokens))
            continue

        current_chunk.add(node)

    if current_chunk.nodes:
        chunks.append(current_chunk)

    # Merge adjacent small chunks (avoid fragmentation)
    chunks = merge_small_chunks(chunks, min_tokens=target_tokens // 3)

    return chunks
```

Each chunk receives a relevance score based on multiple signals:

``` php
def compute_relevance(chunk: Chunk, query: str, context: TaskContext) -> float:
    """
    Multi-signal relevance scoring for context selection.
    Returns float in [0.0, 1.0].
    """
    signals = {}

    # Signal 1: Direct symbol match (highest weight)
    if any(sym in query for sym in chunk.exported_symbols):
        signals['direct_match'] = 0.9

    # Signal 2: Import graph proximity
    distance = shortest_import_path(chunk.module, context.target_module)
    signals['import_proximity'] = max(0, 1.0 - distance * 0.3)

    # Signal 3: Embedding similarity (semantic)
    signals['semantic'] = cosine_similarity(
        embed(chunk.compressed_repr),
        embed(query)
    )

    # Signal 4: Recent modification (temporal relevance)
    if chunk.last_modified_hours < 24:
        signals['recency'] = 0.3

    # Signal 5: Git blame overlap with modified files
    if chunk.file in context.modified_files:
        signals['modified'] = 0.7

    # Weighted combination
    weights = {
        'direct_match': 0.30,
        'import_proximity': 0.25,
        'semantic': 0.20,
        'recency': 0.10,
        'modified': 0.15
    }

    score = sum(signals.get(k, 0) * w for k, w in weights.items())
    return min(1.0, score)
```

This multi-signal approach is dramatically more accurate than pure vector similarity for code. A function that's semantically dissimilar to the query but sits in the same module as the target file will still score high due to import proximity and modification signals.

Here's where the architecture gets philosophically interesting. The primary complaint about AI coding agents is dishonesty — they hallucinate APIs, invent imports, and confidently write code that references nonexistent functions. The token-first architecture addresses this through **grounded compression**.

Traditional agents hallucinate because of context dilution. When a model sees 80,000 tokens of context, it cannot maintain precise recall of every function signature. Under pressure to produce output, it fills gaps with plausible-sounding hallucinations:

``` python
# Model hallucination example (traditional agent)
from myapp.utils import parse_config  # DOES NOT EXIST
result = process_data(config)          # process_data doesn't accept config param
```

The compressed context is **exhaustive for interfaces**. Every function, class, and exported symbol in the dependency graph appears in the compressed representation with its exact signature. There is no gap for the model to hallucinate into.

```
# What the model actually sees (compressed but complete):

Available functions in scope:
  fn parseConfig(path: string) -> AppConfig  [src/config/parser.ts]
  fn loadEnv(file?: string) -> Record<string,string>  [src/config/env.ts]
  fn process_data(input: InputData, opts?: ProcessOpts) -> Output  [src/core/engine.ts]

  NOTE: These are ALL exported functions in the dependency graph.
  If you need a function not listed here, it does not exist.
```

The architecture explicitly tells the model: *this is the complete set of available functions*. There is no implicit knowledge, no retrieval gap. If the model needs something that isn't listed, it must ask or admit it doesn't exist.

The system prompt includes an explicit honesty contract:

```
HONESTY CONTRACT:
1. You are given the COMPLETE list of available functions, types, and modules.
2. If you need something not in this list, state: "Not available in current context"
3. Never invent function names, parameters, or import paths.
4. If uncertain whether something exists, request clarification.
5. Your compressed context is authoritative — it reflects the actual codebase state.
```

This combination of exhaustive compressed interfaces plus explicit honesty instructions dramatically reduces hallucination. In benchmarks, the rate of invented function calls drops from ~12% (traditional RAG) to ~2% (token-first compression).

The token budget allocator is the central control mechanism that balances cost, quality, and completeness. It operates as a real-time optimization problem:

``` python
class TokenBudgetAllocator:
    def __init__(self, model_max_context: int, cost_per_token: float, budget_limit: float):
        self.max_context = model_max_context
        self.cost_per_token = cost_per_token
        self.budget_limit = budget_limit

        # Fixed allocations
        self.system_prompt_tokens = 2048
        self.safety_margin = 512  # room for model reasoning

        # Dynamic allocation (the interesting part)
        self.available = model_max_context - self.system_prompt_tokens - self.safety_margin

    def allocate(self, components: list[ContextComponent]) -> dict[str, int]:
        """
        Distribute available tokens across context components
        using utility-maximizing allocation.
        """
        # Score each component by utility density (info per token)
        for comp in components:
            comp.utility_density = comp.expected_usefulness / comp.token_cost
            comp.compressed_cost = comp.token_cost / comp.compression_ratio

        # Greedy allocation by utility density
        components.sort(key=lambda c: c.utility_density, reverse=True)

        allocation = {}
        remaining = self.available

        for comp in components:
            # Allocate compressed version first
            comp_allocation = min(comp.compressed_cost, remaining)
            allocation[comp.name] = int(comp_allocation)
            remaining -= comp_allocation

            if remaining <= 0:
                break

        # Validate against cost budget
        total_cost = sum(
            allocation[c.name] * c.cost_per_token
            for c in components
        )
        if total_cost > self.budget_limit:
            # Scale down proportionally, preserving highest-utility components
            scale = self.budget_limit / total_cost
            for comp in components:
                allocation[comp.name] = int(allocation[comp.name] * scale)

        return allocation
```

The allocator uses a tiered priority system:

| Tier | Component | Allocation Priority | Compression Level | 
|---|---|---|---|
| 0 | Current task description | Always full | None (raw) | 
| 1 | Modified files (this session) | High | Light (2-3x) | 
| 2 | Direct dependencies | Medium-High | Moderate (4-6x) | 
| 3 | Indirect dependencies | Medium | Heavy (6-10x) | 
| 4 | Conversation summary | Medium | Progressive (varies) | 
| 5 | Project config / conventions | Low | Maximum (10-15x) | 
| 6 | Documentation / README | Low | Maximum or omit | 

Let's walk through a concrete implementation of the compression pipeline for a TypeScript project.

``` js
// src/compression/ast-parser.ts
import { parse, Node, FunctionDeclaration, ClassDeclaration, ExportStatement } from 'typescript';

export interface CompressedSymbol {
  name: string;
  kind: 'function' | 'class' | 'interface' | 'const' | 'enum';
  signature: string;
  dependencies: string[];
  isExported: boolean;
  tokenCostOriginal: number;
  tokenCostCompressed: number;
}

export function extractSymbols(source: string, fileName: string): CompressedSymbol[] {
  const sf = parse(source, { fileName });
  const symbols: CompressedSymbol[] = [];

  sf.forEachChild(node => {
    if (isExportStatement(node)) {
      node.forEachChild(inner => handleExported(inner, symbols));
    } else if (isFunctionDeclaration(node) || isClassDeclaration(node)) {
      handleExported(node, symbols);
    } else if (isVariableStatement(node)) {
      // Check if it's an exported const
      if (hasModifier(node, 'export')) {
        handleExported(node, symbols);
      }
    }
  });

  return symbols;
}

function handleExported(
  node: Node,
  symbols: CompressedSymbol[]
): void {
  if (isFunctionDeclaration(node)) {
    const fn = node as FunctionDeclaration;
    const params = fn.parameters?.map(p => 
      `${p.name.getText()}: ${p.type?.getText() || 'any'}`
    ) || [];
    const returnType = fn.type?.getText() || 'void';

    // Extract called functions
    const deps = extractCalledFunctions(fn.body);

    symbols.push({
      name: fn.name.getText(),
      kind: 'function',
      signature: `${fn.name.getText()}(${params.join(', ')}) -> ${returnType}`,
      dependencies: deps,
      isExported: true,
      tokenCostOriginal: countTokens(fn.getText()),
      tokenCostCompressed: countTokens(formatCompressedSymbol(symbols[symbols.length - 1]))
    });
  }
  // ... handle classes, interfaces, etc.
}

function extractCalledFunctions(body: Node): string[] {
  const called = new Set<string>();

  function visit(node: Node) {
    if (isCallExpression(node)) {
      const expr = node.expression;
      if (isIdentifier(expr)) {
        called.add(expr.text);
      } else if (isPropertyAccessExpression(expr)) {
        called.add(expr.getText());
      }
    }
    node.forEachChild(visit);
  }

  visit(body);
  return [...called];
}
js
// src/compression/engine.ts
import { extractSymbols, CompressedSymbol } from './ast-parser';
import { estimateTokenCount } from './tokenizer';

export interface CompressionResult {
  compressed: string;
  originalTokens: number;
  compressedTokens: number;
  ratio: number;
  symbolsIncluded: number;
  symbolsOmitted: number;
}

export class CompressionEngine {
  private compressionProfile: CompressionProfile;

  constructor(profile: CompressionProfile) {
    this.compressionProfile = profile;
  }

  compress(
    source: string,
    options: {
      budget: number;
      focusSymbols?: string[];
      includeDependencies?: boolean;
    }
  ): CompressionResult {
    const symbols = extractSymbols(source, options.fileName || 'unknown');
    const originalTokens = estimateTokenCount(source);

    // Score symbols by relevance
    const scored = symbols.map(sym => ({
      symbol: sym,
      score: this.scoreSymbol(sym, options.focusSymbols || [])
    }));

    // Sort by score, then greedily allocate budget
    scored.sort((a, b) => b.score - a.score);

    const included: CompressedSymbol[] = [];
    let usedTokens = 0;

    for (const { symbol } of scored) {
      const compressed = this.formatSymbol(symbol);
      const cost = estimateTokenCount(compressed);

      if (usedTokens + cost <= options.budget) {
        included.push(symbol);
        usedTokens += cost;
      }
    }

    // Build output
    const compressed = included.map(s => this.formatSymbol(s)).join('\n');
    const compressedTokens = estimateTokenCount(compressed);

    return {
      compressed,
      originalTokens,
      compressedTokens,
      ratio: originalTokens / compressedTokens,
      symbolsIncluded: included.length,
      symbolsOmitted: symbols.length - included.length
    };
  }

  private scoreSymbol(sym: CompressedSymbol, focus: string[]): number {
    let score = 0;

    if (sym.isExported) score += 0.5;
    if (focus.includes(sym.name)) score += 1.0;
    if (sym.kind === 'function') score += 0.3;
    if (sym.kind === 'interface') score += 0.4;

    // Penalize huge symbols (likely complex internals)
    const complexityPenalty = Math.min(
      0.5,
      sym.tokenCostOriginal / 1000
    );
    score -= complexityPenalty;

    return score;
  }

  private formatSymbol(sym: CompressedSymbol): string {
    const prefix = sym.isExported ? 'export ' : '';
    let result = `${prefix}${sym.kind} ${sym.signature}`;

    if (sym.dependencies.length > 0) {
      result += `\n  // calls: ${sym.dependencies.join(', ')}`;
    }

    return result;
  }
}
js
// src/agent/context-builder.ts
import { CompressionEngine } from '../compression/engine';
import { TokenBudgetAllocator } from '../budget/allocator';

export class ContextBuilder {
  private compressor: CompressionEngine;
  private allocator: TokenBudgetAllocator;
  private dependencyGraph: DependencyGraph;

  async buildContext(
    task: string,
    modifiedFiles: string[],
    model: string
  ): Promise<AssembledContext> {
    const modelConfig = getModelConfig(model);

    // Step 1: Identify relevant modules via dependency graph
    const relevantModules = this.dependencyGraph.getTransitiveDependencies(
      modifiedFiles,
      { depth: 2 }
    );

    // Step 2: Allocate token budget across components
    const budget = this.allocator.allocate({
      task,
      modifiedFiles,
      relevantModules,
      totalBudget: modelConfig.maxContext - 2048 // reserve for system prompt
    });

    // Step 3: Compress each component within its allocation
    const compressed = await Promise.all(
      relevantModules.map(mod => {
        const source = this.readFile(mod.path);
        return this.compressor.compress(source, {
          budget: budget[mod.path] || 500,
          focusSymbols: this.extractFocusSymbols(task, mod.path)
        });
      })
    );

    // Step 4: Assemble final context
    return {
      systemPrompt: this.buildSystemPrompt(compressed),
      totalTokensUsed: compressed.reduce((sum, c) => sum + c.compressedTokens, 0),
      compressionRatio: this.computeOverallRatio(compressed),
      honestyGuarantees: this.generateHonestyContract(compressed)
    };
  }
}
```

The token-first architecture has been benchmarked across multiple coding tasks. Here are representative results:

| Task | Traditional Agent (tokens) | Token-First (tokens) | Savings | Quality Delta | 
|---|---|---|---|---|
| Understand 50-file module | 78,400 | 9,200 | 88% | +3.2% accuracy | 
| Refactor auth system | 62,100 | 14,800 | 76% | +5.1% accuracy | 
| Add new API endpoint | 45,300 | 11,400 | 75% | +2.8% accuracy | 
| Debug failing test | 38,700 | 8,900 | 77% | +7.4% accuracy | 
| Cross-module rename | 91,200 | 18,600 | 79% | +4.6% accuracy | 

| Metric | Traditional RAG | Token-First | 
|---|---|---|
| Invented function calls | 12.3% | 1.8% | 
| Wrong import paths | 8.7% | 0.9% | 
| Incorrect parameter types | 15.1% | 3.2% | 
| Nonexistent method calls | 9.4% | 1.1% | 
| Overall hallucination rate | 11.4% | 1.8% | 

The compression pipeline adds minimal latency:

This is negligible compared to the 2-30 second LLM inference time it enables by reducing input size.

No architecture is universally superior. The token-first approach has specific failure modes:

When debugging a subtle off-by-one error or a race condition, the model needs to see the *exact* implementation, not a compressed summary. The architecture must detect these scenarios and expand compression selectively.

If the task requires creating a pattern that doesn't exist in the codebase (e.g., "implement a circuit breaker pattern"), compressed context provides no reference material. The system needs to either fetch external knowledge or operate in a "generation mode" with minimal context.

For tasks like "optimize this for throughput," the model needs to reason about specific implementation details — loop structures, allocation patterns, cache behavior. Heavy compression loses these details.

The dependency graph grows superlinearly in monorepos with 10,000+ modules. The graph traversal itself becomes expensive, and the compressed representation of "all dependencies" may exceed the context window even after compression.

```
// Adaptive compression: detect when full fidelity is needed
function detectCompressionLevel(task: string, context: TaskContext): CompressionLevel {
  if (/(debug|race|off-by-one|deadlock|memory leak)/i.test(task)) {
    return 'minimal'; // Near-full source
  }
  if (/(refactor|rename|add endpoint|new feature)/i.test(task)) {
    return 'aggressive'; // Heavy compression OK
  }
  if (/(optimize|performance|throughput|latency)/i.test(task)) {
    return 'moderate'; // Keep implementation details
  }
  return 'standard'; // Default balanced compression
}
```

If you want to implement a token-first architecture for your own coding agent, here's the recommended build order:

``` js
// A minimal token-first agent in ~100 lines
import { parse } from 'typescript';
import OpenAI from 'openai';

const client = new OpenAI();

async function tokenFirstAgent(task: string, codebase: Map<string, string>) {
  // Step 1: Compress all files to signatures
  const compressed = new Map<string, string>();

  for (const [path, source] of codebase) {
    const sf = parse(source, { fileName: path });
    const symbols: string[] = [];

    sf.forEachChild(node => {
      const text = node.getText();
      // Keep only exported declarations' signatures
      if (text.startsWith('export ')) {
        // Truncate to first 200 chars (signature area)
        const signature = text.split('\n').slice(0, 5).join('\n');
        symbols.push(signature);
      }
    });

    compressed.set(path, symbols.join('\n'));
  }

  // Step 2: Build compressed context
  const context = [...compressed.entries()]
    .map(([path, sigs]) => `// ${path}\n${sigs}`)
    .join('\n\n');

  // Step 3: Call model with compressed context
  const response = await client.chat.completions.create({
    model: 'gpt-4',
    messages: [
      {
        role: 'system',
        content: `
You have access to the following codebase interfaces:

${context}

HONESTY RULES:
- These are the ONLY available functions/classes in scope
- If you need something not listed, say it's not available
- Never invent function names or import paths
- Request clarification if uncertain
        `.  rim()
      },
      { role: 'user', content: task }
    ]
  });

  return response.choices[0].message.content;
}
```

Larger context windows are *necessary but not sufficient*. Attention mechanisms have diminishing returns beyond ~30-40K tokens of actual context — the model can't maintain precise recall across that much information. Token-first compression achieves better accuracy *within* a smaller window than uncompressed context in a larger window. The combination of compression + larger windows is ideal, but compression provides independent value even with fixed window sizes.

The architecture works best with languages that have mature parsers (TypeScript, Python, Go, Rust, Java). For languages without standard parsers, you can use tree-sitter as a universal parser backend. The compression quality will be lower (you lose type information), but structural compression still provides 2-3x reduction.

Yes, and arguably more so. Open-source models typically have smaller context windows (8K-32K), making token budget management even more critical. The compressed representations also tend to be more compatible with smaller models because they reduce the cognitive load — the model doesn't need to parse and understand verbose source code, just work with clean interface descriptions.

Build an evaluation harness that compares outputs from compressed vs. uncompressed contexts on a held-out set of tasks. Key metrics: (1) Does the output compile? (2) Does it pass existing tests? (3) Does it reference only real functions? (4) Is the implementation semantically correct? If compression causes quality degradation on specific task types, tune the compression level for those scenarios using the adaptive system described above.

The token-first architecture represents a fundamental shift in how we think about AI coding agents. Instead of treating context as a passive container to fill, it treats it as an active resource to manage — compressing, prioritizing, and guaranteeing honesty at the architectural level. The 74K-star traction reflects something real: developers are frustrated with agents that hallucinate, waste tokens, and produce unreliable output. A compression-first approach addresses all three problems simultaneously, and the engineering is accessible enough that you can build a working version in under a week.

For more technical deep-dives on AI architecture and developer tooling, explore [Tamiz's Insights](https://tamiz.pro/insights).
