Taming Context Bloat: How to Scale AI Agent Memory Without Breaking the Token Bank A developer proposes a solution to context bloat in AI agents by decoupling ephemeral dialogue from persistent conversational state. The approach uses a sliding window for chat history and injects structured state JSON directly into the system prompt, ensuring token usage remains bounded regardless of session length. The developer provides a Python context manager to implement this pattern. Stop dumping raw message arrays into LLMs and start using structured state with sliding windows. The most common mistake when deploying AI agents is treating chat history as an append-only log. In early prototypes, appending every user turn, tool response, and raw JSON blob directly into the messages array works fine. In production, this pattern collapses after twenty turns. Token usage scales linearly with conversation depth, driving up API latency and inference costs. Worse, models experience "lost-in-the-middle" degradation, forgetting early constraints or crashing altogether due to token limit errors. The naive anti-pattern: unbounded list growth messages.append {"role": "user", "content": user input} messages.append {"role": "assistant", "content": llm response} 30 turns later: 15,000 tokens wasted on stale tool payloads response = client.chat.completions.create model="gpt-4o", messages=messages Dumping unpruned histories into your LLM turns your database into an expensive latency trap. The solution is decoupling ephemeral dialogue from persistent conversational state . Instead of forcing the LLM to re-parse the entire conversation history on every turn to understand what happened ten minutes ago, we split context into two distinct layers: Incoming User Turn │ ▼ ┌─────────────────────────────────────────┐ │ Context Assembler │ │ ─────────────────────────────────────── │ │ 1. Static System Prompt Identity │ │ 2. Current State JSON Facts & Goals │ │ 3. Sliding Window Buffer Last N Turns │ └─────────────────────────────────────────┘ │ ▼ LLM Inference Bounded & Predictable This ensures your token payload stays flat whether a session lasts 3 turns or 300 turns. Here is a lightweight context manager you can drop directly into your backend service pipeline. python from typing import Any, Dict, List def build bounded context system prompt: str, raw history: List Dict str, str , state payload: Dict str, Any , max turns: int = 6 - List Dict str, str : """Assemble a token-bounded context payload with structured state.""" Enforce strict sliding window on ephemeral chat history trimmed history = raw history -max turns: if len raw history max turns else raw history Inject current state directly as a system-level context injection state injection = { "role": "system", "content": f"CURRENT SESSION STATE: {state payload}" } return {"role": "system", "content": system prompt}, state injection + trimmed history This pattern provides deterministic context bounds. Your backend guarantees that the context size passed to the provider never exceeds your calculated budget: If an agent needs to update persistent state like a shipping address or user intent , extract that state asynchronously or via tool calls, store it in your database, and inject the clean JSON dictionary on the next invocation.