Stop dumping raw message arrays into LLMs and start using structured state with sliding windows.
The most common mistake when deploying AI agents is treating chat history as an append-only log. In early prototypes, appending every user turn, tool response, and raw JSON blob directly into the messages
array works fine.
In production, this pattern collapses after twenty turns. Token usage scales linearly with conversation depth, driving up API latency and inference costs. Worse, models experience "lost-in-the-middle" degradation, forgetting early constraints or crashing altogether due to token limit errors.
messages.append({"role": "user", "content": user_input})
messages.append({"role": "assistant", "content": llm_response})
response = client.chat.completions.create(model="gpt-4o", messages=messages)
Dumping unpruned histories into your LLM turns your database into an expensive latency trap.
The solution is decoupling ephemeral dialogue from persistent conversational state.
Instead of forcing the LLM to re-parse the entire conversation history on every turn to understand what happened ten minutes ago, we split context into two distinct layers:
Incoming User Turn
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββ
β Context Assembler β
β βββββββββββββββββββββββββββββββββββββββ β
β 1. Static System Prompt (Identity) β
β 2. Current State JSON (Facts & Goals) β
β 3. Sliding Window Buffer (Last N Turns) β
βββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
LLM Inference (Bounded & Predictable)
This ensures your token payload stays flat whether a session lasts 3 turns or 300 turns.
Here is a lightweight context manager you can drop directly into your backend service pipeline.
from typing import Any, Dict, List
def build_bounded_context(
system_prompt: str,
raw_history: List[Dict[str, str]],
state_payload: Dict[str, Any],
max_turns: int = 6
) -> List[Dict[str, str]]:
"""Assemble a token-bounded context payload with structured state."""
trimmed_history = raw_history[-max_turns:] if len(raw_history) > max_turns else raw_history
state_injection = {
"role": "system",
"content": f"CURRENT_SESSION_STATE: {state_payload}"
}
return [{"role": "system", "content": system_prompt}, state_injection] + trimmed_history
This pattern provides deterministic context bounds. Your backend guarantees that the context size passed to the provider never exceeds your calculated budget:
If an agent needs to update persistent state (like a shipping address or user intent), extract that state asynchronously or via tool calls, store it in your database, and inject the clean JSON dictionary on the next invocation.