# Introducing Contextpress: The Python Library That Refactors Your LLM Context

> Source: <https://pub.towardsai.net/introducing-contextpress-the-python-library-that-refactors-your-llm-context-c57965617edb?source=rss----98111c9905da---4>
> Published: 2026-08-10 04:55:05+00:00

By Taha Azizi — AI Engineer | Data Scientist | Tech Writer

Your LLM is not forgetting things because it’s dumb. It’s forgetting because you’re sending it too much.

Over the past two years, nearly every dimension of large language models has improved drastically. But one thing has stubbornly resisted meaningful progress: The Context Window.[contextpress]is trying to address and solve this issue.

Every chat session accumulates noise fast, filler acknowledgements, questions you rephrased three different ways, threads that resolved six turns ago and are still sitting there eating tokens. By turn 30, the model’s attention is spread thin across all of it. The research has a name for what happens next: “lost in the middle.” The useful parts of your history are still there. They’re just buried.

I personally hit this wall constantly. After long writing iteration, the chat starts going sideways, the model repeats itself, forgets a constraint I gave it early on, starts hedging when it wasn’t before. I know the feeling now. At that point I’d manually start a new chat, paste in only what still mattered, and pick up from there. It was a tedious context management done by hand.

That’s the gap [ contextpress ](https://pypi.org/project/contextpress/)fills. It’s a small Python library (Apache 2.0, Python 3.10+) by myself that runs a deterministic NLP pipeline over your message history before it reaches the model , stripping filler, deduplicating near-identical turns, collapsing resolved threads, enforcing a hard token ceiling. Good news and by design, no API key is necessary.

You pass it a list of standard LLM-style message dicts. You get back a shorter list.

Between input and output, five stages run in order:

**Filler** drops low-semantic turns. the “Sure! Happy to help.” and “Got it, thanks!” exchanges that carry zero information.

**Repetition** uses TF-IDF cosine similarity to find near-duplicate turns and keeps only the most recent version.

**Resolution** collapses threads where agreement was clearly reached into a single RESOLVED: system message.

**Recency** extractively compresses older turns while keeping recent ones intact.

**Budget** enforces your hard token limit via tiktoken, always protecting the system prompt and the last exchange.

You don’t have to run all five. Three presets (low, medium, high) control *how aggressively the pipeline runs*, and you can pass an explicit stages= list if you need exact control.

**Tier 1** is what you get by default: fully local, no credentials, deterministic. Ideal for production pipelines where reproducibility matters and you don’t want another model in the critical path.

**Tier 2** is optional. You can attach an LLMBackend (OpenAI, Anthropic, or even local Ollama models) that runs after Tier 1 for semantic deduplication and summarization. If the LLM call fails, the library falls back to Tier 1 output silently; your app doesn't crash.

Let’s start implementing:

```
mkdir contextpress-demo && cd contextpress-demopython -m venv .venv
source .venv/bin/activate      # macOS / Linux.venv\Scripts\activate       # Windows
pip install contextpress
```

First run downloads NLTK data once, silently. Everything after that is local. Open the folder in VS Code (code .) and run the examples below.

```
# save as demo_01_hello.py# run: python demo_01_hello.py
python
from contextpress import ContextManager
messages = [    {"role": "user", "content": "Hello!"},    {"role": "assistant", "content": "Hi there! Sure, I am happy to help you today."},]
out = ContextManager().compress(messages, token_budget=500)print("turns in:", len(messages))print("turns out:", len(out))print("result:", out)
```

Small input, not much to strip. This just confirms everything is wired up correctly and shows you the output format, same dict structure, possibly fewer entries.

This is the one to run first. It builds the kind of history that shows up in real sessions: Filler, repeated questions phrased differently, a resolved thread still sitting in the history.

python

```
# save as demo_02_compress.py# run: python demo_02_compress.py
python
from contextpress import ContextManager
messages = [    {"role": "system", "content": "You are a helpful assistant."},    {"role": "user", "content": "Can you help me with my Python project?"},    {"role": "assistant", "content": "Sure! Happy to help. What do you need?"},    {"role": "user", "content": "I need to read a CSV file."},    {"role": "assistant", "content": "Got it! You can use pandas for that. Sure thing."},    {"role": "user", "content": "How do I read a CSV file in Python?"},    {"role": "assistant", "content": "Use pandas: import pandas as pd; df = pd.read_csv('file.csv')"},    {"role": "user", "content": "OK great. And how do I read a CSV file?"},    {"role": "assistant", "content": "As I mentioned, use pd.read_csv('file.csv')."},    {"role": "user", "content": "Got it. Now I need to filter rows where age > 30."},    {"role": "assistant", "content": "Use df[df['age'] > 30] to filter rows by age."},    {"role": "user", "content": "Perfect, thanks! That works."},    {"role": "assistant", "content": "Great! Glad that helped. Let me know if you need anything else."},    {"role": "user", "content": "One more thing — how do I save the result back to CSV?"},    {"role": "assistant", "content": "Use df.to_csv('output.csv', index=False)."},]
python
def rough_token_count(msgs):    return sum(len(m["content"].split()) * 4 // 3 for m in msgs)
before_tokens = rough_token_count(messages)
cm = ContextManager(type="chat", compression="high")out = cm.compress(messages, token_budget=200)
after_tokens = rough_token_count(out)
php
print(f"Turns:   {len(messages)} -> {len(out)}")print(f"~Tokens: {before_tokens} -> {after_tokens}")print()for m in out:    preview = m["content"][:80].replace("\n", " ")    print(f"  [{m['role']}] {preview}")
```

The repetition stage catches both phrasings of the CSV question and keeps only the later one. Filler drops “Sure! Happy to help.” and “Got it!” as content-free turns. Budget enforces the 200-token ceiling. Turn count and token estimate both drop visibly.

```
# save as demo_03_presets.py# run: python demo_03_presets.py
python
from contextpress import ContextManager
messages = [    {"role": "system", "content": "You are a coding assistant."},    {"role": "user", "content": "What is a list comprehension?"},    {"role": "assistant", "content": "Sure! A list comprehension is a compact way to create lists. Happy to explain."},    {"role": "user", "content": "Can you explain list comprehensions?"},    {"role": "assistant", "content": "Of course! [x*2 for x in range(10)] doubles each number."},    {"role": "user", "content": "Thanks, got it. Now what about dict comprehensions?"},    {"role": "assistant", "content": "Same idea: {k: v for k, v in items.items()} builds a dict."},]
cm = ContextManager(type="chat")
for preset in ("low", "medium", "high"):    out = cm.compress(messages, token_budget=300, compression=preset)    print(f"[{preset:6}] {len(messages)} turns -> {len(out)} turns")
print()
# stages= overrides the preset entirelyout = cm.compress(    messages,    token_budget=300,    stages=["filler", "repetition", "budget"],)print(f"[custom] {len(messages)} turns -> {len(out)} turns")
```

low runs filler and repetition only. medium adds recency. high adds resolution. Pass stages= explicitly and the preset is ignored entirely — you get exactly what you list.

The three types (chat, rag_doc, agent) change how stages behave internally, not just which ones run.

```
# save as demo_04_types.py# run: python demo_04_types.py
python
from contextpress import ContextManager
# chat: standard conversational compressionchat_msgs = [    {"role": "system", "content": "You are a helpful assistant."},    {"role": "user", "content": "What's the capital of France?"},    {"role": "assistant", "content": "Sure! The capital of France is Paris."},    {"role": "user", "content": "And what's the capital of France again?"},    {"role": "assistant", "content": "Paris."},]out = ContextManager(type="chat").compress(chat_msgs, token_budget=300)print(f"chat:    {len(chat_msgs)} -> {len(out)} turns")
# rag_doc: resolution stays off; recency weights by query relevance, not chat orderdoc_chunks = [    {"role": "user", "content": "What does the report say about Q3 revenue?"},    {"role": "assistant", "content": "Chunk 1: Q3 revenue reached $4.2M, up 12% YoY."},    {"role": "assistant", "content": "Chunk 2: The board approved a dividend increase in Q3."},    {"role": "assistant", "content": "Chunk 1 again: Q3 revenue was $4.2M — strong performance."},]out = ContextManager(type="rag_doc").compress(doc_chunks, token_budget=300)print(f"rag_doc: {len(doc_chunks)} -> {len(out)} turns")
# agent: filler rules preserve tool-call turns; resolution triggers on task completionagent_msgs = [    {"role": "system", "content": "You are an agent with tool access."},    {"role": "user", "content": "Run the test suite."},    {"role": "assistant", "content": "Sure! Running tests now."},    {"role": "assistant", "content": "Tool call: run_tests()"},    {"role": "assistant", "content": "Tests passed. Task complete."},    {"role": "user", "content": "Great, now run the tests again."},    {"role": "assistant", "content": "Running tests again."},    {"role": "assistant", "content": "Tool call: run_tests()"},    {"role": "assistant", "content": "Tests passed again."},]out = ContextManager(type="agent").compress(agent_msgs, token_budget=400)print(f"agent:   {len(agent_msgs)} -> {len(out)} turns")
```

For rag_doc, resolution is intentionally off even at high,document chunks don't have conversational agreement patterns. For agent, filler detection is tool-aware and won't drop turns containing tool calls.

No cloud key needed here either — just Ollama running locally.

```
# Prerequisites:# 1. Install Ollama from https://ollama.com# 2. ollama serve# 3. ollama pull llama3.2# 4. pip install ollama# save as demo_05_ollama.py# run: python demo_05_ollama.py
python
from contextpress import ContextManagerfrom contextpress.llm.adapters import OllamaBackend
messages = [    {"role": "system", "content": "You are a helpful assistant."},    {"role": "user", "content": "Tell me about neural networks."},    {"role": "assistant", "content": "Sure! Neural networks are computing systems inspired by the brain. Happy to explain more."},    {"role": "user", "content": "What are neural networks exactly?"},    {"role": "assistant", "content": "They are layered systems of nodes that learn patterns from data via backpropagation."},    {"role": "user", "content": "And what is backpropagation?"},    {"role": "assistant", "content": "Backprop computes gradients layer by layer and updates weights to reduce loss."},]
backend = OllamaBackend(model="llama3.2")cm = ContextManager(    type="chat",    llm_backend=backend,    llm_min_input_chars=200,    llm_max_summary_tokens=512,)
php
out = cm.compress(messages, token_budget=600)print(f"turns: {len(messages)} -> {len(out)}")for m in out:    print(f"  [{m['role']}] {m['content'][:120]}")
```

Tier 2 runs after Tier 1 completes. It deduplicates non-system turns, then summarizes if the transcript is long enough (llm_min_input_chars). System turns are never touched. If Ollama isn't running, Tier 1 output is returned with a warning — nothing breaks.

**Chat apps** — keeps sessions from ballooning between turns without requiring a full memory architecture. **RAG pipelines**,rag_doc mode deduplicates overlapping retrieved chunks before they reach the model, which matters more than most people realize when retrieval hits the same content from different documents. **Agent loops,** agent mode trims the observation-action log while keeping tool calls intact. **Research, **deterministic Tier 1 means your compression step is reproducible run-to-run, and the package includes a CITATION.cff if you need to cite it.

It isn’t a substitute for proper long-term memory in a multi-session agent. It won’t do deep semantic reasoning the way a summarizing LLM would. What it does is eliminate the predictable, mechanical noise that every real conversation accumulates — and it does that with a single import and no external dependencies by default.

The maintainer is direct about this on PyPI: stable for its intended scope, low maintenance cadence, PRs welcome with a 2–4 week review window. If you need a feature urgently, forking is the recommended path. The library is typed (py.typed included), Apache 2.0 licensed, and the codebase is small enough that reading it takes an afternoon.

Bug reports go to [github.com/Taha-azizi/contextpress/issues](https://github.com/Taha-azizi/contextpress/issues). A minimal reproduction case gets you a faster response than a description alone.

```
pip install contextpress
```

Start with Example 2. Swap in a history that looks like one of yours — the long ones where the chat started going sideways — and see what comes out. If something behaves unexpectedly, the issue tracker is the right place.

1- The raw numbers have grown — 128K, 200K, even 1M tokens in some announcements. [But the practical number remain almost the same around 60k to 100k.](https://arxiv.org/pdf/2410.18745) The research on this is consistent and a little uncomfortable: the longer your context, the more the model’s attention dilutes. Facts buried in the middle of a long conversation get partially or fully ignored, a failure mode researchers literally named “[lost in the middle.](https://arxiv.org/abs/2307.03172)” [Retrieval accuracy drops.](https://aclanthology.org/2024.tacl-1.9/) The model starts confusing earlier context with later context.

[Introducing Contextpress: The Python Library That Refactors Your LLM Context](https://pub.towardsai.net/introducing-contextpress-the-python-library-that-refactors-your-llm-context-c57965617edb) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.
