cd /news/developer-tools/introducing-contextpress-the-python-… · home topics developer-tools article
[ARTICLE · art-89937] src=pub.towardsai.net ↗ pub= topic=developer-tools verified=true sentiment=↑ positive

Introducing Contextpress: The Python Library That Refactors Your LLM Context

AI engineer Taha Azizi released Contextpress, a Python library (Apache 2.0, Python 3.10+) that refactors LLM context by running a deterministic NLP pipeline over message history to strip filler, deduplicate near-identical turns, collapse resolved threads, and enforce a hard token ceiling. The library, available on PyPI, requires no API key and offers optional LLM backends for semantic deduplication, with a fallback to local processing if the LLM call fails.

read9 min views1 publishedAug 10, 2026

By Taha Azizi — AI Engineer | Data Scientist | Tech Writer

Your LLM is not forgetting things because it’s dumb. It’s forgetting because you’re sending it too much.

Over the past two years, nearly every dimension of large language models has improved drastically. But one thing has stubbornly resisted meaningful progress: The Context Window.[contextpress]is trying to address and solve this issue.

Every chat session accumulates noise fast, filler acknowledgements, questions you rephrased three different ways, threads that resolved six turns ago and are still sitting there eating tokens. By turn 30, the model’s attention is spread thin across all of it. The research has a name for what happens next: “lost in the middle.” The useful parts of your history are still there. They’re just buried.

I personally hit this wall constantly. After long writing iteration, the chat starts going sideways, the model repeats itself, forgets a constraint I gave it early on, starts hedging when it wasn’t before. I know the feeling now. At that point I’d manually start a new chat, paste in only what still mattered, and pick up from there. It was a tedious context management done by hand.

That’s the gap contextpress fills. It’s a small Python library (Apache 2.0, Python 3.10+) by myself that runs a deterministic NLP pipeline over your message history before it reaches the model , stripping filler, deduplicating near-identical turns, collapsing resolved threads, enforcing a hard token ceiling. Good news and by design, no API key is necessary.

You pass it a list of standard LLM-style message dicts. You get back a shorter list.

Between input and output, five stages run in order:

Filler drops low-semantic turns. the “Sure! Happy to help.” and “Got it, thanks!” exchanges that carry zero information.

Repetition uses TF-IDF cosine similarity to find near-duplicate turns and keeps only the most recent version.

Resolution collapses threads where agreement was clearly reached into a single RESOLVED: system message.

Recency extractively compresses older turns while keeping recent ones intact.

Budget enforces your hard token limit via tiktoken, always protecting the system prompt and the last exchange.

You don’t have to run all five. Three presets (low, medium, high) control how aggressively the pipeline runs, and you can pass an explicit stages= list if you need exact control.

Tier 1 is what you get by default: fully local, no credentials, deterministic. Ideal for production pipelines where reproducibility matters and you don’t want another model in the critical path.

Tier 2 is optional. You can attach an LLMBackend (OpenAI, Anthropic, or even local Ollama models) that runs after Tier 1 for semantic deduplication and summarization. If the LLM call fails, the library falls back to Tier 1 output silently; your app doesn't crash.

Let’s start implementing:

mkdir contextpress-demo && cd contextpress-demopython -m venv .venv
source .venv/bin/activate      # macOS / Linux.venv\Scripts\activate       # Windows
pip install contextpress

First run downloads NLTK data once, silently. Everything after that is local. Open the folder in VS Code (code .) and run the examples below.

python
from contextpress import ContextManager
messages = [    {"role": "user", "content": "Hello!"},    {"role": "assistant", "content": "Hi there! Sure, I am happy to help you today."},]
out = ContextManager().compress(messages, token_budget=500)print("turns in:", len(messages))print("turns out:", len(out))print("result:", out)

Small input, not much to strip. This just confirms everything is wired up correctly and shows you the output format, same dict structure, possibly fewer entries.

This is the one to run first. It builds the kind of history that shows up in real sessions: Filler, repeated questions phrased differently, a resolved thread still sitting in the history.

python

python
from contextpress import ContextManager
messages = [    {"role": "system", "content": "You are a helpful assistant."},    {"role": "user", "content": "Can you help me with my Python project?"},    {"role": "assistant", "content": "Sure! Happy to help. What do you need?"},    {"role": "user", "content": "I need to read a CSV file."},    {"role": "assistant", "content": "Got it! You can use pandas for that. Sure thing."},    {"role": "user", "content": "How do I read a CSV file in Python?"},    {"role": "assistant", "content": "Use pandas: import pandas as pd; df = pd.read_csv('file.csv')"},    {"role": "user", "content": "OK great. And how do I read a CSV file?"},    {"role": "assistant", "content": "As I mentioned, use pd.read_csv('file.csv')."},    {"role": "user", "content": "Got it. Now I need to filter rows where age > 30."},    {"role": "assistant", "content": "Use df[df['age'] > 30] to filter rows by age."},    {"role": "user", "content": "Perfect, thanks! That works."},    {"role": "assistant", "content": "Great! Glad that helped. Let me know if you need anything else."},    {"role": "user", "content": "One more thing — how do I save the result back to CSV?"},    {"role": "assistant", "content": "Use df.to_csv('output.csv', index=False)."},]
python
def rough_token_count(msgs):    return sum(len(m["content"].split()) * 4 // 3 for m in msgs)
before_tokens = rough_token_count(messages)
cm = ContextManager(type="chat", compression="high")out = cm.compress(messages, token_budget=200)
after_tokens = rough_token_count(out)
php
print(f"Turns:   {len(messages)} -> {len(out)}")print(f"~Tokens: {before_tokens} -> {after_tokens}")print()for m in out:    preview = m["content"][:80].replace("\n", " ")    print(f"  [{m['role']}] {preview}")

The repetition stage catches both phrasings of the CSV question and keeps only the later one. Filler drops “Sure! Happy to help.” and “Got it!” as content-free turns. Budget enforces the 200-token ceiling. Turn count and token estimate both drop visibly.

python
from contextpress import ContextManager
messages = [    {"role": "system", "content": "You are a coding assistant."},    {"role": "user", "content": "What is a list comprehension?"},    {"role": "assistant", "content": "Sure! A list comprehension is a compact way to create lists. Happy to explain."},    {"role": "user", "content": "Can you explain list comprehensions?"},    {"role": "assistant", "content": "Of course! [x*2 for x in range(10)] doubles each number."},    {"role": "user", "content": "Thanks, got it. Now what about dict comprehensions?"},    {"role": "assistant", "content": "Same idea: {k: v for k, v in items.items()} builds a dict."},]
cm = ContextManager(type="chat")
for preset in ("low", "medium", "high"):    out = cm.compress(messages, token_budget=300, compression=preset)    print(f"[{preset:6}] {len(messages)} turns -> {len(out)} turns")
print()

low runs filler and repetition only. medium adds recency. high adds resolution. Pass stages= explicitly and the preset is ignored entirely — you get exactly what you list.

The three types (chat, rag_doc, agent) change how stages behave internally, not just which ones run.

python
from contextpress import ContextManager

For rag_doc, resolution is intentionally off even at high,document chunks don't have conversational agreement patterns. For agent, filler detection is tool-aware and won't drop turns containing tool calls.

No cloud key needed here either — just Ollama running locally.

python
from contextpress import ContextManagerfrom contextpress.llm.adapters import OllamaBackend
messages = [    {"role": "system", "content": "You are a helpful assistant."},    {"role": "user", "content": "Tell me about neural networks."},    {"role": "assistant", "content": "Sure! Neural networks are computing systems inspired by the brain. Happy to explain more."},    {"role": "user", "content": "What are neural networks exactly?"},    {"role": "assistant", "content": "They are layered systems of nodes that learn patterns from data via backpropagation."},    {"role": "user", "content": "And what is backpropagation?"},    {"role": "assistant", "content": "Backprop computes gradients layer by layer and updates weights to reduce loss."},]
backend = OllamaBackend(model="llama3.2")cm = ContextManager(    type="chat",    llm_backend=backend,    llm_min_input_chars=200,    llm_max_summary_tokens=512,)
php
out = cm.compress(messages, token_budget=600)print(f"turns: {len(messages)} -> {len(out)}")for m in out:    print(f"  [{m['role']}] {m['content'][:120]}")

Tier 2 runs after Tier 1 completes. It deduplicates non-system turns, then summarizes if the transcript is long enough (llm_min_input_chars). System turns are never touched. If Ollama isn't running, Tier 1 output is returned with a warning — nothing breaks.

Chat apps — keeps sessions from ballooning between turns without requiring a full memory architecture. RAG pipelines,rag_doc mode deduplicates overlapping retrieved chunks before they reach the model, which matters more than most people realize when retrieval hits the same content from different documents. Agent loops, agent mode trims the observation-action log while keeping tool calls intact. **Research, **deterministic Tier 1 means your compression step is reproducible run-to-run, and the package includes a CITATION.cff if you need to cite it.

It isn’t a substitute for proper long-term memory in a multi-session agent. It won’t do deep semantic reasoning the way a summarizing LLM would. What it does is eliminate the predictable, mechanical noise that every real conversation accumulates — and it does that with a single import and no external dependencies by default.

The maintainer is direct about this on PyPI: stable for its intended scope, low maintenance cadence, PRs welcome with a 2–4 week review window. If you need a feature urgently, forking is the recommended path. The library is typed (py.typed included), Apache 2.0 licensed, and the codebase is small enough that reading it takes an afternoon.

Bug reports go to github.com/Taha-azizi/contextpress/issues. A minimal reproduction case gets you a faster response than a description alone.

pip install contextpress

Start with Example 2. Swap in a history that looks like one of yours — the long ones where the chat started going sideways — and see what comes out. If something behaves unexpectedly, the issue tracker is the right place.

1- The raw numbers have grown — 128K, 200K, even 1M tokens in some announcements. But the practical number remain almost the same around 60k to 100k. The research on this is consistent and a little uncomfortable: the longer your context, the more the model’s attention dilutes. Facts buried in the middle of a long conversation get partially or fully ignored, a failure mode researchers literally named “lost in the middle.Retrieval accuracy drops. The model starts confusing earlier context with later context.

Introducing Contextpress: The Python Library That Refactors Your LLM Context was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

── more in #developer-tools 4 stories · sorted by recency
── more on @taha azizi 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/introducing-contextp…] indexed:0 read:9min 2026-08-10 ·