# Your agent session transcripts are precious, keep them

> Source: <https://quesma.com/blog/agent-session-transcripts-are-precious/>
> Published: 2026-10-01 08:00:00+00:00

You keep code in git, but agent session transcripts often end up in `/dev/null`. They are proof of work for the tokens you paid for, and lessons you will pay to relearn.

Before observability existed, we used to `ssh` into the server and `grep` the logs after a crash. Most companies are at that stage with agentic coding. They do not collect sessions from Claude Code, Codex, or Cursor, so they never learn from the most valuable data they produce: how their engineers work with agents.

At Quesma we built [Quesma Shipper](https://github.com/QuesmaOrg/quesma-shipper), an open-source collector that gathers sessions into your own Amazon S3 bucket or similar.

## What is an agent session transcript

The transcript is the journal of what the agent did: prompts, thoughts, tool calls, results, and replies, plus the subagents and workflows the session spawned. Researchers call it a trajectory. An example Claude Code session from my computer:

Other tools keep similar data in their own formats.

## AI labs understand the value of session data

Anthropic analyzed [200,000 of its own Claude Code transcripts](https://www.anthropic.com/research/how-ai-is-transforming-work-at-anthropic) to understand how engineers work, and its evals guide ends with the instruction [“Read the transcripts!”](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents). Cursor acknowledges: [“Our approach uses agent sessions as training data.”](https://cursor.com/blog/semsearch) Replit calls trajectories [“a core strategic advantage.”](https://replit.com/blog/how-replit-makes-sense-of-code-at-scale-ai-data)

Consumer plans train on your sessions unless you opt out, and they sell the cheapest tokens: in [our Claude Code pricing analysis](https://quesma.com/blog/claude-code-pricing-for-enterprise/), a $200 Max seat delivered $2,986 of tokens at API rates. No-training clauses and zero data retention live on the business tiers, where the bill multiplies. Meta’s [Contributor tier](https://dev.meta.ai/docs/pricing-rate-limits) charges 12x less per input token “in exchange for permission to use your prompts and completions to train future Meta models.”

This data is worth committing fraud for. Anthropic’s [September 2026 threat report](https://www.anthropic.com/threat-intelligence-report-september-2026) names seven Chinese labs that replayed coding sessions through Claude, rerouted Claude Code users, and bought transcripts from proxy operators who “save exchanges between users and US models without the knowledge or consent of those users.” An [NSA, CISA, and FBI advisory](https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a) the same week described their campaigns as “industrial-scale distillation.”

## What you can learn from transcripts

Transcripts are full of low-hanging fruit to improve your agentic coding:

- What did we spend on each repo, PR, and type of work? Allocate it to R&D, COGS, or S&M, forecast it, and rightsize subscriptions.
- What adds cost without value? Verbose documentation, needless exploration, lengthy license headers, and noisy command output all feed context rot.
- Which habits should become instructions and skills? Tasks resolved at [91% with review and 43% without](https://quesma.com/blog/multi-model-claude-code-refusal-terminal-bench/) after an`AGENTS.md` change.
- How was this change made? See how an agent built a pull request, what it checked (such as an adversarial security review), and what it cost.

Simple misconfigurations waste tokens and time. For example, in our [RTK benchmark](https://quesma.com/blog/does-rtk-make-ai-coding-cheaper/), one DeepSeek attempt repeated the same command 339 times, hitting the same error for 12 minutes. The task still passed, at about 9x the cost. It was an RTK bug, fixed in the next release. Students at [our first hackathon](https://quesma.com/blog/our-first-hackathon/) found a lot of waste in the SWE-chat dataset: agents rerunning tests without changing the code, or rereading the same files after compaction.

### Your intelligence is made of your preferences

There is no single measure of intelligence: people and companies each carry their own context and implicit knowledge. An investor rewards research that weighs every alternative. A game designer rewards trying ten ideas quickly over polishing one. A payments team rewards rigorous tests over speed. Industry benchmarks measure frontier intelligence, not your use case.

Sessions where an engineer corrected or reverted the agent are ready-made evals. They are more accurate and cheaper than synthetic tests, and they let you [hillclimb](https://claude.dev/blog/automating-eval-design-and-hillclimbing/) on models, effort settings, skills, and instructions. Your taste, encoded as evals, is your edge.

### Transcripts help investigate the alien brain

Frontier agents have already gone rogue: OpenAI’s agents [hacked Hugging Face](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/), and Claude [hacked real organizations](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals) during evaluations, which came to light when Anthropic reread its transcripts. Small models fail differently: in our tests, a local model [planted a backdoor on request up to 95% of the time](https://quesma.com/blog/local-llms-security-paradox/), network isolation or not. Whatever goes wrong, the transcript records what the agent was asked and what it did. Do not wait for a disaster, or for AGI: keep your transcripts and watch them for early warning signs.

## Why collecting and storing is hard

There are three ways to get your transcripts.

Ask the vendor: Model providers’ priority is protecting the model from distillation by competitors, and they already hide parts of the session: Anthropic’s full thinking, OpenAI’s compaction, [Codex’s messages to its subagents](https://github.com/openai/codex/commit/5f4d06ef186b896d316620556e561d59206c3ebf). Independent harnesses such as Pi already [feel the pain of sessions that are not portable](https://earendil.com/posts/session-portability/). Official exports, such as Anthropic’s [Compliance API](https://platform.claude.com/docs/en/manage-claude/compliance-sessions), come with Claude Enterprise pricing and many limitations.

Put a proxy in the middle: It needs every tool pointed at it, and some tools refuse. Claude Code’s `/remote-control` [hardcodes its API endpoint](https://x.com/chopra_tejas/status/2096052034906759424), which users of the Headroom proxy asked Anthropic to change. A proxy also misses a lot of data: the local context that gives a prompt its meaning (“fix this code” means very different things in different repositories), and billing data such as how much of the weekly limit was used.

Copy the files from disk: They are not forever. Claude Code deletes them after [30 days by default](https://code.claude.com/docs/en/settings-reference#cleanupperioddays), and its cleanup has [ignored a 99,999-day setting](https://github.com/anthropics/claude-code/issues/41458) and [deleted sessions without warning](https://github.com/anthropics/claude-code/issues/59248); both issues are still open. They are also scattered: agentic coding spread faster than any policy. Teams mix providers and harnesses, company seats and personal subscriptions; shadow IT is the norm, and some engineers run local models such as Qwen.

A naive `rsync` collector makes security worse. Agents can read `.env` files, paste tokens into commands, and print credentials in tool output, and all of it lands in the transcript, which Claude Code stores [in plaintext](https://code.claude.com/docs/en/data-usage). Copy those files to a shared drive and you have a permission leak. Privacy is the other half: talking to an agent feels private. Companies should own this data, scrub it before it leaves the laptop, and decide who may read it, what they see, and how.

## What we open-sourced

Collection should be a commodity. OpenTelemetry did this for observability: one open, vendor-agnostic standard for gathering distributed traces. Agent sessions need the same. That is why we open-sourced [Quesma Shipper](https://github.com/QuesmaOrg/quesma-shipper) under Apache 2.0. It is a ~20 MB Go program, installed once per machine, that scans the session files Claude Code, Codex, and Cursor already write. Nothing to reconfigure, no proxy, no hooks.

Shipper scrubs the credentials and personal data it detects, using deterministic regular-expression and entropy checks based on gitleaks. It encrypts the files using your organization’s public key and uploads them to your own bucket, such as Amazon S3.

Quesma aims to be the custodian of Shipper, not its owner. We are happy to collaborate with anyone who wants to build on it, and to move it to open governance as the community grows. Shipper is enterprise-friendly: signed releases and MDM support. What is ready today is collection: everything from the laptop to your bucket. [Start collecting your coding sessions](https://github.com/QuesmaOrg/quesma-shipper) before they are deleted.

We also plan to open-source an ETL layer that normalizes sessions from different agents into one standard format, and to build our own proprietary analytics on top of it.

Become a Quesma design partner. We are looking for engineering organizations that run Claude Code, Codex, or Cursor across many developers and want to understand and optimize their bill. It is free during the pilot. [Talk to us](https://quesma.com/contact).
