cd /news/ai-agents/parsing-claude-code-session-transcri… · home › topics › ai-agents › article
[ARTICLE · art-149102] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Parsing Claude Code Session Transcripts in Python: Turns, Tool Calls, and Results

A developer documented how Claude Code silently writes each session to an append-only JSONL transcript under ~/.claude/projects/<project-slug>/<session-id>.jsonl, and showed how to parse those files in Python to extract conversation turns, tool_use blocks, and their matching tool_result blocks by id. The approach uses two passes to join calls to results by tool_use_id and skips torn final lines so live sessions can still be analyzed.

by read5 min views7 publishedOct 11, 2026

Every time I run Claude Code, it quietly writes the entire session to disk. No flag, no setup, no instrumentation. The file is already there when I want it. That single fact is the reason I have been able to build small analysis tools on top of my own agent runs without ever touching the agent itself. In this tutorial I want to show you how to read those files and pull real structure out of them: the turns, the tool calls, and the results those calls produced.

I learned this the boring way, by reading a real 34MB transcript line by line, so everything below is grounded in shapes I actually saw on disk rather than shapes I hoped were there.

Claude Code stores one file per session under:

~/.claude/projects/<project-slug>/<session-id>.jsonl

The slug is your project path with slashes turned into dashes. The file is JSONL: one JSON object per line, appended as the session runs. That append-only nature matters later.

Finding them is a one-liner:

from pathlib import Path

def find_sessions(root: Path | None = None) -> list[Path]:
    root = root or (Path.home() / ".claude" / "projects")
    if not root.exists():
        return []
    return sorted(root.glob("**/*.jsonl"))

The first thing you learn reading a live transcript is that the last line can be torn. If Claude Code is mid-write while you read, the final line may be half a JSON object. The right move is not to crash the whole analysis over one bad line, it is to skip it:

import json
from typing import Iterator

def records(path: Path) -> Iterator[dict]:
    with path.open(errors="replace") as fh:
        for line in fh:
            line = line.strip()
            if not line:
                continue
            try:
                yield json.loads(line)
            except json.JSONDecodeError:
                continue

Not every line is a conversation turn. In one real file the type field held values like user, assistant, queue-operation, attachment, file-history-snapshot, mode, and a few others. The two you care about for turns are user and assistant. Each of those carries a message object, and inside it a content field that is a list of typed blocks.

The blocks are where the interesting stuff is. A text block looks like {"type": "text", "text": "..."}. A tool call is a tool_use block. A tool's output is a tool_result block. So the mental model is:

A tool_use block has these keys: type, id, name, input, and sometimes caller. The id is the anchor. When the tool finishes, a later line carries a tool_result block whose tool_use_id matches that id. That pairing, call to result by id, is the whole game.

Here is the core extraction. I make two passes: collect every call and every result keyed by id, then join them. Two passes because results do not always appear in a tidy order relative to calls, and scanning a file twice is cheap compared to getting the pairing subtly wrong.

def tool_calls(path: Path) -> list[dict]:
    uses: dict[str, dict] = {}
    results: dict[str, dict] = {}

    for rec in records(path):
        msg = rec.get("message") or {}
        content = msg.get("content")
        if not isinstance(content, list):
            continue
        for block in content:
            if not isinstance(block, dict):
                continue
            btype = block.get("type")
            if btype == "tool_use":
                uses[block["id"]] = {
                    "name": block.get("name"),
                    "input": block.get("input") or {},
                    "ts": rec.get("timestamp"),
                }
            elif btype == "tool_result":
                tid = block.get("tool_use_id")
                if tid:
                    results[tid] = {
                        "content": block.get("content"),
                        "ts": rec.get("timestamp"),
                        "is_error": bool(block.get("is_error")),
                    }

    joined = []
    for tid, use in uses.items():
        res = results.get(tid, {})
        joined.append({
            "id": tid,
            "tool": use["name"],
            "input": use["input"],
            "output": flatten(res.get("content")),
            "started_at": use.get("ts"),
            "ended_at": res.get("ts"),
            "is_error": res.get("is_error", False),
        })
    return joined

tool_result content is not always a string. Sometimes it is a plain string, and sometimes it is a list of typed blocks like [{"type": "text", "text": "..."}]. If you assume one form, half your results come out as None or as ugly repr strings. Normalize it:

def flatten(content) -> str:
    if isinstance(content, str):
        return content
    if isinstance(content, list):
        parts = []
        for c in content:
            if isinstance(c, dict):
                parts.append(c.get("text") or json.dumps(c))
            else:
                parts.append(str(c))
        return "\n".join(parts)
    return "" if content is None else str(content)

Records carry an ISO-8601 timestamp with a trailing Z. Python's fromisoformat wants an explicit offset, so swap the Z before parsing. Once a call and its result both have timestamps, the difference is your tool latency, no profiler required:

from datetime import datetime

def ts(value: str | None) -> datetime | None:
    if not value:
        return None
    try:
        return datetime.fromisoformat(value.replace("Z", "+00:00"))
    except ValueError:
        return None

Once you have this, small tools fall out of it almost for free. In agentrace I filter tool calls down to the ones named Agent (subagent delegations), pair each delegation's prompt with the report that came back, and time how long each subagent ran. In ctxlens I walk the same turns to see how context accumulates across a session. Neither tool needed any hooks or wrappers around Claude Code. The data was already on disk; I just had to read it correctly.

This is an undocumented, internal format. It is not a stable public API, and it can change between Claude Code versions. New record types can appear, keys can be added, block shapes can shift. That is exactly why the code above is defensive at every step: skip lines it does not understand, guard every .get, and normalize content that comes in more than one shape. Write your parser to tolerate the unexpected rather than to assume today's shape is forever, and it will survive most format drift. When something does change, the fix is usually a print loop over type and block keys to see what is new.

If you want a worked example of this technique applied to subagent observability, my project agentrace does exactly that on top of these same files: github.com/AgentPostmortem/agentrace. Clone it, point it at your own ~/.claude/projects, and you will be reading your agent's history in a couple of minutes.

── more in #ai-agents 4 stories · sorted by recency
── more on @claude code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/parsing-claude-code-…] indexed:0 read:5min 2026-10-11 · —