# Parsing Claude Code Session Transcripts in Python: Turns, Tool Calls, and Results

> Source: <https://dev.to/royalpinto007/parsing-claude-code-session-transcripts-in-python-turns-tool-calls-and-results-3p3l>
> Published: 2026-10-11 09:30:29+00:00

Every time I run Claude Code, it quietly writes the entire session to disk. No flag, no setup, no instrumentation. The file is already there when I want it. That single fact is the reason I have been able to build small analysis tools on top of my own agent runs without ever touching the agent itself. In this tutorial I want to show you how to read those files and pull real structure out of them: the turns, the tool calls, and the results those calls produced.

I learned this the boring way, by reading a real 34MB transcript line by line, so everything below is grounded in shapes I actually saw on disk rather than shapes I hoped were there.

Claude Code stores one file per session under:

```
~/.claude/projects/<project-slug>/<session-id>.jsonl
```

The slug is your project path with slashes turned into dashes. The file is JSONL: one JSON object per line, appended as the session runs. That append-only nature matters later.

Finding them is a one-liner:

``` php
from pathlib import Path

def find_sessions(root: Path | None = None) -> list[Path]:
    root = root or (Path.home() / ".claude" / "projects")
    if not root.exists():
        return []
    return sorted(root.glob("**/*.jsonl"))
```

The first thing you learn reading a live transcript is that the last line can be torn. If Claude Code is mid-write while you read, the final line may be half a JSON object. The right move is not to crash the whole analysis over one bad line, it is to skip it:

``` php
import json
from typing import Iterator

def records(path: Path) -> Iterator[dict]:
    with path.open(errors="replace") as fh:
        for line in fh:
            line = line.strip()
            if not line:
                continue
            try:
                yield json.loads(line)
            except json.JSONDecodeError:
                # A transcript being appended to while we read it can produce a
                # torn final line. Skipping it lets us still analyse a live session.
                continue
```

Not every line is a conversation turn. In one real file the `type` field held values like `user`, `assistant`, `queue-operation`, `attachment`, `file-history-snapshot`, `mode`, and a few others. The two you care about for turns are `user` and `assistant`. Each of those carries a `message` object, and inside it a `content` field that is a list of typed blocks.

The blocks are where the interesting stuff is. A text block looks like `{"type": "text", "text": "..."}`. A tool call is a `tool_use` block. A tool's output is a `tool_result` block. So the mental model is:

A `tool_use` block has these keys: `type`, `id`, `name`, `input`, and sometimes `caller`. The `id` is the anchor. When the tool finishes, a later line carries a `tool_result` block whose `tool_use_id` matches that `id`. That pairing, call to result by id, is the whole game.

Here is the core extraction. I make two passes: collect every call and every result keyed by id, then join them. Two passes because results do not always appear in a tidy order relative to calls, and scanning a file twice is cheap compared to getting the pairing subtly wrong.

``` php
def tool_calls(path: Path) -> list[dict]:
    uses: dict[str, dict] = {}
    results: dict[str, dict] = {}

    for rec in records(path):
        msg = rec.get("message") or {}
        content = msg.get("content")
        if not isinstance(content, list):
            continue
        for block in content:
            if not isinstance(block, dict):
                continue
            btype = block.get("type")
            if btype == "tool_use":
                uses[block["id"]] = {
                    "name": block.get("name"),
                    "input": block.get("input") or {},
                    "ts": rec.get("timestamp"),
                }
            elif btype == "tool_result":
                tid = block.get("tool_use_id")
                if tid:
                    results[tid] = {
                        "content": block.get("content"),
                        "ts": rec.get("timestamp"),
                        "is_error": bool(block.get("is_error")),
                    }

    joined = []
    for tid, use in uses.items():
        res = results.get(tid, {})
        joined.append({
            "id": tid,
            "tool": use["name"],
            "input": use["input"],
            "output": flatten(res.get("content")),
            "started_at": use.get("ts"),
            "ended_at": res.get("ts"),
            "is_error": res.get("is_error", False),
        })
    return joined
```

`tool_result` content is not always a string. Sometimes it is a plain string, and sometimes it is a list of typed blocks like `[{"type": "text", "text": "..."}]`. If you assume one form, half your results come out as `None` or as ugly repr strings. Normalize it:

``` php
def flatten(content) -> str:
    if isinstance(content, str):
        return content
    if isinstance(content, list):
        parts = []
        for c in content:
            if isinstance(c, dict):
                parts.append(c.get("text") or json.dumps(c))
            else:
                parts.append(str(c))
        return "\n".join(parts)
    return "" if content is None else str(content)
```

Records carry an ISO-8601 `timestamp` with a trailing `Z`. Python's `fromisoformat` wants an explicit offset, so swap the `Z` before parsing. Once a call and its result both have timestamps, the difference is your tool latency, no profiler required:

``` php
from datetime import datetime

def ts(value: str | None) -> datetime | None:
    if not value:
        return None
    try:
        return datetime.fromisoformat(value.replace("Z", "+00:00"))
    except ValueError:
        return None
```

Once you have this, small tools fall out of it almost for free. In agentrace I filter tool calls down to the ones named `Agent` (subagent delegations), pair each delegation's prompt with the report that came back, and time how long each subagent ran. In ctxlens I walk the same turns to see how context accumulates across a session. Neither tool needed any hooks or wrappers around Claude Code. The data was already on disk; I just had to read it correctly.

This is an undocumented, internal format. It is not a stable public API, and it can change between Claude Code versions. New record types can appear, keys can be added, block shapes can shift. That is exactly why the code above is defensive at every step: skip lines it does not understand, guard every `.get`, and normalize content that comes in more than one shape. Write your parser to tolerate the unexpected rather than to assume today's shape is forever, and it will survive most format drift. When something does change, the fix is usually a print loop over `type` and block keys to see what is new.

If you want a worked example of this technique applied to subagent observability, my project agentrace does exactly that on top of these same files: [github.com/AgentPostmortem/agentrace](https://github.com/AgentPostmortem/agentrace). Clone it, point it at your own `~/.claude/projects`, and you will be reading your agent's history in a couple of minutes.
