Remote 200, Local Exit 1. Store the Turn Id on Both Sides. A developer proposes storing a shared turn_id on both remote inference logs and local tool spans to close a gap where an agent's remote log recorded HTTP 200 while the local tool span recorded exit code 1, yet the agent summary still reported the edit landed cleanly. The note offers a join-record schema, a validation checker, four mismatch classes and a debug loop, with code presented as an unexecuted example rather than a benchmark. A coding agent writes a three-file patch during a long run. The remote inference log records HTTP 200 for that turn. The local tool span records exit code 1 for the patch. The agent summary still reports that the edit landed cleanly. No shared identifier connects those three records together. The failure sits in the gap between the two logs. Recent community posts spend their energy on model choice. This note does not rank models or repeat those arguments. It asks whether both logs can name the same attempt. Local traces usually name tools, arguments, and exit codes. Remote logs usually name model calls, status, and latency. Each side keeps its own clock and its own identifier space. A retry writes a second model call under a new request id. The tool span can still point at the first attempt only. A file diff will not show that the join itself broke. This split appears on long agent runs with several tools. It appears again when the runner and the server sit far apart. Delay makes the timelines look unrelated even when they match. This note gives you a join record and a small checker. It also gives you four mismatch classes and a debug loop. The code is a proposed example and was not executed here. It is not a benchmark of any model or any server. It does not quote a token allowance or a hardware spec. Treat fixture counts as labels for the method, not as measurements. Store one join row for every model turn and every tool span. trace id identifies the whole agent run from start to finish. turn id identifies one model request and the matching response. tool call id identifies one tool call, or stays null on model rows. parent turn id names the model turn that requested that tool. attempt is an integer that starts at 1 for each logical turn. side is either model or tool , and no third value is valid. server request id stores the remote id, or null when that id is absent. status is a short code such as ok or error . local mono ns stores local monotonic time, counted in nanoseconds. Do not store raw prompts or tool output in this join file. Store a content hash when you need an equality check later. Keep secrets out of the file before you export it anywhere. A checker that guesses parents will hide emitter bugs. Fail closed when a required field is missing or contradictory. Do not repair a row by matching the nearest timestamp. python Proposed example. Not executed in this draft. from dataclasses import dataclass @dataclass frozen=True class JoinRow: trace id: str turn id: str tool call id: str | None parent turn id: str | None attempt: int side: str server request id: str | None status: str local mono ns: int def validate row: JoinRow - list str : errors: list str = if not row.trace id or not row.turn id: errors.append "missing ids" if row.attempt < 1: errors.append "bad attempt" if row.side not in {"model", "tool"}: errors.append "bad side" if row.side == "tool" and not row.parent turn id: errors.append "tool without parent turn" if row.side == "tool" and not row.tool call id: errors.append "tool without tool id" if row.side == "model" and row.tool call id is not None: errors.append "model row has tool id" if row.local mono ns < 0: errors.append "bad clock" return errors Run validate on every row before the join step starts. Stop the report when any row returns a schema error. A partial join over dirty input looks precise and is not. Index model rows by trace id, turn id, and attempt. Look up each tool span by its parent turn id. Return one class string, and do not attach a score. python Proposed example. Not executed in this draft. def index models rows: list JoinRow - dict tuple str, str, int , JoinRow : indexed = {} for row in rows: if row.side = "model": continue indexed row.trace id, row.turn id, row.attempt = row return indexed def classify tool: JoinRow, models: dict - str: if tool.side = "tool": return "not a tool row" if not tool.parent turn id: return "orphan tool" key = tool.trace id, tool.parent turn id, tool.attempt parent = models.get key if parent is None: return "missing model row" if parent.server request id and tool.server request id: if parent.server request id = tool.server request id: return "request id mismatch" if parent.status = "ok" and tool.status == "ok": return "tool ok after model error" return "joined" Empty parent ids should die in validation, not in classify. The orphan class remains for checkers that skipped that gate. The checker does not rank models and does not score patches. It only reports whether the two logs agree on identity. A joined pair can still contain a wrong edit. Agreement is a precondition, not a verdict on the patch. | Class | Local evidence | Remote evidence | First check | |---|---|---|---| | orphan tool | tool span, empty parent | nothing required | validation should already have failed | | missing model row | parent id is set | no matching model row | dropped log, wrong id, or sampling | | request id mismatch | server id A on the tool | server id B on the model | a retry wrote a new call | | tool ok after model error | tool status is ok | model status is not ok | summary trusted the wrong side | Read the table from the class column toward the check column. Fix the emitter before you rewrite the prompt or the tool. A missing remote row is not proof that the call never ran. Use a tiny fixture so the four classes stay reviewable. Label it as synthetic data inside the test file itself. Do not paste these counts into a status report as results. request id mismatch , missing model row , and joined . php Proposed fixture. Not executed in this draft. def test fixture classes - None: rows = JoinRow "run 1844", "turn a", None, None, 1, "model", "req 1", "ok", 10 , JoinRow "run 1844", "turn a", None, None, 2, "model", "req 2", "ok", 20 , JoinRow "run 1844", "tool row 1", "tool 1", "turn missing", 1, "tool", "req x", "error", 30 , JoinRow "run 1844", "tool row 2", "tool 2", "turn a", 1, "tool", "req 9", "ok", 12 , JoinRow "run 1844", "tool row 3", "tool 3", "turn a", 2, "tool", "req 2", "ok", 21 , assert all validate row == for row in rows models = index models rows got = classify row, models for row in rows if row.side == "tool" assert sorted got == "joined", "missing model row", "request id mismatch" The fixture expects those mismatch classes to appear. The gate command below is for a later run, after the emitter fix. Do not point the gate at this fixture and call the failure a product bug. Proposed commands. Paths are examples only. python join check.py --local runs/trace.jsonl --remote runs/inference.jsonl --out runs/join-report.json python join check.py --local runs/trace.jsonl --remote runs/inference.jsonl --expect-zero orphan tool,request id mismatch The second command is a regression gate for the emitter. It should fail while those two mismatch classes still appear. Keep missing model rows out of that gate until retention is known. Write the checker output as a small JSON object. Mark synthetic runs so nobody quotes them as production data. Include the input file hash so a later replay can detect edits. { "trace id": "run 1844", "fixture": true, "schema errors": 0, "input sha256": "replace-with-real-hash", "counts": { "joined": 1, "orphan tool": 0, "missing model row": 1, "request id mismatch": 1, "tool ok after model error": 0 } } Keep both input files immutable during that replay. Write each report to a new path with the trace hash in the name. That habit stops a later edit from rewriting the evidence. Disclosure: This article was prepared as part of MonkeyCode's product outreach. The operator describes MonkeyCode as an open-source agent project. The same brief says free model access and a free server option exist. This draft does not state a token cap or a model list. It also omits machine size and any uptime promise. Those access details change, and stale numbers mislead planning. Confirm the current limits in the project documentation before a long run. Do not treat a blog post, including this one, as the pricing source. A free server helps this loop in one narrow way. You can place the agent loop on that server instead of the laptop. You still keep the local join file under your control. That file remains the record you hold when remote retention is short. Free access does not remove the two-log problem. A remote log can rotate before you export it. A rate limit can create attempt 2 with no matching tool span. Record attempt on both sides or you will join the wrong call. Echo the server request id into the local row when one is returned. If the server cannot echo the turn id, stop and fix that contract. Use the free option as a second environment, not as an oracle. Run the same schema in both places and diff the class counts. Diff those class-count reports, not the product pages. A lower count means the emitter improved, not that the model improved. Do not upload raw prompts just to fill an empty join field. Redact secrets before any export to a shared or free server. If redaction is unclear, keep the join file on the runner. Skip it when one process already writes one complete log. Skip it when you cannot change the emitter that mints ids. Skip it when the question is model quality rather than trace integrity. A fixed task harness answers quality questions better than this checker. This checker answers whether the two logs describe the same attempt. Mixing those questions produces a confident report about the wrong fault. Also skip the remote half when the server log is heavily sampled. Say that in the report so a reader does not chase empty rows. Local validation can still run alone in that sampled case. Pick one failed run that already has a local JSONL trace. Add the turn id and the parent turn id at the emitter. Classify once, and keep that report beside the trace as a baseline. If you use MonkeyCode's described free server option, copy the request id. Put that id into the same local join row before you export. Check current access limits in the project docs before the second run. Repeat the six steps in that second environment. Stop when the gated classes hit zero, then reread the patch. The join is ready only when the same attempt is named on both sides.