cd /news/ai-agents/trace-agent-tool-calls-on-a-free-ser… · home topics ai-agents article
[ARTICLE · art-115049] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Trace Agent Tool Calls on a Free Server: A 10M-Token Debug Loop

A developer detailed a debugging workflow that uses tool-call traces to diagnose failures in LLM agents, leveraging MonkeyCode's free server tier and 10 million free tokens. The approach involves wrapping every tool call with a decorator to log timestamps, arguments, results, and diffs, then analyzing the traces to identify suspicious calls. The developer shared code examples and a seven-step loop to catch common failure patterns.

read4 min views1 publishedAug 29, 2026

At 2 AM, my agent rewrote a config file. Tests passed locally. The deployment failed silently.

The logs showed no error. The agent called read_file

and write_file

. The diff looked correct. But the service crashed.

I needed tool-call traces. Every input, output, and diff. Not just metrics.

MonkeyCode is an open-source project. It offers a free server tier and 10 million free tokens for LLM calls. That's enough to run a trace-analysis pipeline for a real debugging session. Disclosure: This article was prepared as part of MonkeyCode's product outreach.

Here is the workflow I use.

LLM agents hide their reasoning. You see the final patch. You do not see the bad assumption.

Tool calls are the ground truth. They show what the model actually did. Which file it read. Which command it ran. Which value it wrote.

Diffs show the change. Without them, you cannot tell if the agent edited the right lines.

Wrap every tool in a small decorator. Record the timestamp, tool name, arguments, return value, and token cost.

import json, time
from functools import wraps

TRACE_LOG = 'traces.jsonl'

def traced(name):
    def decorator(fn):
        @wraps(fn)
        def wrapper(*args, **kwargs):
            start = time.time()
            result = fn(*args, **kwargs)
            entry = {
                'time': start,
                'tool': name,
                'args': kwargs,
                'result': result,
                'cost_tokens': kwargs.get('max_tokens', 0)
            }
            with open(TRACE_LOG, 'a') as f:
                f.write(json.dumps(entry) + chr(10))
            return result
        return wrapper
    return decorator

This gives you a JSONL file. One line per call. Easy to parse.

MonkeyCode's free server hosts a small parser. Upload the log file first.

scp traces.jsonl user@your-free-server:~/agent-runs/

Then run the analysis script.

python analyze_trace.py traces.jsonl

Capture the file state before and after each call. Then use difflib.unified_diff

.

import difflib

def compute_diff(before, after):
    before_lines = before.splitlines()
    after_lines = after.splitlines()
    return chr(10).join(difflib.unified_diff(
        before_lines, after_lines, lineterm=''
    ))

Store the before

snapshot in each trace entry. Add it to your decorator.

Here is the core logic. It uses MonkeyCode's free model to label each call.

import json, sys

def analyze(path):
    with open(path) as f:
        traces = [json.loads(line) for line in f]
    for t in traces:
        diff = compute_diff(t['before'], t['result'])
        label = monkeycode.complete(
            prompt='Did this diff preserve intent?',
            trace=t,
            diff=diff,
            model='free'
        )
        if label == 'suspicious':
            print(t['time'], t['tool'])

analyze(sys.argv[1])

This script answers one question. “Which tool call likely caused the failure?”

My loop has seven steps.

Repeat until no call gets flagged.

Here is the table I use for every trace.

Field Why it matters
timestamp call order
tool name action taken
arguments model's belief
result actual outcome
diff code change
token cost budget leak

The combination of diff and arguments catches most failures.

I see three patterns again and again.

Symptom Likely cause Check
wrong file changed bad tool argument diff
empty output truncated context result
repeated call misread error timestamp

The debug loop catches all three. It just needs a few good traces.

Last week, an agent ran a rename operation. It called move_file(src, dst)

. The diff showed old content overwritten.

The trace revealed a third argument. The tool schema changed. The agent used an outdated description.

The debug loop caught it in minutes. No paid telemetry needed.

The 10M token pool is finite. Plan how far it goes.

Assume one analysis costs 200 tokens. That gives 50,000 analyses. A few failing runs produce hundreds of traces. The free tier lasts a long time.

The math changes if you call the model per tool call. Batch multiple calls into one prompt. You save tokens.

This approach needs a wrappable agent. Some agents use black-box functions.

The free server may not handle high concurrency. Do not run production telemetry there.

The 10M token allowance is generous. But it is not for high-frequency online summarization. Batch your traces.

My capture script is example code. Adjust it to your own agent framework.

If you need real-time alerting, choose a hosted observability service.

If your agent spawns many parallel tools, the free server might drop requests.

If you store sensitive data, do not upload traces to a remote server. Run a local parser instead.

Tool-call traces bridge logs and outcomes. You do not need a big budget to start.

MonkeyCode's free tier let me test this pipeline. You can try it too. Start with one failing run. Trace it. Fix it. Repeat.

The next time your agent breaks at 2 AM, you will know exactly which call to blame.

── more in #ai-agents 4 stories · sorted by recency
── more on @monkeycode 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/trace-agent-tool-cal…] indexed:0 read:4min 2026-08-29 ·