{"slug": "claude-parallel-tool-use-my-agent-loop-crashed-on-388-of-1200-runs", "title": "Claude Parallel Tool Use: My Agent Loop Crashed on 388 of 1,200 Runs", "summary": "A developer's codebase Q&A agent failed 388 of 1,200 runs with HTTP 400 errors because its loop only handled the first tool_use block, while Claude's parallel tool use returns multiple tool calls per turn. The initial fix of stripping orphaned tool_use blocks eliminated crashes but dropped the agent's eval score to 38/50; executing every tool_use block and returning all tool_results together in one user message raised it to 46/50.", "body_md": "My codebase Q&A agent ran 1,200 times in its first week. 388 of those runs died with the same HTTP 400. The other 812 worked fine, which is the worst possible ratio: too low to look like a broken deploy, too high to ignore.\n\nThe cause was one line I'd copied from my own earlier prototype: `next(b for b in resp.content if b.type == \"tool_use\")`. It grabs the first tool call. **Claude parallel tool use** means there often isn't just one.\n\nThen I \"fixed\" it in a way that made the crashes disappear and quietly made the agent dumber. This post covers all of it, plus the numbers from the version that finally worked.\n\n`tool_use` blocks`tool_result` in the `asyncio.gather`), return `is_error: true` instead of dropping them.`disable_parallel_tool_use` when tool order truly matters. It cost me 0.8 extra turns per run.\nIt answers questions about a mid-sized Python monorepo. Three tools: `read_file`, `grep`, `list_dir`. Python SDK, a Claude Sonnet model, a plain `while stop_reason == \"tool_use\"` loop. Nothing exotic.\n\nA typical question is \"where do we retry failed webhook deliveries?\" The model greps, reads two or three files, answers. On a good run that's 3 or 4 round trips.\n\nHere's the loop body I shipped on day one:\n\n```\nresp = client.messages.create(model=MODEL, max_tokens=2048,\n                              tools=TOOLS, messages=messages)\n\nif resp.stop_reason == \"tool_use\":\n    block = next(b for b in resp.content if b.type == \"tool_use\")\n    result = run_tool(block.name, block.input)\n    messages.append({\"role\": \"assistant\", \"content\": resp.content})\n    messages.append({\"role\": \"user\", \"content\": [\n        {\"type\": \"tool_result\", \"tool_use_id\": block.id, \"content\": result}\n    ]})\n```\n\nIt passed every manual test I ran. My manual tests were all simple questions, and simple questions get one tool call at a time.\n\nBecause it can, and it's usually the smart move. When the model already knows it needs `retry.py` and `webhooks/sender.py`, asking for both in one turn saves a full round trip. Parallel tool use is on by default in the Messages API.\n\nSo a response to \"compare how the two webhook senders handle timeouts\" looks like this:\n\n```\ncontent: [\n  text:     \"I'll read both sender implementations.\"\n  tool_use: id=toolu_01A..., name=read_file, input={path: \"webhooks/sender.py\"}\n  tool_use: id=toolu_01B..., name=read_file, input={path: \"webhooks/legacy_sender.py\"}\n]\nstop_reason: \"tool_use\"\n```\n\nMy loop ran `toolu_01A`, appended the full assistant content (both blocks), and sent back one result. The next request failed with:\n\n```\n400 invalid_request_error: messages.4: `tool_use` ids were found without\n`tool_result` blocks immediately after: toolu_01B...\n```\n\nThat's the rule: every `tool_use` in an assistant message needs a `tool_result` with the matching `tool_use_id` in the next user message. No partial credit.\n\nMy first patch was the obvious shortcut. If the API complains about orphaned `tool_use` blocks, remove them before appending the assistant turn:\n\n```\nfirst = next(b for b in resp.content if b.type == \"tool_use\")\nkept = [b for b in resp.content if b.type != \"tool_use\" or b.id == first.id]\nmessages.append({\"role\": \"assistant\", \"content\": kept})\n```\n\nCrashes went to zero. I felt great for about a day.\n\nThen I ran my 50-question eval set (hand-labeled, each with a known correct file and line range). Score: 38/50. The v4 loop described below scores 46/50 on the same set.\n\nReading transcripts made it obvious. From the model's point of view, its own history now said it had asked for one file. So it did one of two things:\n\nEditing the assistant's own past turns is gaslighting your agent. It plans based on what it believes it already did.\n\nExecute every `tool_use` block, then return every result together in a single user message. Here's the loop I run now:\n\n``` python\nimport asyncio\nfrom anthropic import AsyncAnthropic\n\nclient = AsyncAnthropic()\n\nasync def run_one(block):\n    try:\n        out = await run_tool(block.name, block.input)\n        return {\"type\": \"tool_result\", \"tool_use_id\": block.id, \"content\": out}\n    except Exception as e:\n        return {\"type\": \"tool_result\", \"tool_use_id\": block.id,\n                \"content\": f\"{type(e).__name__}: {e}\", \"is_error\": True}\n\nasync def step(messages):\n    resp = await client.messages.create(model=MODEL, max_tokens=2048,\n                                        tools=TOOLS, messages=messages)\n    messages.append({\"role\": \"assistant\", \"content\": resp.content})\n    if resp.stop_reason != \"tool_use\":\n        return resp\n\n    calls = [b for b in resp.content if b.type == \"tool_use\"]\n    results = await asyncio.gather(*(run_one(b) for b in calls))\n\n    assert {r[\"tool_use_id\"] for r in results} == {c.id for c in calls}\n    messages.append({\"role\": \"user\", \"content\": list(results)})\n    return None\n```\n\nThree details matter more than they look.\n\n**One user message, not several.** I briefly tried sending one user message per result. Beyond being awkward, Anthropic's tool use docs warn that splitting results across messages teaches the model to stop making parallel calls in that conversation. You lose the speedup you just paid to support.\n\n**`tool_result` blocks go first.** If you want to add text to that user message (I append a short \"N tool calls remaining in budget\" note), it goes after the results. Text first and the API rejects it.\n\n**Errors are results.** A `FileNotFoundError` on one of three reads used to kill the whole turn. Now the model gets `is_error: true` with the message, and it usually corrects the path on the next turn. That alone recovered 2 of my 50 eval questions.\n\nThe `assert` is cheap insurance. If someone later adds a tool that silently returns `None` and gets filtered out, the loop fails loudly in my process instead of with a 400 from the API.\n\nIn my setup, about one tool turn in seven. After the fix, I logged every turn across another 1,200 runs:\n\n`stop_reason: \"tool_use\"`\n`read_file` across six test fixtures)`grep` + That 14.2% is the number that explains the 32% crash rate. A run has several tool turns, so the odds that at least one of them goes parallel are much higher than the per-turn rate. Your number will differ with your tools and prompts. Read-heavy tools with obvious fan-out invite more parallelism than a single \"run SQL\" tool would.\n\nOnly if your tools have ordering dependencies. Setting `tool_choice={\"type\": \"auto\", \"disable_parallel_tool_use\": True}` makes the model emit at most one `tool_use` per turn, which makes the naive loop technically correct.\n\nI tested it as a fourth variant on the same question set. All four, side by side:\n\n| Loop version | Crashed runs | Eval score | Turns/run | Input tokens/run | Median latency | \n|---|---|---|---|---|---|\n| v1: first block only | 388 / 1,200 | n/a | n/a | n/a | n/a | \n| v2: strip extra blocks | 0 | 38/50 | 4.9 | ~41K | 26.5s | \n| v3: disable_parallel_tool_use | 0 | 45/50 | 4.4 | ~36K | 24.9s | \n| v4: all results, gathered | 0 | 46/50 | 3.6 | ~29K | 18.2s | \n\nThe token column surprised me most. Every extra turn re-sends the whole conversation as input. Fewer round trips means less re-reading, so v4 used about 29% fewer input tokens per run than v2, on top of being faster because the file reads overlap.\n\nWhere I *would* turn parallelism off: an agent with `write_file` and `run_tests`. If the model fires both in one turn, `asyncio.gather` gives you no ordering guarantee, and tests may run against the old file. For side-effecting tools, either disable parallel use or execute the blocks sequentially in the order they appear.\n\nGrep your code for these patterns. Each one is a version of my bug:\n\n`next(b for b in ... if b.type == \"tool_use\")`` resp.content[-1]` or `resp.content[1]` used as \"the tool call\"`tool_use_id` variable that's a single string instead of a list`try/except` around tool execution that `continue` s without appending a result`resp.content` before appending it to history\nIf your SDK version ships a tool runner helper, it handles the matching for you. I still like owning the loop, because then I can log the parallel rate, which turned out to be the most useful number in this whole debugging session.\n\nClaude parallel tool use returns multiple `tool_use` blocks in a single response, and the API requires a `tool_result` for every one of them, matched by `tool_use_id`, in the next user message. Loops that handle only the first block crash with a 400; loops that strip the extra blocks stop crashing but feed the model a false history and give worse answers. Execute every call, return all results together in one user message with failures marked `is_error: true`, and reserve `disable_parallel_tool_use` for tools whose order matters. In my agent that meant zero crashes, 46/50 on eval instead of 38/50, and runs that were 31% faster.\n\n*Written by the developer behind [Preterview](https://preterview.com/en), an interview prep platform.*", "url": "https://wpnews.pro/news/claude-parallel-tool-use-my-agent-loop-crashed-on-388-of-1200-runs", "canonical_source": "https://dev.to/ji_ai/claude-parallel-tool-use-my-agent-loop-crashed-on-388-of-1200-runs-4i28", "published_at": "2026-10-02 09:02:05+00:00", "updated_at": "2026-10-02 09:07:11.632622+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "developer-tools"], "entities": ["Claude", "Anthropic", "Claude Sonnet"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/claude-parallel-tool-use-my-agent-loop-crashed-on-388-of-1200-runs", "markdown": "https://wpnews.pro/news/claude-parallel-tool-use-my-agent-loop-crashed-on-388-of-1200-runs.md", "text": "https://wpnews.pro/news/claude-parallel-tool-use-my-agent-loop-crashed-on-388-of-1200-runs.txt", "jsonld": "https://wpnews.pro/news/claude-parallel-tool-use-my-agent-loop-crashed-on-388-of-1200-runs.jsonld"}}