cd /news/ai-agents/codeaf-crew-mode-tested-do-worker-pl… · home › topics › ai-agents › article
[ARTICLE · art-146891] src=mindstudio.ai ↗ pub= topic=ai-agents verified=true sentiment=· neutral

CodeAF Crew Mode Tested: Do Worker, Planner, Checker Models Deliver?

A hands-on test of CodeAF, a Go-based coding agent harness distributed as a single binary, found that its parallel task splitting works as advertised and that crew mode assigns separate worker, planner, and checker models to a task, but local model setup via Ollama has rough edges. CodeAF's benchmark claims come from its own GitHub repo, based on 113 real GitHub issues rather than SWE-bench, where DeepSeek v4 Flash reportedly scored highest among 10 coding tools tested under CodeAF's "developer mode." The tester concluded the tool shows real promise in parallel task handling and granular built-in cost tracking but has enough interface friction that it is not yet a clear daily-driver replacement for more established harnesses.

by read8 min views1 publishedOct 6, 2026
CodeAF Crew Mode Tested: Do Worker, Planner, Checker Models Deliver?
Image: Mindstudio (auto-discovered)

CodeAF splits coding tasks across worker, planner, and checker models. We tested crew mode, parallel tasks, and cost tracking against its own claims.

What is CodeAF and what does “crew mode” actually do? #

CodeAF is a Go-based coding agent harness distributed as a single binary, installed with a curl command much like other CLI coding tools. Its core pitch is task splitting: you hand it a chunk of work, it breaks that work into smaller tasks, runs them, checks the results, and merges whatever passes. The standout feature is “crew” mode, which assigns three distinct roles to models working on a task: a worker that does the actual implementation, a planner that structures the approach, and a checker that reviews the output before it’s accepted. This is different from a single model doing everything in one context window. In theory, splitting responsibilities this way should catch more bugs and produce more reliable code, since the checker model has one job: find what’s wrong before the result gets merged.

TL;DR #

  • Task splitting works as advertised : when given multiple requests at once (like fixing an add function, an off-by-one error, and a third bug), CodeAF visibly spun up parallel tasks and ran them side by side with a real-time log panel.
  • Crew mode lets you mix and match models by role , assigning a worker, planner, and checker separately, and you can filter to free/open models only with a single command.
  • Cost tracking is built in and granular : a spend command shows exactly what a task cost and which model handled it, useful for anyone trying to control API spend across multiple providers.
  • The benchmark claims come from CodeAF’s own GitHub repo , based on a set of 113 real GitHub issues (not SWE-bench) where DeepSeek v4 Flash reportedly scored highest among 10 coding tools tested under CodeAF’s “developer mode.”
  • Local model setup via Ollama has rough edges : initial connection defaults to OpenRouter unless you explicitly override it, and model selection in the menu requires typing the name since pressing Enter on a highlighted option doesn’t always register.
  • A “redo with stronger model” option exists for escalating a result to a more capable model like Claude Opus, provided you’ve set up the relevant API key.
  • The tool is young , with real promise in its parallel task handling and cost visibility, but enough interface friction that it’s not yet a clear daily-driver replacement for more established harnesses.

Seven tools to build an app. Or just Remy. #

Editor, preview, AI agents, deploy — all in one tab. Nothing to install.

How does CodeAF’s parallel task splitting perform in practice? #

The clearest demonstration of CodeAF’s core claim came from a simple test: asking it to fix three separate issues in one prompt (an add function, an off-by-one bug, and a third fix). Rather than tackling them sequentially in a single thread, CodeAF split the request into multiple tasks and ran them in parallel, visible as separate panes on screen. The tasks completed quickly, and a left-side log let you inspect what each task actually did, in real time.

This is the feature most coding agents don’t do well. Most harnesses process requests linearly within one conversation, meaning independent fixes queue up and wait their turn even though there’s no dependency between them. Watching CodeAF spin up two or more tasks simultaneously and finish them independently is a genuine point in its favor, assuming the underlying model is fast enough to keep up. In this test, switching away from a known “overthinking” model to a faster API-based option made a visible difference in how quickly tasks returned.

What is crew mode and how do worker, planner, and checker models work together? #

Crew mode is accessed with a /crew command, which opens a selection screen showing available models split across roles. You can assign a different model to each role (worker, planner, checker) or use a shortcut to default to open/free models only. The interface shows pricing for each model inline, so you can see at a glance whether you’re about to spend real money or stick to zero-cost local and open options.

In one test, a crew was assembled using GLM and a locally running Ollama model, then given a straightforward task: add a word count function to a string utility file along with a test. The run went through all three roles in sequence (plan, implement, check) before returning a result, and the system correctly logged which model had been used and what it cost.

The practical value here is clear: instead of trusting one model to plan, write, and self-review its own code (which tends to produce blind spots when the model doesn’t catch its own mistakes), you get a second and third model acting as checks. Whether this meaningfully reduces bugs over a large number of tasks wasn’t something a short hands-on test could fully confirm, but the mechanism itself works as described.

Does CodeAF’s cost tracking hold up? #

Yes, and this is one of the more immediately useful features for anyone running multiple models across providers. A spend command reports exactly how much a given task cost, broken down by the model that handled it. In the test run, this showed the cost of a task handled by an open model, tied to the task’s identifier. There’s also a home view (triggered by a shortcut) that aggregates spend and status across every project, and a “wall” view showing all open conversations as live tiles, which is useful if you’re running several tasks or projects concurrently and want a single dashboard rather than switching between terminal windows.

For developers juggling free local models against paid API calls, this kind of transparent, per-task cost accounting removes a lot of guesswork that other coding agents leave opaque.

How do CodeAF’s benchmark claims hold up to scrutiny? #

CodeAF’s GitHub repository includes a chart presenting its own benchmark results: a set of 113 real GitHub issues, not the widely used SWE-bench, tested across 10 coding tools using the same underlying model. According to this chart, DeepSeek v4 Flash scored highest when run through CodeAF’s own “developer mode” configuration. These are numbers CodeAF itself produced and published, not results from an independent third party.

That distinction matters. A benchmark published by the tool’s own developers, using a custom issue set rather than a standard public benchmark, should be treated as a marketing claim until verified independently. It’s not inherently false, but it’s also not something to take at face value just because it’s presented as a chart on a GitHub page. The honest takeaway is that CodeAF’s parallel task handling and crew mode are features you can verify yourself in minutes, while the specific performance numbers claiming DeepSeek v4 Flash’s superiority require more rigorous, independent testing to confirm.

Is CodeAF worth using as a local coding agent today? #

For local Ollama-based use, CodeAF currently has friction. On first run, it defaults to trying to connect to OpenRouter rather than asking which provider you want, which is an odd default for a tool that supports local model connections. Connecting Ollama explicitly requires a connect command, and the model selection menu doesn’t always behave as expected (highlighting a model and pressing Enter sometimes fails to select it, requiring the model name to be typed out manually). There was also at least one instance of a request timing out on first use, requiring the local model to be reloaded before it responded. None of these are fatal flaws, but they add up to a tool that still feels early. The core ideas, splitting tasks into parallel runs, assigning distinct roles to different models, and tracking spend transparently, are genuinely useful and not common across competing coding harnesses. Whether CodeAF becomes a reliable daily driver depends on whether these rough edges get fixed as the project matures.

Frequently Asked Questions #

What does CodeAF’s crew mode do differently from a single-model coding agent?

Crew mode splits a coding task across three roles, a worker model that writes the code, a planner model that structures the approach, and a checker model that reviews the result before accepting it, rather than relying on one model to do all three jobs in the same context.

Can CodeAF run entirely on local models?

Yes, it supports connecting to Ollama for local model use, though the default setup flow initially tries to connect to OpenRouter and requires an explicit command to switch to a local provider.

Are CodeAF’s benchmark claims independently verified?

No. The published benchmark, based on 113 real GitHub issues and claiming DeepSeek v4 Flash as the top performer, comes from CodeAF’s own GitHub repository rather than an independent or widely adopted benchmark like SWE-bench.

Does CodeAF show how much each task costs?

Other agents ship a demo. Remy ships an app. #

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

Yes, a built-in spend command reports the cost of individual tasks along with which model handled them, and a home view aggregates spend across all active projects.

Is CodeAF ready to replace other coding agent harnesses?

Not clearly yet. It has promising features like parallel task splitting and crew-based model assignment, but interface issues like inconsistent model selection and default provider connections suggest it’s still an early-stage tool.

── more in #ai-agents 4 stories · sorted by recency
openalternative.co · · #ai-agents
IronClaw
── more on @codeaf 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/codeaf-crew-mode-tes…] indexed:0 read:8min 2026-10-06 · —