{"slug": "codeaf-coding-agent-running-it-locally-with-ollama-tested", "title": "CodeAF Coding Agent: Running It Locally With Ollama, Tested", "summary": "CodeAF, a Go-based coding agent harness distributed as a single binary, installs via a one-line curl command and connects to Ollama, OpenRouter, Code X, Qwen, and MiniMax, according to hands-on testing published by the project's reviewers. The tool's self-reported benchmark, built on 113 real GitHub issues with ten coding tools given the same model, placed CodeAF's \"developer mode\" highest using DeepSeek v4 Flash, but that figure comes from CodeAF's own GitHub repository rather than an independent third party. Testing found local Ollama support works but is rough — the first prompt timed out until the model finished loading in Ollama — while the task-splitting feature and a \"crew\" system assigning a worker, planner, and checker across different models were the standout capabilities, and every call reportedly routes through DeepSeek v4 regardless of the model displayed.", "body_md": "# CodeAF Coding Agent: Running It Locally With Ollama, Tested\n\nA hands-on look at CodeAF, the Go-based coding agent harness, tested locally with Ollama models and OpenRouter, including its crew and task-split features.\n\n## What is CodeAF?\n\nCodeAF is a Go-based coding agent harness distributed as a single binary. It installs with a one-line curl command, similar to other coding agent tools on the market, and connects to multiple model providers including Ollama, OpenRouter, Code X, Qwen, and MiniMax. Its stated pitch is simple: hand it a piece of work, it splits that work into smaller tasks, runs them, checks the results, and merges whatever passes. That task-splitting and verification loop is the feature CodeAF leans on most heavily in its own marketing.\n\nThe project cites a benchmark built on 113 real GitHub issues, where ten coding tools were given the same model and CodeAF’s own “developer mode” scored highest using DeepSeek v4 Flash. That number comes from CodeAF’s own GitHub repository, not an independent third party, so it’s worth treating as a starting claim rather than a settled fact. Hands-on testing with local Ollama models and OpenRouter is the only way to see whether the tool backs up the chart.\n\n## TL;DR\n\n- **CodeAF installs fast** through a single curl command and ships as one Go binary, with no heavyweight dependencies to set up before first launch.\n- **It defaults to OpenRouter on first run** , which some users may find pushy since it tries to connect there before giving a choice of other providers like Ollama.\n- **Local Ollama support works but is rough** , with at least one case of the first prompt timing out until the model finished loading in Ollama itself.\n- **The task-splitting feature is the standout** , visibly spinning up parallel tasks on screen and letting users watch a log of what each task did in real time.\n- **The “crew” system assigns three roles** (a worker, a planner, and a checker) across different models, including an option to restrict the crew to free, open models only.\n- **Cost tracking is built in** , showing exactly how much a task cost and which model handled it, plus a “redo with stronger model” option for unsatisfying results.\n- **The benchmark numbers CodeAF publishes come from its own repository** , so they describe the project’s self-reported performance rather than independently verified results.\n\n## Other agents ship a demo. Remy ships an app.\n\nReal backend. Real database. Real auth. Real plumbing. Remy has it all.\n\n## How do you install CodeAF?\n\nInstallation is a single curl command, the same pattern used by many modern CLI coding tools. Once it finishes, CodeAF exists as a standalone Go binary. There’s no separate runtime or package manager step involved.\n\nAfter installing, the sensible move is to create a clean working directory for testing rather than running it against a real project right away. Initializing that directory with `git init` gives CodeAF a clean git state to work from, since the tool’s task and merge features interact with version control.\n\nOn first launch, CodeAF tries to connect to OpenRouter automatically. There’s no prompt asking which provider to use first, it just attempts the connection. Interrupting that with Ctrl+C and running `codeaf connect` brings up a provider list including Ollama, Code X, Qwen, and MiniMax. Running `codeaf connect ollama` connects it to a local Ollama instance. From there, `codeaf models` lists what’s available, though the documentation notes that every call ultimately routes through DeepSeek v4 regardless of what’s displayed, which is a point of friction worth knowing about before relying on local-only model selection.\n\n## How does the model and crew selection work?\n\nInside a running CodeAF session, `/model` opens a model picker. Typing the model name works, but simply highlighting an option and pressing Enter doesn’t always select it, requiring the name to be typed out manually. That’s a small but real usability gap for a tool otherwise built around speed.\n\nThe more interesting system is `/crew`. A crew consists of three model roles: a worker that does the actual job, a planner that plans the task, and a checker that reviews the result. Running `crew models open` restricts the crew to free, open models, for example pairing a model like GLM with a local Ollama model. The crew selector also displays pricing for each model combination before committing, which matters if a mix of free and paid models is in play.\n\nOnce a task runs through a configured crew, `/spend` shows exactly what the task cost and which model handled it. If the output isn’t good enough, a “redo with stronger model” option exists, swapping in a more capable (and typically paid) model like Claude Opus, provided an API key for that provider is set.\n\n## What does the task-splitting feature actually look like in practice?\n\nThis is the feature CodeAF is built around, and it’s also the most visibly functional part of the tool in testing. Giving it a single instruction that bundles multiple fixes, for example “fix the add function, fix the off-by-one error, and” a third task, triggers CodeAF to break the request into separate tasks and run them in parallel. During testing, two tasks appeared running side by side on screen, with a live log on the left panel showing what each task was doing as it happened.\n\n- ✕a coding agent\n- ✕no-code\n- ✕vibe coding\n- ✕a faster Cursor\n\nThe one that tells the coding agents what to build.\n\nAdditional tasks can be queued into the same session. There’s also a home view for checking every project running under CodeAF along with its cost, and a tiled “wall” view that shows every open conversation as live tiles at once. For anyone running long or multiple coding sessions in parallel, that visibility is useful for keeping track of what’s actually happening without flipping between terminal windows.\n\nThere’s also a standing-order feature: telling CodeAF something like “never commit straight to main” sets a persistent instruction that applies across the session, useful for long-running or continuous coding tasks where a guardrail needs to stay in effect without repeating it.\n\n## Is CodeAF worth using as a daily coding harness?\n\nAs it stands, CodeAF is young and has noticeable rough edges. The default push toward OpenRouter on first run, the model picker that doesn’t always register a highlighted selection, and at least one timeout when loading a local Ollama model all point to a tool that hasn’t fully smoothed out its first-run experience yet.\n\nThat said, the core ideas behind it are genuinely useful. Watching a single instruction get split into parallel tasks with a visible log of progress is a feature that larger, more established coding agent harnesses don’t always make this transparent. The crew system’s division of labor between a worker, planner, and checker model, combined with real-time cost tracking per task, gives a level of visibility into both execution and spend that’s worth having regardless of which harness ends up being a daily driver.\n\nWhether it replaces an existing setup depends on tolerance for an early-stage tool. For local Ollama users specifically, CodeAF works, but the DeepSeek v4 routing note in its own documentation suggests local models may not always be doing as much of the actual work as the model picker implies.\n\n## Frequently Asked Questions\n\n### What is CodeAF?\n\nCodeAF is a Go-based coding agent harness, distributed as a single binary, that splits incoming coding tasks into smaller subtasks, runs them, checks the results, and merges what passes. It supports multiple model providers including Ollama and OpenRouter.\n\n### Can CodeAF run fully with local Ollama models?\n\nYes, CodeAF connects to a local Ollama instance through `codeaf connect ollama`, and models already pulled in Ollama show up in its model list. However, its documentation notes that calls may still route through DeepSeek v4 behind the scenes, and one tested run timed out until the Ollama model finished loading.\n\n### What is the “crew” feature in CodeAF?\n\nCrew is CodeAF’s multi-model task system. It assigns three roles, a worker that performs the task, a planner that plans it, and a checker that reviews the output, and lets users choose which models fill each role, including an option to restrict the crew to free, open models only.\n\n### How accurate is CodeAF’s benchmark claim?\n\nCodeAF cites a self-run benchmark using 113 real GitHub issues across ten coding tools, where its developer mode scored highest using DeepSeek v4 Flash. That figure comes from the project’s own GitHub repository rather than an independent evaluation, so it should be treated as a self-reported result pending outside verification.\n\n### Does CodeAF show how much a task costs?\n\nYes. Running `/spend` after a task displays its cost and which model handled it. This works whether using paid API-based models through OpenRouter or a mix of free and paid models in a configured crew.", "url": "https://wpnews.pro/news/codeaf-coding-agent-running-it-locally-with-ollama-tested", "canonical_source": "https://www.mindstudio.ai/blog/codeaf-coding-agent-ollama-local/", "published_at": "2026-10-06 00:00:00+00:00", "updated_at": "2026-10-07 14:49:51.294731+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "large-language-models"], "entities": ["CodeAF", "Ollama", "OpenRouter", "DeepSeek v4 Flash", "Code X", "Qwen", "MiniMax", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/codeaf-coding-agent-running-it-locally-with-ollama-tested", "markdown": "https://wpnews.pro/news/codeaf-coding-agent-running-it-locally-with-ollama-tested.md", "text": "https://wpnews.pro/news/codeaf-coding-agent-running-it-locally-with-ollama-tested.txt", "jsonld": "https://wpnews.pro/news/codeaf-coding-agent-running-it-locally-with-ollama-tested.jsonld"}}