cd /news/large-language-models/claude-code-s-ultrareview-vs-dromeas… · home topics large-language-models article
[ARTICLE · art-104658] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Claude Code's ultrareview vs Dromeas Code Review with LLM council

Dromeas, a code review tool, compared its LLM council-based review against Claude Code's ultrareview on a large, real pull request from the open-source project openclaw. Dromeas's three-model council produced 17 verified findings, while ultrareview found only one nit-level issue, with no overlap. The comparison highlights differences in scope and cost, with Dromeas costing about $5 per run versus $5–25 for ultrareview.

read2 min views2 publishedAug 20, 2026

Originally published at

[dromeas.ai] We heard about Claude Code's ultrareview and got excited — a cloud-run, multi-agent deep review sounded like exactly the kind of thing worth building a workflow around.

So we pointed it at changes in our own repo and compared it against Dromeas code review: three analyzers (quality, security, compliance) cross-checked by an LLM council. Dromeas held up well in that first pass.

That result was interesting enough that we wanted a harder, more neutral test: a large, real, independently-approved pull request from a codebase neither tool had any stake in. So we picked openclaw/openclaw — a public, actively-developed agentic coding tool — and went looking for its biggest recently-merged, genuinely-reviewed PR. That led us to openclaw#124250, 31 files changed, approved by a human reviewer, and we ran the same head-to-head again.

"Preserve ClawHub external source identity and expose only supported actions" — merged, approved by a human reviewer (not a bot self-merge), XL size: 31 files changed, +1,064/−116 lines, spanning the Control UI, macOS, iOS, and Android clients plus the backend that serves them.

The bug it fixes: ClawHub's search API returns each result's source under a nested install.reference

field, but the client code expected a flat installRef

. Every external search result silently fell through to a synthesized @owner/slug

reference — quietly pointing installs at a different publisher's skill than the one the operator actually picked. An identity-spoofing bug in a skill-installation flow, fixed across five client surfaces.

ultrareview: 1 finding, nit severity — a duplicate test assertion in an Android test file, unrelated to the identity-spoofing bug the PR exists to fix.

Dromeas's LLM council: 29 candidate findings raised, 17 kept after cross-verification. Three models (Opus 5, DeepSeek V4 Pro, GPT-5.6 Terra) independently analyzed the diff, then a decider cross-checked each finding. All 12 quality findings and all 5 security findings held up; 12 compliance findings were flagged as duplicates of already-caught security issues or dropped outright, with the report explaining why for each.

None of Dromeas's 17 kept findings overlap with ultrareview's one — not because ultrareview did a bad job reading the diff, but because questions like "is this credential field masked" or "does this action get an audit trail" were never in its scope. Full breakdown, cost comparison (~$5 for the full council run vs. $5–25 typical for ultrareview), and the four findings flagged for manual triage are in the full post →

── more in #large-language-models 4 stories · sorted by recency
── more on @dromeas 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-code-s-ultrar…] indexed:0 read:2min 2026-08-20 ·