cd /news/artificial-intelligence/big-pickle-on-swe-atlas-codebase-qna · home topics artificial-intelligence article
[ARTICLE · art-98361] src=github.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Big Pickle on SWE Atlas – Codebase QnA

OpenCode Zen's free stealth model big-pickle resolved 50.8% (63/124) of Scale AI's SWE Atlas Codebase QnA benchmark tasks on 2026-08-11, outperforming all official Mini-SWE-Agent scaffold entries and both Codex-scaffold GPT models, trailing only Opus 5 (63.17%) and Opus 4.8 (57.26%) on the official leaderboard. The run used the official harness, task data, and judge model, with zero command timeouts or OOM kills, but Scale did not verify the result and the model's identity is unconfirmed.

read5 min views1 publishedAug 16, 2026
Big Pickle on SWE Atlas – Codebase QnA
Image: Michielbdejong (auto-discovered)

Task Resolve Rate: 50.8% (63/124) big-pickle, the free stealth model on OpenCode Zen, evaluated on

Scale AI's SWE AtlasCodebase QnA benchmark using the mini-swe-agent scaffold.

Run on 2026-08-11 with the official open-source harness, task data, and judge model.

Against the official SWE Atlas QnA leaderboard (updated 2026-07-28):

Model (scaffold) Task Resolve Rate
Opus 5 (Claude Code, xHigh) 63.17
Opus 4.8 (Claude Code, xHigh) 57.26
big-pickle (Mini-SWE-Agent) — this run
50.81
GLM 5.2 (Mini-SWE-Agent) 48.12
GPT-5.6-Sol (Codex, xHigh) 46.00
GPT 5.5 (Codex, xHigh) 45.43

Within the Mini-SWE-Agent scaffold class — the apples-to-apples comparison — this run outscores every entry on the official leaderboard, and it also tops the Codex-scaffold GPT entries. Only the two Claude models running on their native Claude Code scaffold score higher. Note the caveats below before treating this as a leaderboard-equivalent number.

Language Resolved Rate
TypeScript 18/31 58.1%
Python 16/29 55.2%
Go 19/38 50.0%
C 10/26 38.5%
Category Resolved Rate
Code Onboarding 17/28 60.7%
Architecture & system design 23/44 52.3%
Root-cause analysis 17/37 45.9%
Security 5/11 45.5%
API & library usage / integration 1/4 25.0%

Everything follows Scale's published protocol as closely as budget allowed:

Tasks: all 124 Codebase QnA tasks fromscaleapi/SWE-Atlas(Apache-2.0), unmodified — including Scale's shippedmswea_qa_config.yaml

agent configuration (system/instance templates,step_limit: 250

).Harness:Harborv0.18.0 with Modal sandboxes, per the SWE-Atlas README.** Scaffold:**mini-swe-agentpinned to 2.4.6 — the same minimal bash-only scaffold Scale uses for non-first-party models on the leaderboard.Model:big-pickle

via OpenCode Zen's OpenAI-compatible endpoint (https://opencode.ai/zen/v1

), litellm routeopenai/big-pickle

. Total consumption:674M input / 4.3M output tokens, at $0(the model is free during its stealth period).** Judge:**claude-opus-4-5-20251101

— the exact judge model Scale specifies — accessed through Anthropic's OpenAI-compatible endpoint (https://api.anthropic.com/v1

) withEVAL_MODEL

overridden to the bare Anthropic model ID.Scoring: the benchmark's own rubric-based verifier, unmodified. A task resolves only if every scored must-have rubric passes.

Read these before quoting the number:

Single trial per task(-k 1

). The official protocol runs 3 trials and reports the mean. At n=124, the single-trial standard error is ≈ ±4.5 points — comparable to the leaderboard's own reported error bars (±5).Reduced sandbox resources. Tasks declare 16 CPU / 16 GB; this run used 4 CPU / 8 GB to fit a personal budget. Slower command execution can only depress an agent's score (via command timeouts or OOM kills), not inflate it. Empirically it appears to have had no effect here: a scan of all 124 agent trajectories foundzero command timeouts and zero exit-137 kills— no command ever hit the 900s ceiling or the memory limit.** Self-reported.**Scale did not run or verify this evaluation. The full per-task verifier logs in this repo allow independent auditing, and the run is reproducible from the configs here plus the public SWE-Atlas repo.Model identity unknown. big-pickle is officially unconfirmed; leaked provider errors and API response signatures suggest it is currently served by DeepSeek infrastructure. The underlying model may change without notice, so this result is a snapshot of whatever was behind the alias on 2026-08-11.Data exposure. OpenCode states that prompts to big-pickle during its free period may be used to improve the model. The benchmark's task content (already public, canary-marked by Scale) was necessarily sent to that endpoint.Two resolved tasks had unscored rubrics. Ontask-...ba9ad

(5 of 11 rubrics) andtask-...baa1d

(1 rubric), the judge returned unparseable output through all 8 retries; the benchmark's verifier excludes unscored rubrics from the pass computation by design. Treating unscored-as-fail instead gives a strict-lower-bound of61/124 = 49.2%— still above every Mini-SWE-Agent leaderboard entry. All verifier logs are included so you can apply either convention.

git clone https://github.com/scaleapi/SWE-Atlas && cd SWE-Atlas
git clone --branch v0.18.0 --depth 1 https://github.com/laude-institute/harbor.git
uv tool install ./harbor --with modal && uv tool install modal && modal setup

./preflight.sh
bash run_config/qa/big-pickle_smoke.sh    # 3-task smoke test first
bash run_config/qa/big-pickle_miniswe.sh  # full 124-task run

Hard-won gotchas the configs already handle:

Do not pass— Harbor silently switches mini-swe-agent to the OpenAI Responses API, which chat-completions-only endpoints like Zen don't serve.--ak reasoning_effort

with anopenai/

-prefixed modelKeep agent and judge credentials separate. The judge reads hostOPENAI_API_KEY

/OPENAI_API_BASE

(via each task's[verifier.env]

); the agent's Zen credentials go through--ae

per-agent overrides.Pass secrets to Harbor redacts literal secrets to--ae

as${VAR}

templates, not literals.****

when persisting job state, which breaksharbor job resume

with instant 401s. Templates round-trip and re-resolve from the host env.- Expect a few % of trials to die to Modal Failed to read exec stdio stream

errors;harbor job resume -f <ErrorType> ...

re-runs them cleanly.

Approximate cost for the full QnA run: ~$70 of Modal compute (at reduced sandbox resources; roughly 2–3× that at the declared 16 CPU/16 GB), ~$25 of Anthropic API for judging, $0 for the model.

results/per_task_results.csv

— task ID, category, language, resolved, aggregate rubric score, rubrics passed/totalresults/summary.json

— headline numbers and breakdownsresults/verifier_logs/

— the judge's full per-rubric output for every task (audit trail). Notes like(flipped from raw=0)

are the benchmark's own shipped verifier logic (evaluate_answer.py

inverts rubrics marked negative-polarity), not post-hoc re-scoring.run_config/

— the exact Harbor run scripts used (QnA smoke + full, plus untested Test Writing / Refactoring variants)preflight.sh

— endpoint/auth checks for both the model and the judge

SWE Atlasbenchmark © Scale AI, Apache-2.0 — paper:arXiv:2605.08366. Per the authors' request, please treat SWE Atlas as a held-out signal of progress rather than a training target.Harbor(Laude Institute) andmini-swe-agent(SWE-agent team).- big-pickle is served by OpenCode Zen.

Evaluation configs and results in this repo are MIT-licensed.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @opencode zen 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/big-pickle-on-swe-at…] indexed:0 read:5min 2026-08-16 ·