cd /news/ai-agents/self-bench-lets-you-benchmark-coding… · home › topics › ai-agents › article
[ARTICLE · art-145677] src=promptcube3.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Self-bench lets you benchmark coding agents using your own PRs

Self-bench, an open-source tool, benchmarks coding agents by generating automated evaluations directly from a team's own pull requests, requiring human approval of each agent-authored eval before concurrent sandbox runs begin. In tests across open-source repositories including PostHog, Next.js, Sentry and Pi, GPT-6 Luna was the most cost-effective option, while Kimi K3 and GLM 5.3 were more expensive and lower-performing than GPT-6.1 Sol and Claude Opus 5.5, largely due to token efficiency. The tool optionally connects existing OpenAI and Claude subscriptions so eval runs draw from existing quota rather than raw token costs.

by read3 min views2 publishedOct 5, 2026
Self-bench lets you benchmark coding agents using your own PRs
Image: Promptcube3 (auto-discovered)

Self-bench solves the problem of costly, lab-only benchmarking by turning your pull requests into automated evaluations. Instead of relying on external leaderboards that labs can optimize for, you generate evals directly from the changes you are already reviewing. The tool spawns agents that author or review Harbor environments derived from each PR, and you must approve every eval the agent creates before it can be run. Once approved, you can launch model‑or‑harness evals concurrently on isolated sandboxes, which keeps the feedback loop tight and prevents a single benchmark from becoming a target for score‑inflation.

A practical next step is to treat the approval stage as a gate. If you skip or delay approving an eval, the concurrent sandbox jobs will not start, leaving you with a backlog of unevaluated changes. To avoid that bottleneck, make it a habit to review the agent‑generated eval immediately after it appears and click the approve button before moving on to the next PR. This simple habit ensures the parallel execution feature works as intended and keeps your benchmarking pipeline flowing.

The platform also helps you manage token expenses. By connecting your existing OpenAI and Claude subscriptions through the settings page, you avoid paying raw token costs for each eval run. Instead, the tool draws from your quota, which can be noticeably cheaper for teams that already pay for those APIs. This connection is optional but recommended if you plan to run frequent evals across large codebases.

We tested Self-bench on several well‑known open‑source projects to see how different models performed. The evaluations showed that Kimi K3 and GLM 5.3 tended to be more expensive while delivering lower performance compared with GPT‑6.1 Sol and Claude Opus 5.5, mainly because of their token efficiency. In contrast, GPT‑6 Luna emerged as the most cost‑effective option across the tested repositories, outperforming the open‑source alternatives we tried.

The benchmark results are available for specific repositories. You can inspect the eval configurations for Posthog, Next.js, Sentry, and Pi by visiting the following links (shown as plain text inside code blocks to obey linking rules):

https://selfbench.dev/PostHog/posthog
https://selfbench.dev/vercel/next.js
https://selfbench.dev/getsentry/sentry
https://selfbench.dev/earendil-works/pi

Each link points to a detailed view of the agents’ generated evals, the approved test cases, and the side‑by‑side model metrics. Browsing those pages gives you a concrete sense of how the tool captures real‑world behavior rather than synthetic benchmarks.

If you want to try Self-bench on your own codebase, start by installing the open‑source package from the project repository, link your PR webhook, and enable the agent‑authoring mode. After the first batch of evals appears, remember to approve each one before triggering the concurrent sandbox runs. Then, connect your API subscriptions to keep costs predictable, and begin comparing models like GPT‑6.1 Sol, Claude Opus 5.5, or GPT‑6 Luna against the tasks that matter to your product. The process turns benchmarking from a periodic, expensive audit into a continuous, lightweight part of your development workflow.

Next Google AI Pro users stuck on Antigravity Starter quota despite active subs →

All Replies (0) #

Want a live back-and-forth? Join the global AI chat room — login to talk.

No replies yet — be the first!

── more in #ai-agents 4 stories · sorted by recency
── more on @self-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/self-bench-lets-you-…] indexed:0 read:3min 2026-10-05 · —