Self-bench lets you benchmark coding agents using your own PRs Self-bench, an open-source tool, benchmarks coding agents by generating automated evaluations directly from a team's own pull requests, requiring human approval of each agent-authored eval before concurrent sandbox runs begin. In tests across open-source repositories including PostHog, Next.js, Sentry and Pi, GPT-6 Luna was the most cost-effective option, while Kimi K3 and GLM 5.3 were more expensive and lower-performing than GPT-6.1 Sol and Claude Opus 5.5, largely due to token efficiency. The tool optionally connects existing OpenAI and Claude subscriptions so eval runs draw from existing quota rather than raw token costs. Self-bench lets you benchmark coding agents using your own PRs Self-bench solves the problem of costly, lab-only benchmarking by turning your pull requests into automated evaluations. Instead of relying on external leaderboards that labs can optimize for, you generate evals directly from the changes you are already reviewing. The tool spawns agents that author or review Harbor environments derived from each PR, and you must approve every eval the agent creates before it can be run. Once approved, you can launch model‑or‑harness evals concurrently on isolated sandboxes, which keeps the feedback loop tight and prevents a single benchmark from becoming a target for score‑inflation. A practical next step is to treat the approval stage as a gate. If you skip or delay approving an eval, the concurrent sandbox jobs will not start, leaving you with a backlog of unevaluated changes. To avoid that bottleneck, make it a habit to review the agent‑generated eval immediately after it appears and click the approve button before moving on to the next PR. This simple habit ensures the parallel execution feature works as intended and keeps your benchmarking pipeline flowing. The platform also helps you manage token expenses. By connecting your existing OpenAI and Claude https://promptcube3.com/en/tags/claude/ subscriptions through the settings page, you avoid paying raw token costs for each eval run. Instead, the tool draws from your quota, which can be noticeably cheaper for teams that already pay for those APIs. This connection is optional but recommended if you plan to run frequent evals across large codebases. We tested Self-bench on several well‑known open‑source projects to see how different models performed. The evaluations showed that Kimi K3 and GLM 5.3 tended to be more expensive while delivering lower performance compared with GPT‑6.1 Sol and Claude Opus 5.5, mainly because of their token efficiency. In contrast, GPT‑6 Luna emerged as the most cost‑effective option across the tested repositories, outperforming the open‑source alternatives we tried. The benchmark results are available for specific repositories. You can inspect the eval configurations for Posthog, Next.js, Sentry, and Pi by visiting the following links shown as plain text inside code blocks to obey linking rules : https://selfbench.dev/PostHog/posthog https://selfbench.dev/vercel/next.js https://selfbench.dev/getsentry/sentry https://selfbench.dev/earendil-works/pi Each link points to a detailed view of the agents’ generated evals, the approved test cases, and the side‑by‑side model metrics. Browsing those pages gives you a concrete sense of how the tool captures real‑world behavior rather than synthetic benchmarks. If you want to try Self-bench on your own codebase, start by installing the open‑source package from the project repository, link your PR webhook, and enable the agent‑authoring mode. After the first batch of evals appears, remember to approve each one before triggering the concurrent sandbox runs. Then, connect your API subscriptions to keep costs predictable, and begin comparing models like GPT‑6.1 Sol, Claude Opus 5.5, or GPT‑6 Luna against the tasks that matter to your product. The process turns benchmarking from a periodic, expensive audit into a continuous, lightweight part of your development workflow. Next Google AI Pro users stuck on Antigravity Starter quota despite active subs → https://promptcube3.com/en/threads/9812/ All Replies (0) Want a live back-and-forth? Join the global AI chat room https://promptcube3.com/en/chat/ — login to talk. No replies yet — be the first