Show HN: Self-bench – build SWE-bench style evals from private repos Mupt AI released self-bench, an open-source tool that automatically converts private, completed pull requests into held-out coding-agent evals, generating validated Harbor datasets with a single CLI command. The tool has been used to build evals for Next.js and Vite, and aims to help developers compare open-weight models, frontier models, and routers on performance, speed, and cost. Hey HN We all know that most public evals are saturated and are hard to trust. What matters is whether a model reliably works in your codebase in your actual day-to-day-work. To solve this, we built self-bench https://github.com/mupt-ai/self-bench https://github.com/mupt-ai/self-bench , an open-source repo that automatically turns your private, completed PRs into held-out coding-agent evals. It's fully open-source and you can run it on your own machine or on Modal sandboxes With one CLI command, we give you a validated Harbor dataset based on real changes from your repository. We've used self-bench to build evals for popular OSS projects, including Next.js https://huggingface.co/datasets/dari-ai/nextjs-selfbench https://huggingface.co/datasets/dari-ai/nextjs-selfbench and Vite https://huggingface.co/datasets/dari-ai/vite-selfbench https://huggingface.co/datasets/dari-ai/vite-selfbench . With self-bench, we want to make it easy to compare open-weight models, frontier models, and routers on performance, speed and cost. Try out the repo or book a call if you want to chat through using it within your repos https://calendly.com/avyay-dari/30min https://calendly.com/avyay-dari/30min Comments URL: https://news.ycombinator.com/item?id=49300607 https://news.ycombinator.com/item?id=49300607 Points: 2 Comments: 0