Hey HN!
We all know that most public evals are saturated and are hard to trust. What matters is whether a model reliably works in your codebase in your actual day-to-day-work.
To solve this, we built self-bench (https://github.com/mupt-ai/self-bench), an open-source repo that automatically turns your private, completed PRs into held-out coding-agent evals. It's fully open-source and you can run it on your own machine or on Modal sandboxes!
With one CLI command, we give you a validated Harbor dataset based on real changes from your repository. We've used self-bench to build evals for popular OSS projects, including Next.js (https://huggingface.co/datasets/dari-ai/nextjs-selfbench) and Vite (https://huggingface.co/datasets/dari-ai/vite-selfbench).
With self-bench, we want to make it easy to compare open-weight models, frontier models, and routers on performance, speed and cost.
Try out the repo or book a call if you want to chat through using it within your repos! [https://calendly.com/avyay-dari/30min](https://calendly.com/avyay-dari/30min)
Comments URL: [https://news.ycombinator.com/item?id=49300607](https://news.ycombinator.com/item?id=49300607)
Points: 2