Benchmarking AI models on real hardware is way harder than it looks. Between AMD, NVIDIA, Apple Silicon, Intel, and Qualcomm, plus hundreds of open source models and quants, getting the test bench right is a challenge.
So we're making it dead simple. We're open sourcing our internal testing harness. Anyone can run open source models on their own hardware and submit results to the public leaderboard.
It's live now, with a few hundred submissions already.
If you want to know how a specific model performs on a given hardware, chances are the data is already there. This is the same tool we use internally to track model and chip performance. Try it out and tell us what to improve
Comments URL: [https://news.ycombinator.com/item?id=49737278](https://news.ycombinator.com/item?id=49737278)
Points: 2