Put together an extensive open source test suite for voice assistants. You can view it including current leaderboard at: https://git.cicero.sh/aquila/ha-voice-test-suite/
Tests are reproduceable, with clear instructions on how to run them on your machine there.
I only have a GPU with 4GB vRAM, hence only capable of testing the small LLMs like Qwen3 4B Instruct. Have Gemma 4 running right now, but it's insanely slow and probably another 24 - 48 hours before it finishes.
Tried cloud models like Claude Sonnet 5 and Grok 4.5, but was rate limited each time, and got fed up so dropped it. Maybe someone out there uses AI more than me and has proper rate limits and wouldn't mind running the test suite on some frontier models as I'm assuming folks would appreciate seeing the results. I'm already confident they'll get low 90s with 60 - 90 minute duration, so not overly worried about it.
If anyone has larger GPU and wouldn't mind giving it a spin on larger 20B+ local LLMs that would be awesome. Running the tests against any model is quite straight forward, instructions in the readme. Enjoy!
Comments URL: [https://news.ycombinator.com/item?id=49150636](https://news.ycombinator.com/item?id=49150636)
Points: 1