Aquila 语音助手测试套件用于家庭助理

1作者: aquila4165 天前原帖
为语音助手整理了一个全面的开源测试套件。您可以在以下链接查看,包括当前的排行榜:https://git.cicero.sh/aquila/ha-voice-test-suite 测试是可重复的,并提供了清晰的说明,指导您如何在自己的机器上运行它们。 我只有一块4GB显存的GPU,因此只能测试像Qwen3 4B Instruct这样的小型大语言模型。目前正在运行Gemma 4,但速度非常慢,可能还需要24到48小时才能完成。 我尝试过像Claude Sonnet 5和Grok 4.5这样的云模型,但每次都受到速率限制,感到厌烦所以放弃了。也许有其他人比我更频繁地使用AI,并且有合适的速率限制,愿意在一些前沿模型上运行测试套件,我相信大家会很欣赏看到这些结果。我已经很有信心它们会在60到90分钟的测试中获得低90分的成绩,所以对此并不太担心。 如果有人有更大的GPU,并愿意在更大的20B+本地大语言模型上试试,那就太好了。对任何模型运行测试都相当简单,说明在自述文件中。 祝您愉快!
查看原文
Put together an extensive open source test suite for voice assistants. You can view it including current leaderboard at: https:&#x2F;&#x2F;git.cicero.sh&#x2F;aquila&#x2F;ha-voice-test-suite&#x2F;<p>Tests are reproduceable, with clear instructions on how to run them on your machine there.<p>I only have a GPU with 4GB vRAM, hence only capable of testing the small LLMs like Qwen3 4B Instruct. Have Gemma 4 running right now, but it&#x27;s insanely slow and probably another 24 - 48 hours before it finishes.<p>Tried cloud models like Claude Sonnet 5 and Grok 4.5, but was rate limited each time, and got fed up so dropped it. Maybe someone out there uses AI more than me and has proper rate limits and wouldn&#x27;t mind running the test suite on some frontier models as I&#x27;m assuming folks would appreciate seeing the results. I&#x27;m already confident they&#x27;ll get low 90s with 60 - 90 minute duration, so not overly worried about it.<p>If anyone has larger GPU and wouldn&#x27;t mind giving it a spin on larger 20B+ local LLMs that would be awesome. Running the tests against any model is quite straight forward, instructions in the readme.<p>Enjoy!