问HN:有没有人对构建一个仅限于马具的基准测试感兴趣?
目前有很多大型语言模型(LLM)的基准测试,但很少有,甚至没有,专门针对“工具使用”的基准测试。我认为这将是一个非常好的社区项目,可以建立这样一个基准测试。
最终目标:建立一个关于“工具使用”性能的排行榜(多维度),涵盖一系列多样化的现实世界任务[1],并按基础模型和推理努力进行分组。任何人都可以贡献结果。
任务标准、测量方法、基础框架等可以由一个团队而不是单个人决定。
如果有足够的兴趣,我将创建一个Discord频道。
披露:我是一款名为Dirac的编码代理的维护者(https://github.com/dirac-run/dirac),因此我不会影响最终基准测试的样子,以避免任何利益冲突。我只是想让这个项目实现。
[1] 多样化的现实世界任务是指那些复杂程度足够高、贡献者曾经遇到过的任务,最好来自开源代码库。
查看原文
There are a lot of LLM benchmarks but few, if any, harness benchmarks. I am thinking this would be a really good community project to build one.<p>End goal: a leaderboard of harness performance (multiple axis) on a set of diverse real world tasks[1], grouped by underlying models and reasoning efforts. Anyone can contribute results.<p>The task criteria, measurements, underlying framework et al can be decided by a group rather than a single person.<p>If there is sufficient interest, I will create a discord.<p>Disclosure: I am the maintainer of a coding agent called Dirac (https://github.com/dirac-run/dirac) so I will not influence what the final benchmark should look like to avoid any conflict of interest. I just want to make this happen.<p>[1] Diverse real world tasks meaning sufficiently complex tasks that the contributors have encountered, preferably from an opensource repo.