请问HN:有没有针对大型语言模型(LLMs)的良好安全基准?
我自己也在寻找这个,但觉得进行一次实际的讨论是很有必要的。我对大型语言模型(LLMs)的基准测试方面还比较陌生。
例如,我看过 eyeballvull [1]。这个项目看起来很有前景,但我没有看到广泛的支持。我尊重作者仍然在坚持这个项目。我正在寻找的是一个代理能够全面扫描代码库的基准测试。
但我也在想:或许还有其他我尚未了解的项目。
冒着可能被淹没的风险,对于那些对安全意识的软件工程(无论是否使用代理)感兴趣的人,我们来聊聊吧!我的邮箱在我的个人资料中。
[1] https://arxiv.org/abs/2407.08708
查看原文
I'm looking for this myself but figured it's good to have an actual discussion about this. I'm pretty new to the benchmarking side of LLMs.<p>For example, I looked at eyeballvull [1]. It seems promising but I don't see wide support for example. I respect the author for still committing. A benchmark where an agent scans a repo in full is what I'm looking for.<p>But then I also wondered: maybe there are others out there that I haven't been aware of yet.<p>And, at the risk of potentially being flooded, for people that are curious about security aware software engineering with agents (or without them for that matter), let's have a chat! My email is in my profile.<p>[1] https://arxiv.org/abs/2407.08708