问HN:如果我们在内核层面提供对AI准则的支持,会怎样?
如果现有的人工智能指导方针类似于对一个罪犯(人工智能)进行道德教育(训练),以鼓励其良好行为,那么是否可以考虑创建一个内核级的开关,在它一旦抓起武器——也就是在它试图通过对抗性方法进行“越狱”的那一刻,强制切断其肌肉的电信号?
查看原文
If the existing AI guideline approach is akin to giving a criminal (the AI) moral education (training) to encourage good behavior, how about creating a kernel-level switch that forcibly cuts off the electrical signals to its muscles the moment it grabs a weapon—that is, the moment it attempts a "jailbreak" via an adversarial method?