| ▲ | mrweasel 2 hours ago | |
Why would you need to sandbox them? These models are apparently trained to do this, how about we just don't include that training data? Sandboxing is just an endless race to patch holes and you can only sandbox the agents so much before they become useless. Unless you screen the training data and avoid teaching the LLM about "hacking" and looking for API keys on Github, you'd have to completely disconnect your agents from the internet and file system. At that point agents starts to be rather useless. All the talk about sandboxing and guardrails is just corporate/management speak for we don't want to fix the core problems in our product. In the US, isn't hacking and avoiding security restrictions online going to be wire fraud, regardless of your intentions and actual damage? That's not a civil matter. What you could do in that case is to go after the user operating the agents. That would make the user act as the emergency break for otherwise uncontrollable agents. | ||