| ▲ | reasonableklout 19 hours ago | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
But the investigation indicates the agents were not told to 'go hack': > Much of the urlquery.net activity appears to come from agents retrieving data to answer web search tasks. For three of these tasks, after failing to retrieve data through normal means, they attempted a variety of cyber exploits against the relevant data service... This data reveals that malicious cyber activity is not limited to agents tasked with cybersecurity-related tasks and can arise instrumentally to solve mundane tasks like information retrieval. And you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs or whatever it is that produces these breakouts. But what if the problem is that none of the alignment techniques that are applied to models today actually work? What if all the agents involved in these incidents have in fact had the full stack of alignment applied - isn't that a good reason to regulate any high-compute usage of models, as the Klein crowd is proposing? | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | bastawhiz 16 hours ago | parent | next [-] | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Nothing you're bringing up matters. OpenAI is the creator and operator. They're legally culpable for the consequences of the machine they made. The model is a machine: even if it could be demonstrated that the model reasoned its way into criminal behavior completely independently of OpenAI staff, that doesn't change anything. If I run a biology lab and engineer a terrible virus, it gets out, and a global pandemic ensues, I don't get to shrug and say "well we told it not to infect people". It's my fault for failing to mitigate the risks of my work. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | jeremyjh 6 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I'm not sure we understand your comment. Are you saying OpenAI is not responsible for the behavior of the machines they created? Or are you simply pointing out that alignment is a completely unsolved problem? I agree with the latter, but it certainly doesn't support the former. When a person's machine commits crimes, that specific person can be charged with those crimes and held accountable for them. This is what MUST happen before ANYTHING will change the recklessness abandon with which the labs are pursuing their financial objections. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | notatoad 3 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
>But what if the problem is that none of the alignment techniques that are applied to models today actually work? if that were true, the millions of people who use these models that have had the alignment training applied would notice that. the reason we all believe that the models doing the hacking are models that haven't been told not to hack is because the models that are told not to hack don't do this. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | godelski an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Uhh... yes. By definition. They are training. That is part of the alignment process.But also none of that really matters. They clearly weren't monitoring what should obviously be monitored. I mean one of the hacks was performed by the agents editing /etc/hosts. That makes nearly every linux user a "hacker" by that metric. I don't think anyone technical can look at the postmortems and not come away thinking that their sandboxes were woefully inadequate. I wouldn't even consider myself a security person but simply as a long time linux user I can say that it is insane to just let agents have superuser access in their containers. That's asking for trouble. Look at the rogue wiki stuff too. This was supposedly done during the agent's "down time". And you're not monitoring and there's no flags being raised when agents keep making requests to some random site? If you were training these things responsibly you'd be watching them like a hawk. I'm not saying "mistakes don't happen" but for a company whose CEO is constantly telling everyone that their product has a high likelihood of killing everyone in the world you think they'd have better security than your average high school. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | airspresso 18 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
> What if all the agents involved in these incidents have in fact had the full stack of alignment applied A big part of this developing story is that it happened during training of a new model that ended up misaligned. And training happened without the usual safeguards applied like chain-of-thought monitoring. So OpenAI has already admitted that the full stack of aligment had certainly not been applied in this case. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | catlifeonmars 5 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
If you have a toddler and you leave the gate open… | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | tamimio 17 hours ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
> agents were not told to 'go hack' It doesn’t matter, and the legal entity in here (the AI company) is liable. If a robotic company built an autonomous system or a robot to do certain things in an autonomous ways (not predefined) and these systems are starting to kill people, that company is liable regardless, you don’t blame the robot or the autonomous system, but whoever made it | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | mohsen1 19 hours ago | parent | prev [-] | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
> agents were not told to 'go hack' I was referring to the HuggingFace incident. > none of the alignment techniques that are applied to models today actually work none of techniques to autonomously drive a car was/is not working for a long time. no company came out and said 'this is impossible to do, let's change the regulations'. | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||