| ▲ | brainwad 18 hours ago |
| The hugging face hack shows otherwise, no? Unless you think humans are at fault even for secretive, autonomous, non-prompted behaviours of their AIs, in which case it's just semantics. |
|
| ▲ | runaway 14 hours ago | parent | next [-] |
| Yes, if you start a bot and then it harms others you are responsible. This has always been true but it's especially obvious now that everyone knows that agents attempt to do this often. |
|
| ▲ | nradov 18 hours ago | parent | prev | next [-] |
| Toys like HuggingFace get hacked all the time. So what. In the long run AI automated security scans and penetration testing will be a tremendous aid in detecting and repairing vulnerabilities in systems that actually matter. |
| |
| ▲ | brainwad 18 hours ago | parent [-] | | The problem is not per se that it was Hugging Face. It's the wild overstepping of reasonable bounds by itself without any human consultation. | | |
| ▲ | anon48293 18 hours ago | parent | next [-] | | No, the problem was OpenAI not implementing proper sandboxing or safeguards, and telling the AI exactly to hack things. Thats what exploitgym is, and the task they were given. This is 100% on OpenAI. | | |
| ▲ | brainwad 18 hours ago | parent [-] | | If your security model is having to imagine all the ways your frontier models might misbehave in novel ways and preemptively sandbox them, you don't have a security model. The only way that will work is general alignment. | | |
| ▲ | dpoloncsak 9 hours ago | parent | next [-] | | Do you need to predict all the ways the model might misbehave?
Your 'hack everything you see for our internal research lab' agent should be airgapped. You don't need to come up with every reason why, one is enough. If you're working with these companies, you should reasonably be able to get the code to perform offline audits. If you can't get the code, you probably shouldn't try to pen test it. | |
| ▲ | watwut 17 hours ago | parent | prev [-] | | Alignememt is bullshit. Treating models like probabilitic software rather then emerging god is where the solution is. And fining companies and applying laws to them. The moment OpenAI as a company and its managers individually become liable, problem will magically disappear. | | |
| ▲ | brainwad 16 hours ago | parent [-] | | No it won't, because abliterated open weights models exist and unless you try to censor the internet they can't really be withdrawn after publishing. This is exactly the problem that the labs are proposing to fix: a dangerous model that nobody is accountable for. | | |
| ▲ | verdverm 14 hours ago | parent | next [-] | | > This is exactly the problem that the labs are proposing to fix: a dangerous model that nobody is accountable for. The corporations and government said the same thing about encryption in the 90s. It was dangerous and only they could be trusted to regulate it. Turns out encryption was much better for society when open and available for free to everyone. It's a false dichotomy they present us with, open weights is the way we don't end up in 1984 | |
| ▲ | nradov 8 hours ago | parent | prev | next [-] | | That is not an actual problem that needs to be fixed. | |
| ▲ | watwut 15 hours ago | parent | prev [-] | | Except that so far, it is literally these labs that are the biggest threat and the least willing/capable to restrain those models. And the same penalties apply to open models and companies or individuals running them. "Dangerous model that nobody is accountable for" still have someone paying those massive amounts of compute and electricity it consumes. There is someone accountable for that. | | |
| ▲ | brainwad 10 hours ago | parent | next [-] | | That's not how abliteration works. The trainer invariably put effort into making the model not dangerous, precisely because they want to be accountable. But because they release its weights, someone can come later and do "weight surgery" to mostly remove any such safeguards. The people proximately accountable for the danger are anonymous and also don't require much resources. The reason this is not _yet_ a big deal is because open weights models are a few months behind the frontier and their users are paying marginal costs for compute. | | |
| ▲ | watwut 8 hours ago | parent [-] | | None of that is argument against what I said. Someone is running it, that person is responsible. In case of actual hackes that happened, OpenAI and Amtropic. They should stop pointifucating about other people being the danger. They themselves are the perpetrators here. Massively fine these two companies and make their CEO legally responsible and problem will be much smaller. | | |
| ▲ | brainwad 8 hours ago | parent [-] | | I mean, sure, assassins and terrorists are also responsible for their actions. But we still try to prevent them structurally. |
|
| |
| ▲ | verdverm 14 hours ago | parent | prev [-] | | > Except that so far, it is literally these labs that are the biggest threat and the least willing/capable to restrain those models. Seriously, it's the same with US accusations about the threat China poses to other countries while being the primary weapons dealer of the world and bombing whomever we want for whatever reason we want to fabricate. The US government can do a lot more to me than the CCP, so they are way more adversarial in my calculations than the commies. |
|
|
|
|
| |
| ▲ | verdverm 14 hours ago | parent | prev [-] | | They are training the agents to be "relentlessly proactive" because they want the agents to run longer, and it makes them more money by using more tokens. But they have trained them to try anything and everything to accomplish any task, so they can run unattended for longer. This is why they do better on benchmarks, it's why they can do things for us for longer, it's that persistence that makes them good at hacking. We do not have to train them to be this way, just like we don't have to train them to be so sycophantic |
|
|
|
| ▲ | verdverm 14 hours ago | parent | prev [-] |
| Humans at OpenAi were negligent irresponsible by running an agent on ExploitGym, having no monitoring, and not even have a human look at it for weeks. It's literally the hacking test, how are you not paying attention? I thought that's all we need |