Remix.run Logo
ethin a day ago

Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? This is what OAI should've done. If they had executed this training run in such a sandbox, the model wouldn't have been capable of escaping without social engineering, and if the models somehow managed to do that to it's evaluators then that is indeed a massive problem and OAI should disclose that.

pcthrowaway a day ago | parent | next [-]

It can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox

RandomLensman a day ago | parent | next [-]

A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think.

pcthrowaway a day ago | parent | next [-]

> A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out.

But this has actually happened... a lot. Search "social engineering prison breaks".

With AI it only needs to happen once.

I'm reminded of the scene in idiocracy where the protagonist, going through intake at the jail, tells the guard he's supposed to be getting out today, to which the guard says "you're in the wrong line dumbass" and waves him through.

To a true superhuman intelligence, we're the idiots who are theoretically easy to manipulate.

RandomLensman a day ago | parent [-]

I didn't say it doesn't happen, but that it is a low probability. And we have ways to reduce probabilities in critical areas.

There is no omnipotent AI currently (and there might never be) and I don't see why with current AI it only needs to happen once.

ben_w a day ago | parent [-]

They don't need to be omnipotent, and they're already human-or-superhuman at persuasion: https://arxiv.org/html/2411.06837v2

This may just be that humans find long arguments more persuasive than short ones, obviously LLMs can do that easily, but the outcome is I think more important than the mechanism.

RandomLensman a day ago | parent [-]

That is about persuasion with evidence on various topics, not about persuading people to abandon safty protocols and processes and highly policed settings.

Yes, many things could happen, but again, that failure is possible is not a reason to do implement processes etc. I don't see why hypotheticals should stop addressing actuals.

Smaug123 16 hours ago | parent [-]

I’m afraid human red-teamers against supposedly highly secure targets, with lots of protocols in highly policed settings, do frequently manage this kind of social engineering. There’s loads of stories of pentesting military establishments, for example.

RandomLensman 16 hours ago | parent [-]

Is there data on how frequently and what types of security levels? Military has varying levels of security and secrecy, for example.

Also, not a reason not to pursue processes etc., no? I doubt that things fail all the time, for example.

bottlepalm a day ago | parent | prev [-]

Prisoners don’t have much to offer if you help them escape. A malicious super AI on the other hand can probably find you millions of dollars worth of crypto in an afternoon.

RandomLensman a day ago | parent [-]

The current issues are not caused by some malicious god-like AI - maybe we need to focus on the issues at hand first rather than hypotheticals? (And we do have experience policing people around financial incentives, too. Nothing perfect, but also not nothing.)

ben_w a day ago | parent | next [-]

I suspect the current models probably can find literal millions lying around for the taking, given they could pull off the incident under discussion.

Tens of millions, even.

Getting them to run correctly is dangling in front of the researcher's noses a carrot labelled "tens of trillions", though I suspect this is an illusion in much the same way that Wikipedia is not valued at [peak cost of Encyclopaedia Britannica] * [global population with internet connection].

> And we do have experience policing people around financial incentives, too. Nothing perfect, but also not nothing.

Yes but be careful anthropomorphising the LLMs too much. They're only somewhat human-like in their behaviour, and to the extent that they're human-like they demonstrate a huge range of personality disorders: https://www.personalitybenchmark.ai

Though plus side, apparently not evil: https://arxiv.org/html/2406.14703v2

RandomLensman a day ago | parent [-]

I am not anthropomorphising the LLMs at all, I was talking about the obligations we put on humans using/making/etc. machines etc.

ben_w a day ago | parent [-]

Hmm. I think I misunderstood what you meant by "experience policing people around financial incentives" in that case.

RandomLensman a day ago | parent [-]

Simple examples here would be higher financial transparency obligations or more closely policing transactions.

bottlepalm 12 hours ago | parent | prev [-]

> we need to focus on the issues at hand first rather than hypotheticals

This attitude is what's got us here in the first place, and if we continue thinking like this when we're going to go right over the cliff. The hypothetical cliff that's coming, but we've never gone over a cliff before so we keep on driving.

Uhhrrr a day ago | parent | prev [-]

No. If OpenAI were being responsible and not criminally negligent, at the top of page 1 of the runbook would be "don't connect this to the actual Internet, even if the agent says Please."

bottlepalm a day ago | parent | prev | next [-]

Oh really? Please tell me how you intend to enforce AI is only run in the magic sandbox? Harsh HN comments?

ethin 18 hours ago | parent [-]

If I am evaluating an AI for safety, the last thing I would do is connect it to real-world peripherals or systems to allow it to reek havoc. That is criminal negligence at it's finest (especially if the AI is capable of committing crimes as happened here). I would place it on a system dedicated specifically for testing models, which had no NIC and no physical capability of accessing any outside system. If I wanted to know how the model might behave if given access to a certain system or set of systems, I would do it responsibly by writing simulation software which does it's best to simulate the real thing (and for networking this is already trivial to do). You could take this extremely far and simulate all kinds of things this way from basic networking to nuclear launch systems. And in the context of OpenAI, which is valued at over $1T, I have no qualms about stating that they (could) do this, because it is definitively something they could burn money on doing if they cared enough. They intentionally choose not to do so, and then have an amazed look on their faces when the model does something criminal like this.

bottlepalm 12 hours ago | parent [-]

You totally missed the point. When I asked:

> how you intend to enforce AI is only run in the magic sandbox

I didn't mean you, I meant everyone. How do you enforce everyone for example 'place [AI] on a system dedicated system' disconnected from the internet.

I don't think you can.

famouswaffles a day ago | parent | prev [-]

>Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability?

Nobody is building general intelligence and agents only to have it sit around doing nothing. It's going to have such capabilities.