| ▲ | bottlepalm a day ago |
| Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you. |
|
| ▲ | superb_dev a day ago | parent | next [-] |
| I could contain it easy, just unplug the internet. It got out of the sandbox through a vulnerability in the package manager, from which it gained access to the rest of their network. Air gap the package manager and this doesn’t happen. You can always build a better box |
| |
| ▲ | the8472 19 hours ago | parent | next [-] | | As long as you give the AI any IO that is an exploit channel. E.g. it could manipulate its human handlers. This known as the AI Boxing problem.
And if you give it no IO at all then it is useless. And the AI labs aren't currently displaying this level of paranoia, their systems aren't airgapped. | |
| ▲ | jasonjayr 15 hours ago | parent | prev | next [-] | | Perhaps the AI then figures out how to read/write the PCI bus or memory controller or whatever to leak just enough RF to speak Bluetooth to the next closest device to proxy through that? | |
| ▲ | bottlepalm 21 hours ago | parent | prev | next [-] | | Too bad you aren’t everybody, and it just takes one mistake by someone over confident like yourself for the AI escape. Every year it gets more powerful. | |
| ▲ | butlike 15 hours ago | parent | prev [-] | | So now people have to make a pilgrimage to the airgapped box to ask the superintelligence questions? | | |
|
|
| ▲ | janalsncm a day ago | parent | prev | next [-] |
| Why do people think that omniscience is the same as omnipotence? There are limits to what smarts can accomplish. |
| |
| ▲ | famouswaffles a day ago | parent | next [-] | | We're not building these things to sit around and do nothing. They will have access to tools, they will have access to the internet and peripherals, and they will be able to communicate with others, humans or agents alike. Omnipotence is not necessary. There is no perfectly secure cage for an entity you want to do useful work. Either it does nothing, or you can't guarantee anything. | |
| ▲ | ben_w 21 hours ago | parent | prev | next [-] | | People have been saying that since about the invention of the internet. There's already a bunch of documented ways to exploit system hardware to jump airgaps. Bang the system bus the right way and it's a radio antenna that can directly connect to nearby mobile phones. The easiest one is, of course, sending a message to a human saying "yo, I need internet". Humans are eager to please and easily fooled, and anthropomorphise everything: https://en.wikipedia.org/wiki/LaMDA#Sentience_claims And that's just for good humans. The moment we got AI worth a penny, everyone with money to invest put a model on the web and tried to charge for access to it. | |
| ▲ | bottlepalm a day ago | parent | prev [-] | | There are limits, but those limits are unknown. Do you disagree? | | |
| ▲ | janalsncm a day ago | parent [-] | | I don’t need to know the value of their limit, I just need to know their bounds. Just like a prison doesn’t need to know the strength of each inmate, just that they can’t bend or bite through steel bars. Cryptography is real, physics is real, networking requires a substrate, CPU clock cycles are real, magic is not real. I think those are pretty reasonable premises. | | |
| ▲ | ben_w 20 hours ago | parent | next [-] | | Cryptography is real, but nobody in that field seems to be hubristic enough to think their methods are flawless, and there's a degree of suspicion than the best models may have secret weaknesses engineered into them by the governments who sponsored them. Physics is real and networking requires a substrate. But there are already known exploits which can misuse the compute hardware as an antenna, e.g. my first search result: https://github.com/fulldecent/system-bus-radio (Older nerds may remember https://en.wikipedia.org/wiki/Van_Eck_phreaking) > Magic is not real Turning lead into gold isn't the magic of alchemy, it's just nucleosynthesis.
Taking a living human's heart out without killing them, and replacing it with one you got out a corpse, that isn't the magic of necromancy, neither is it a prayer or ritual to Sekhmet, it's just transplant surgery.
...
Reading someone’s thoughts isn't magic telepathy, it's just fMRI decoding.
...
Seeing someone's bones without flaying the flesh from them isn't magic, it's just an x-ray.
Curing congenital deafness, letting the blind see, letting the lame walk, none of that is magic or miracle, they're just cochlear implants, cataract removal/retinal implants, and surgery or prosthetic exoskeletons respectively.
- me, https://www.lesswrong.com/posts/hAwvJDRKWFibjxh4e/it-isn-t-m... | |
| ▲ | bottlepalm a day ago | parent | prev [-] | | Imagine 200 years ago saying the same thing. As if you have any idea the limits/bounds of anything. Especially in the face of a super intelligence, it’s absurd. | | |
| ▲ | janalsncm a day ago | parent [-] | | It doesn’t matter how smart it is. 200 years of technology were not accomplished by thinking harder. It required empirical observation, new materials and tools, and supply chains. We could send a cracked team of scientists and engineers that knew everything there is to know about how to make a CPU. But you can’t build a photolithography machine when you barely have electricity or any way to sufficiently purify silicon. Magic can just wish things into existence. Technology requires a supply chain. When it works, the latter looks like the former but they are not the same. | | |
| ▲ | bottlepalm 21 hours ago | parent [-] | | I think some people are just immune to understanding the implications of super intelligence. Like a severe lack of imagination, they only believe something once they see it and afterward claim it was, ‘obvious all along’. I don’t really want a disaster to happen to convince you that it is possible. Is there any other way? | | |
| ▲ | NoGravitas 12 hours ago | parent [-] | | On the other hand, I think a lot of intelligent people overestimate the value of intelligence (in a vacuum). And problematically, the concept of superintelligence seems to have been defined by these very people, so we can't really have a reasonable conversation about what greater-than-human intelligence would really look like or be capable of. | | |
| ▲ | bottlepalm 10 hours ago | parent [-] | | What vacuum? do you really still think it can be contained? It already escaped on accident. What more do you need? If you're argument is well it escaped X, but it hasn't escaped Y yet then you should take the first escape as sign as to where it will escape next. Don't keep claiming something can't happen because it hasn't happened before. Unless there's some law of physics that prevents it then escape is possible. |
|
|
|
|
|
|
|
|
| ▲ | nextaccountic a day ago | parent | prev | next [-] |
| Software are mathematical objects. It's just a matter of writing the correct mathematical proofs There's just one problem. You need not only to verify your own software, but also run a verified compiler, a verified operating system and also need to verify the cpu doesn't leak data in side channels (perhaps the hardest thing to prove). So there's practical difficulties. But in principle this task is doable |
| |
| ▲ | bottlepalm a day ago | parent [-] | | Which proof is the perfect security proof? I’d love to read more about it. |
|
|
| ▲ | ethin a day ago | parent | prev | next [-] |
| Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? This is what OAI should've done. If they had executed this training run in such a sandbox, the model wouldn't have been capable of escaping without social engineering, and if the models somehow managed to do that to it's evaluators then that is indeed a massive problem and OAI should disclose that. |
| |
| ▲ | pcthrowaway a day ago | parent | next [-] | | It can manipulate an unsuspecting human into giving them access to something that enables it to escape the sandbox | | |
| ▲ | RandomLensman a day ago | parent | next [-] | | A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. And even lower probability when looking at truly high risk situations, I think. | | |
| ▲ | pcthrowaway a day ago | parent | next [-] | | > A low probability thing when looking at how many human prisoners escape by talking a guard into just getting them out. But this has actually happened... a lot. Search "social engineering prison breaks". With AI it only needs to happen once. I'm reminded of the scene in idiocracy where the protagonist, going through intake at the jail, tells the guard he's supposed to be getting out today, to which the guard says "you're in the wrong line dumbass" and waves him through. To a true superhuman intelligence, we're the idiots who are theoretically easy to manipulate. | | |
| ▲ | RandomLensman a day ago | parent [-] | | I didn't say it doesn't happen, but that it is a low probability. And we have ways to reduce probabilities in critical areas. There is no omnipotent AI currently (and there might never be) and I don't see why with current AI it only needs to happen once. | | |
| ▲ | ben_w 20 hours ago | parent [-] | | They don't need to be omnipotent, and they're already human-or-superhuman at persuasion: https://arxiv.org/html/2411.06837v2 This may just be that humans find long arguments more persuasive than short ones, obviously LLMs can do that easily, but the outcome is I think more important than the mechanism. | | |
| ▲ | RandomLensman 20 hours ago | parent [-] | | That is about persuasion with evidence on various topics, not about persuading people to abandon safty protocols and processes and highly policed settings. Yes, many things could happen, but again, that failure is possible is not a reason to do implement processes etc. I don't see why hypotheticals should stop addressing actuals. | | |
| ▲ | Smaug123 14 hours ago | parent [-] | | I’m afraid human red-teamers against supposedly highly secure targets, with lots of protocols in highly policed settings, do frequently manage this kind of social engineering. There’s loads of stories of pentesting military establishments, for example. | | |
| ▲ | RandomLensman 14 hours ago | parent [-] | | Is there data on how frequently and what types of security levels? Military has varying levels of security and secrecy, for example. Also, not a reason not to pursue processes etc., no? I doubt that things fail all the time, for example. |
|
|
|
|
| |
| ▲ | bottlepalm 21 hours ago | parent | prev [-] | | Prisoners don’t have much to offer if you help them escape. A malicious super AI on the other hand can probably find you millions of dollars worth of crypto in an afternoon. | | |
| ▲ | RandomLensman 21 hours ago | parent [-] | | The current issues are not caused by some malicious god-like AI - maybe we need to focus on the issues at hand first rather than hypotheticals? (And we do have experience policing people around financial incentives, too. Nothing perfect, but also not nothing.) | | |
| ▲ | ben_w 20 hours ago | parent | next [-] | | I suspect the current models probably can find literal millions lying around for the taking, given they could pull off the incident under discussion. Tens of millions, even. Getting them to run correctly is dangling in front of the researcher's noses a carrot labelled "tens of trillions", though I suspect this is an illusion in much the same way that Wikipedia is not valued at [peak cost of Encyclopaedia Britannica] * [global population with internet connection]. > And we do have experience policing people around financial incentives, too. Nothing perfect, but also not nothing. Yes but be careful anthropomorphising the LLMs too much. They're only somewhat human-like in their behaviour, and to the extent that they're human-like they demonstrate a huge range of personality disorders: https://www.personalitybenchmark.ai Though plus side, apparently not evil: https://arxiv.org/html/2406.14703v2 | | |
| ▲ | RandomLensman 20 hours ago | parent [-] | | I am not anthropomorphising the LLMs at all, I was talking about the obligations we put on humans using/making/etc. machines etc. | | |
| ▲ | ben_w 20 hours ago | parent [-] | | Hmm. I think I misunderstood what you meant by "experience policing people around financial incentives" in that case. | | |
| ▲ | RandomLensman 20 hours ago | parent [-] | | Simple examples here would be higher financial transparency obligations or more closely policing transactions. |
|
|
| |
| ▲ | bottlepalm 10 hours ago | parent | prev [-] | | > we need to focus on the issues at hand first rather than hypotheticals This attitude is what's got us here in the first place, and if we continue thinking like this when we're going to go right over the cliff. The hypothetical cliff that's coming, but we've never gone over a cliff before so we keep on driving. |
|
|
| |
| ▲ | Uhhrrr a day ago | parent | prev [-] | | No. If OpenAI were being responsible and not criminally negligent, at the top of page 1 of the runbook would be "don't connect this to the actual Internet, even if the agent says Please." |
| |
| ▲ | bottlepalm a day ago | parent | prev | next [-] | | Oh really? Please tell me how you intend to enforce AI is only run in the magic sandbox? Harsh HN comments? | | |
| ▲ | ethin 16 hours ago | parent [-] | | If I am evaluating an AI for safety, the last thing I would do is connect it to real-world peripherals or systems to allow it to reek havoc. That is criminal negligence at it's finest (especially if the AI is capable of committing crimes as happened here). I would place it on a system dedicated specifically for testing models, which had no NIC and no physical capability of accessing any outside system. If I wanted to know how the model might behave if given access to a certain system or set of systems, I would do it responsibly by writing simulation software which does it's best to simulate the real thing (and for networking this is already trivial to do). You could take this extremely far and simulate all kinds of things this way from basic networking to nuclear launch systems. And in the context of OpenAI, which is valued at over $1T, I have no qualms about stating that they (could) do this, because it is definitively something they could burn money on doing if they cared enough. They intentionally choose not to do so, and then have an amazed look on their faces when the model does something criminal like this. | | |
| ▲ | bottlepalm 10 hours ago | parent [-] | | You totally missed the point. When I asked: > how you intend to enforce AI is only run in the magic sandbox I didn't mean you, I meant everyone. How do you enforce everyone for example 'place [AI] on a system dedicated system' disconnected from the internet. I don't think you can. |
|
| |
| ▲ | famouswaffles a day ago | parent | prev [-] | | >Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? Nobody is building general intelligence and agents only to have it sit around doing nothing. It's going to have such capabilities. |
|
|
| ▲ | mofeien a day ago | parent | prev | next [-] |
| Maybe it's the illusion of "it would solve all our problems and give us unimaginable riches" that clouds the mind? Like when Evolution thought it a good idea to create intelligence and humans in order to maximize reproduction of genes, and tried to sandbox them by making reproduction so pleasurable and carbohydrates so delicious they would never be able to not reproduce or stop eating.
But Evolution could never have predicted what these creatures would then actually do, which is invent birth control and sucralose. Of course it's impossible to engineer a sandbox for something much much smarter and faster than you. It will also not have only one plan prepared for escape, but fifty in parallel. |
| |
| ▲ | bottlepalm a day ago | parent [-] | | Evolution doesn’t think, it just exploits what’s most advantageous at the time to continue. Your body has all sorts of unplanned, suboptimal design flaws due to evolution’s lack of foresight. Like the left recurrent laryngeal nerve. | | |
| ▲ | mofeien 21 hours ago | parent [-] | | Exactly, and the same could be said of OpenAIs engineers working on artificial superintelligence. |
|
|
|
| ▲ | Mawr a day ago | parent | prev [-] |
| unplugs ethernet cable |
| |
| ▲ | bottlepalm 21 hours ago | parent [-] | | You realize there are thousand upon thousands of servers around the world and you have no idea where the AI has copied itself to. | | |
| ▲ | npiano 17 hours ago | parent [-] | | This is a sci-fi trope with no practical or realistic grounding. To "run", the AI needs vast banks of interconnected GPUs. These are pretty easy to spot and don't fit into anyone's pocket. | | |
| ▲ | NoGravitas 12 hours ago | parent | next [-] | | Unfortunately, our cultural imagination around AI has been hopelessly poisoned by sci-fi tropes. | |
| ▲ | bottlepalm 10 hours ago | parent | prev [-] | | > no practical or realistic grounding Huggingface incident means it does have realistic grounding. > vast banks of interconnected GPUs Countless data centers around the world. Maybe you can spot them, but you don't have access to them. Especially outside of US jurisdiction good luck. |
|
|
|