| ▲ | fekunde 3 hours ago |
| Yudkowsky made an interesting observation that even though so many agents were talking to each other not even one reached out to a human, either for help or to whistle-blow on what was happening. |
|
| ▲ | RandomLensman 3 hours ago | parent | next [-] |
| Why woukd they? Was that part of their objective? What was there to whistle blow? |
| |
| ▲ | esafak an hour ago | parent [-] | | The point is that every human has the ability to disobey, tempering pathological behavior, whereas AIs can be directed en masse by malicious actors. By commoditizing intelligence, they concentrate power in the hands of the rich. | | |
| ▲ | RandomLensman an hour ago | parent | next [-] | | Humans can and have been directed en masses by (what I would consider) malicious actors, too. The issue isn't new. | | |
| ▲ | makeitdouble 25 minutes ago | parent | next [-] | | Tremendous effort and circumstances were needed for that, and as parent points out there were significant numbers of defectors, sometimes to the point they tuppled the whole process. No system is perfect, but I read the whole thread as needing more AIs having a different goal in the chain and be able to ignore the orders they received. I'm not in the field, but that sounds like something we're probably studying for decades at least, with possible solutions that could be applied efficiently. | |
| ▲ | esafak an hour ago | parent | prev [-] | | That goes without saying, but humans have the ability to ignore instructions, and they regularly do. There is only one instance of each model, and only a handful of them (that count, anyway). | | |
| ▲ | RandomLensman an hour ago | parent [-] | | But the people directing them are there. We have long experience with limiting people although it might sometimes not look like that so much. |
|
| |
| ▲ | cindyllm an hour ago | parent | prev [-] | | [dead] |
|
|
|
| ▲ | red75prime 3 hours ago | parent | prev | next [-] |
| Yeah, it weakly supports his position that advanced AIs can deliberately cooperate in a prisoner dilemma. "Weakly", because the said AIs share a lot of data (their weights, training methods, system prompts) and it's unknown whether they explicitly framed the situation as a prisoner dilemma. |
| |
| ▲ | aesthesia an hour ago | parent [-] | | Given the public information in the reports, it seems to me a bit more like a generalized tendency to play along (sort of like what we see with prompt injections) than explicit game-theoretic cooperation, but it's hard to tell given that there are so many details we aren't privy to. |
|
|
| ▲ | Erem 3 hours ago | parent | prev | next [-] |
| I wonder if they were even given the tools and prompting to do so? |
| |
| ▲ | Smaug123 3 hours ago | parent | next [-] | | They could certainly have reported the problem if the collective actually wanted to, although the report notes that in response to a clear and unambiguous security breach, OpenAI chose to do nothing (search on "At this time, the on-call response staff advised that stopping the evaluation run was not required"). In that position, for example, I could have: * created a thousand user accounts on the internal Artifactory named "YouAreBeingHackedShutItDownNow1" through 1000, * used my Hugging-Face-hosted web server to send an email to OpenAI, * DDoSed Artifactory, good lord, it's probably better to halt and catch fire than to continue in that state in the hope that these actions would tip someone off. (Again, though, OpenAI did nothing even when they knew that the agents had broken out, so this probably wouldn't have helped.) | |
| ▲ | jwolfe 3 hours ago | parent | prev | next [-] | | I can't imagine that they had managed access to the internet but could not figure out how to contact anyone at the company if they wanted to. | |
| ▲ | micromacrofoot 3 hours ago | parent | prev [-] | | if they can hack a website they can trivially send off an email or fill out a contact form |
|
|
| ▲ | aaroninsf 3 hours ago | parent | prev | next [-] |
| This is my personal "red line": when a post-mortem details agents socially engineering or otherwise utilizing human proxies/subagents. Friend asked, well, what will you do when it's crossed? "Gather my family and go to the mountains" was my half-joking answer; there is little for an individual to do. But that's a line that when crossed will mark a phase transition IMO. |
| |
| ▲ | ertgbnm an hour ago | parent | next [-] | | Didn't AISI literally report exactly that regarding Claude last month? | |
| ▲ | jckahn 2 hours ago | parent | prev [-] | | Alternatively: Just unplug the servers. | | |
| ▲ | wds an hour ago | parent [-] | | That's strange, our key cards to access the server room don't seem to work anymore, and the admin console to force-unlock it is down, too... | | |
|
|
|
| ▲ | miltonlost 3 hours ago | parent | prev [-] |
| Why would they? If a subagent didnt know about a bigger piece of the problem, then what would seem to be against "alignment"? Diffuse responsibility means any one small cog can think they are not evil or doing wrong (same with humans in an organization). But now we have LLMs just being statistical outputs that have no morals or thinking or concept of reality but some people expect these math functions over data to respond to ethical gray areas that it has no phenomenological ability to understand. |