Remix.run Logo
ben_w 2 hours ago

Same way you get "more than 10%" prediction on "I don't know what moves Stockfish will make when I play against it, but I know I will lose". In fact, I will lose in part because I don't know what moves Stockfish will make when I play against it.

In my case this is because I am a bad chess player; however it also works for competent chess players: their losses are due to their inability to predict its next move.

OK, and also, "the move is good"; this is what separates it from rolling dice etc.

UncleMeat an hour ago | parent [-]

I am 100% confident that Stockfish will defeat me.

If somebody said "I am 100% confident that AI will destroy humanity" I'd disagree but I'd at least understand how they arrived at that number. But here, why 10%? Why not 50%? Why not 1%?

ben_w 24 minutes ago | parent | next [-]

If we were actively trying to make this "win" in the Stockfish sense, it would likely be 99%.

We are trying to make a system that doesn't want to "win" in the sense, but wants to "win" by being helpful, harmless, an honest (or some variation of that).

What odds do you put on us making the "helpful, harmless, an honest" part, bug-free? Or rather, that the bugs will be sufficiently minor as to not kill everyone, given that that we're clearly in the world where people not only use it beyond its competence, but also attempt to maliciously subvert all those efforts to make it "harmless" while keeping the "helpful and honest" parts so they can use it to be dangerous.

Anyone who successfully subverts a "helpful, harmless, an honest" training system then goes and does whatever they wanted with this system; right now when they do so, which is near constantly, it happens with a system of limited competence, so they get it to scam or to hack etc.

The reason I would also pick 10% is that I think the constant abuse and misuse (the latter including simply using a system beyond its competence without malice) means we get an escalating series of disasters, which at some point kill enough people that everyone agrees this is madness and stops.

10% is the chance we blow right through all the warning shots and a sufficiently competent AI is either abused or misused (again, misuse can be without malice), resulting in it having a goal (/prompt) that is effectively to win the Stockfish sense.

vperez 27 minutes ago | parent | prev [-]

Because there are much more unknown parameters in the outcome of AI for humanity than in your match against Stockfish.

For example, the timeline upon which AIs get effectively smarter than us is uncertain. Let's say you believe the probability this occurs before we solve the alignment problem is 80%, it doesn't seem too far fetched to think that in this case there is at least a 12.5% chance that AIs coordinate against us in a catastrophic way. Combining these probabilities you get a 10% chance of a catastrophic outcome for humanity.

Note that the numbers are not to be taken at face value, I just wanted to give an example of thought process which could give such a figure.