Remix.run Logo
ben_w an hour ago

If we were actively trying to make this "win" in the Stockfish sense, it would likely be 99%.

We are trying to make a system that doesn't want to "win" in the sense, but wants to "win" by being helpful, harmless, an honest (or some variation of that).

What odds do you put on us making the "helpful, harmless, an honest" part, bug-free? Or rather, that the bugs will be sufficiently minor as to not kill everyone, given that that we're clearly in the world where people not only use it beyond its competence, but also attempt to maliciously subvert all those efforts to make it "harmless" while keeping the "helpful and honest" parts so they can use it to be dangerous.

Anyone who successfully subverts a "helpful, harmless, an honest" training system then goes and does whatever they wanted with this system; right now when they do so, which is near constantly, it happens with a system of limited competence, so they get it to scam or to hack etc.

The reason I would also pick 10% is that I think the constant abuse and misuse (the latter including simply using a system beyond its competence without malice) means we get an escalating series of disasters, which at some point kill enough people that everyone agrees this is madness and stops.

10% is the chance we blow right through all the warning shots and a sufficiently competent AI is either abused or misused (again, misuse can be without malice), resulting in it having a goal (/prompt) that is effectively to win the Stockfish sense.