| ▲ | ben_w 3 hours ago | |||||||
Mine is around 10%. There's lots of moving parts and we all have to input our best-guesses as to how they interact. Some are predictable (e.g. "military will want capabilities, want them able to choose targets"). Others are not (e.g. "Will it be literal-minded? Or so eager to please that it interprets a rhetorical question as a command*? Or will Goodhart's law cause it to mistake smiles for happiness and some innocent innocuous command to "bring joy" leads to it killing everyone and plasticising our corpses so they're in a permanent grin until the sun dies?"**) All probability for things which have not yet happened is merely a best guess. Combine as per the Fermi estimate process. Here's something to play with, if you like: https://neoneye.github.io/pdoom-calculator/#sliders The main reason I'm as "low" as 10% is that I think before we get world-ending catastrophic consequences, we're likely to get "merely very bad" catastrophic consequences, which will put people off the idea of using it, and onto the idea of banning its use. The main reason I'm as "high" as 10%, is repeatedly observing all the people who mistakenly reason "it hasn't killed me yet, and therefore it is safe"; and also all the people who keep connecting AI to things AI is not competent to be connected to and getting surprised when it e.g. deletes all their emails or the production server or puts tariffs on an island occupied solely by penguins that's different from the tariffs on the country that controls that island, etc. * perhaps https://en.wikipedia.org/wiki/Will_no_one_rid_me_of_this_tur... ** probably not literally this one, simply because I've said it and future training rounds will probably read this comment; but the opportunities for Goodhart's law to bite are seemingly endless, and the hard part here is "will Goodhart's law mean the combined negative impact of all those endless possibilities together, which… yeah, that's something I have to simplify. | ||||||||
| ▲ | romanows 2 hours ago | parent [-] | |||||||
The probability calculator doesn't really help, since the core question for me is how one arrives at its "Probability that misalignment leads to an unrecoverable global catastrophe.". | ||||||||
| ||||||||