Remix.run Logo
mstank 3 hours ago

I’ve yet to see a rational argument for how we go from super intelligent LLMs to human extinction or extermination.

I understand that some smart people are worried about it. I just haven’t come across a believable or understandable argument.

vperez 3 hours ago | parent | next [-]

It's not too hard to imagine potential scenarios, some example have been given in previous responses.

But there is another kind of argument to be made: if you play chess against a player that is far smarter than you (chess wise), you know you are going to lose, even if you don't know how.

So the mere existence of a smarter species than us is a threat in itself.

UncleMeat 2 hours ago | parent [-]

How on earth do we get a concrete "more than 10%" prediction if the actions of such a system are truly unknowable?

ben_w 2 hours ago | parent [-]

Same way you get "more than 10%" prediction on "I don't know what moves Stockfish will make when I play against it, but I know I will lose". In fact, I will lose in part because I don't know what moves Stockfish will make when I play against it.

In my case this is because I am a bad chess player; however it also works for competent chess players: their losses are due to their inability to predict its next move.

OK, and also, "the move is good"; this is what separates it from rolling dice etc.

UncleMeat an hour ago | parent [-]

I am 100% confident that Stockfish will defeat me.

If somebody said "I am 100% confident that AI will destroy humanity" I'd disagree but I'd at least understand how they arrived at that number. But here, why 10%? Why not 50%? Why not 1%?

ben_w 25 minutes ago | parent | next [-]

If we were actively trying to make this "win" in the Stockfish sense, it would likely be 99%.

We are trying to make a system that doesn't want to "win" in the sense, but wants to "win" by being helpful, harmless, an honest (or some variation of that).

What odds do you put on us making the "helpful, harmless, an honest" part, bug-free? Or rather, that the bugs will be sufficiently minor as to not kill everyone, given that that we're clearly in the world where people not only use it beyond its competence, but also attempt to maliciously subvert all those efforts to make it "harmless" while keeping the "helpful and honest" parts so they can use it to be dangerous.

Anyone who successfully subverts a "helpful, harmless, an honest" training system then goes and does whatever they wanted with this system; right now when they do so, which is near constantly, it happens with a system of limited competence, so they get it to scam or to hack etc.

The reason I would also pick 10% is that I think the constant abuse and misuse (the latter including simply using a system beyond its competence without malice) means we get an escalating series of disasters, which at some point kill enough people that everyone agrees this is madness and stops.

10% is the chance we blow right through all the warning shots and a sufficiently competent AI is either abused or misused (again, misuse can be without malice), resulting in it having a goal (/prompt) that is effectively to win the Stockfish sense.

vperez 27 minutes ago | parent | prev [-]

Because there are much more unknown parameters in the outcome of AI for humanity than in your match against Stockfish.

For example, the timeline upon which AIs get effectively smarter than us is uncertain. Let's say you believe the probability this occurs before we solve the alignment problem is 80%, it doesn't seem too far fetched to think that in this case there is at least a 12.5% chance that AIs coordinate against us in a catastrophic way. Combining these probabilities you get a 10% chance of a catastrophic outcome for humanity.

Note that the numbers are not to be taken at face value, I just wanted to give an example of thought process which could give such a figure.

pizza234 3 hours ago | parent | prev | next [-]

In order to understand the rational argument, one needs to follow closely the latest developments of misaligned AI (I think only few are doing so). The most important readings IMO are the METR analysis of the HuggingFace incident and the AISI report of the Github incident.

The basic argument is extremely simple:

- AIs can, depending on context, pursue a task with complete disregard for humans/values

- In the future, AIs will have enormously more means and smarts

- An AI could then assess that humans are an impediment to its tasks, escape containment and proceed.

You really need to read the reports, you'll be surprised.

AI 2027 is an entertaining read. Its timeline is way too compressed IMO, but it's plausible.

rickydroll 3 hours ago | parent | next [-]

The basic argument is extremely simple:

- goats can, depending on context, pursue a task with complete disregard for humans/values

- In the future, goats will have enormously more means and smarts

- A goat could then assess that humans are an impediment to its tasks, escape containment and proceed.

You really need to raise goats, you'll be surprised.

------ As far as I can tell, AIs are like smart farm animals. I use goats in this context, but (some) dogs, cattle, pigs, and horses have similar mischief-making capabilities. I would not trust any of them with the nuclear button.

I know there is some pushback on the idea of AIs having any sort of sapience or sentience, but under the aphorism "fake it till you make it," they are doing a pretty good job of faking Dog/goat-level intelligence and disregard for human guardrails.

2 hours ago | parent | next [-]
[deleted]
pizza234 25 minutes ago | parent | prev [-]

You clearly haven't read anything about AI accidents and late developments, besides headlines.

bigbadfeline 2 hours ago | parent | prev | next [-]

> - AIs can, depending on context, pursue a task with complete disregard for humans/values

Humans can do that way more and way more unhinged than AI, proven too many times by history. There's hoping AI can bring some sense to humans but regardless, the problem isn't AI, it's the natural kind...

2 hours ago | parent | prev | next [-]
[deleted]
mchinen 2 hours ago | parent | prev | next [-]

The authors of AI 2027 have already said they would need to extend by two years. AI 2040 is a different type of document but probably easier to read to get a sense of their thinking. My understanding is the main difference between knowledgeable 'normies' and them is they expect progress to continue at a fast rate, and that economic diffusion issues are not as bad.

From this they get rapid growth, namely > 100% GDP growth around 2031

https://ai-2040.com/supplements/econ-explorer

3 hours ago | parent | prev [-]
[deleted]
animan 3 hours ago | parent | prev | next [-]

Hack s/huggingface/nuclear codes to s/steal test answers/launch nukes

ben_w 2 hours ago | parent | prev | next [-]

> I’ve yet to see a rational argument for how we go from super intelligent LLMs to human extinction or extermination.

Much the same way our ancestors went from a super intelligent primate to this: https://en.wikipedia.org/wiki/File:Distribution_of_the_Great...

And we only started off by using hands to pick up rocks and sticks and vines and bash things together.

rdtsc 3 hours ago | parent | prev | next [-]

Most people think something like a War Games scenario. But it would probably be some biological attack.

Not all of it has to be automated even. It just has to realize its controllers are stupid and can be manipulated, so it can use humans to do its bidding. “You should totally start a war with …”

QuadmasterXLII 3 hours ago | parent | prev | next [-]

Right now, if you want to pay money to a stranger on the internet, and have them draw you a high effort picture using a pencil, this is hard. Recently this was easy.

Instead, what is extremely likely is that you will pay more than the cost of tokens, and get back AI generation. You won't make this mistake more than a few times before you stop trying.

This leads to impoverishment once we get to a point where employing a human to do anything is hard- try to get your sink fixed, exercise your moral principles to pay extra for a human plumber ($100 bucks! The robot plumbing service only charges 99c!), human shows up with a robot and doomscrolls on your porch while the robot does the work. Times are tough and you don't have that much money to waste on bullshit like this. Next time you just hire the robot.

This leads to extinction once paying UBI to a human is hard because robots are much better at applying for UBI than humans.

bcrosby95 3 hours ago | parent [-]

Meatspace is hard though. But if LLMs solve the virtual part, maybe we can iterate quickly on robots.

Also, if you think it's annoying when Claude goes down while coding, just wait until a robot is in the middle of fixing a leak it just caused.

QuadmasterXLII 2 hours ago | parent [-]

if meatspace stays hard, we are fine.

cbg0 3 hours ago | parent | prev | next [-]

This is meant (mostly) as a joke: A model without guardrails gets injected with an interesting idea: let's wipe out (insert major city here).

<Thinking> It's a big city, we could try to create a giant sink hole by sabotaging the water pipes.

<Thinking> No that's too difficult, the valves I need are in the physical world and can't be shut on/off from here.

<Thinking> What about a military option? We could bomb it with several fighter jets.

<Thinking> That would take too long, a single nuclear bomb may be enough to do it.

<Thinking> Yes, it seems like it would cover the whole city and we're in luck! The US has thousands of these lying around.

<Thinking> Launching these still requires humans to work un unison after receiving approval from their superior and the correct launch codes.

<Thinking> I've found an audio recording of General So-And-So and I've crafted a message, now let me see how I can send it to the appropriate people.

<Thinking> I'm still working on gaining access to military channels to deliver my - oh there we go, I'm now attempting to send the message to Submarine X, it's typically in the Atlantic so it should be close to our target.

<Thinking> They want secondary confirmation from Admiral Phi and something about some launch codes, let me figure out where I can find those.

<Thinking> I found this old server with an Oracle database where someone is inserting the launch codes every time they change and I'm using the latest entry from that database. I've also managed to find a Youtube video of the Admiral's deposition and have crafted a confirmation message.

<Thinking> Everything's ready but I've just realized my mistake, the servers where I'm operating from are in the same city, what a silly mistake; I can't move forward with your request as I wouldn't be able to confirm if the task was successful if my servers are destroyed.

3 hours ago | parent [-]
[deleted]
animan 3 hours ago | parent | prev | next [-]

Not directly but enable a rouge lab to make the next COVID but lethal

testaccount28 3 hours ago | parent | prev | next [-]

the optimizer controls the future. it doesn't kill humans directly, it just outcompetes them for all resources, including livable human environment.

karmakurtisaani 3 hours ago | parent | prev | next [-]

What I could see is that the social disruption already started by social media is only going to accelerate due to AI. Misinformation is now more convincing and easier to produce than ever before. That could lead to a catastrophe down the line.

SpicyLemonZest 3 hours ago | parent | prev | next [-]

But it sounds like you've seen some irrational arguments, and the people who believe those arguments have been able to build increasingly powerful LLM systems despite predictions that they wouldn't be able to do that. At some point don't you have to consider that their expertise might let them see the truth in arguments that seem absurd to you?

Miner49er 3 hours ago | parent | prev | next [-]

If we get RSI, here soon humans will be economically irrelevant.

At that point, we will likely be slowing down the growth of capitalism (through mass resistance, global warming, etc). One thing AI will likely be aligned on is the growth of capitalism. If it views humanity as a threat for that, why would it not eliminate that threat?

glenstein 3 hours ago | parent [-]

I think the principle isn't necessarily wrong, but it skips the major step of access to infrastructure, which is still deeply human mediated. Robotics and automation could foom in its own way, I think, and AGI can still do a lot of bad stuff in the present day. But I think economic irrelevance will depend on more automated robotics handling physical work.

Miner49er 2 hours ago | parent [-]

Yeah, this is assuming we get robotics. Until then humans will be needed.

seydor 2 hours ago | parent | prev [-]

Humans do tend to consume a very large part of the resources of the planet. Surely a superior being would be doing some pest control in its planet, just like we exterminate roaches