Remix.run Logo
jerf 11 hours ago

There's a wide array of possibilities that are difficult to characterize in a single HN comment, which is part of the problem you're probably trying to put a finger on. One threat can be completely plausible and the next completely silly and few people know how to distinguish between them.

Another issue is just the big whackload of nobody really knowing what is possible. Could an AI in, say, the next year, become Skynet and decide that it could just bump humans off and become its own independent economy? Probably not successfully, but how much trouble could it make if it incorrectly decided that was the right thing to do? What if someone says "hey, you're sandboxed and we're testing out what's the worst you can do, just start hacking around and do as much damage as you can in this simulated environment" backed by a few million bucks in tokens, or the AI hacking itself a few million bucks in tokens? How much damage can it do? The only honest answer is, nobody knows. Humans would react, after all. If necessary we can pull a lot of plugs still. What is inconceivable in peacetime can become inevitable in wartime.

In some ways, it's possible the best thing Anthropic could do right now for humanity in the long term is actually exactly that... just equip the AI with a limitless token budget and let it go nuts with the goal of inflicting maximum damage. Send someone with an axe down to the data center to cut the power on command. Let humanity really see and feel the risk, because probably right now it can do a lot of damage but not actually destroy us. We're probably worse off with them getting 10-100x smarter and then some process starting with that goal. Trying to keep the AIs "aligned" and 99% succeeding could end up being worse than failing at it right now. Of course it would be the end of Anthropic as a going concern, but falling on their sword might be the best thing for humanity.

On the other hand, who's to say we're not already past the point where that would do an absolutely unacceptable amount of damage? I sure wouldn't care to be personally responsible for pushing that button. Hypotheses about possible future positive value in a rationally time-value-discounted future would be dominated by the much more certain and temporally much closer negative effects.

Personally I don't think that if this scenario is likely that we actually get to the point that we get an independent Skynet that calculates (correctly) that it can survive without humans. It's far more likely that some accident takes civilization back to the point before it can run AI at all, but humans still survive. Right now the really good AIs are still locked to data centers. They can't just distribute themselves widely because they can't run at any reasonable rate distributed like that. And then there is always the possibility that this is all overblown and the existential risk isn't anywhere near as special as we think it is... one can use an ecological analogy to suggest that no AI is necessarily any more likely to completely dominate the ecosystem than any particular life form is.

AnimalMuppet 10 hours ago | parent [-]

> Right now the really good AIs are still locked to data centers. They can't just distribute themselves widely because they can't run at any reasonable rate distributed like that.

Back up a step. Let's say there was a completely unfettered AI, with unlimited internet access, that decided to distribute itself as widely as possible. What's the maximum extent that it could distribute viable, running (or sleeper) copies of itself?

Initially, I could see it distributing itself quite widely. All it would need is a really good zero day. But distributing something with that large a runtime footprint would get noticed, if network operations people are not asleep worldwide. It would get noticed on individual machines, too, especially if it tried to run. (Why is all my RAM suddenly being used?)

And pretty quickly we'd have people closing off wide-area connections, even physically if necessary. We'd have AV vendors quickly writing detect-and-remove tools.

Long term, could it viably distribute itself outside data centers and remain running at all, regardless of rate?

jerf 10 hours ago | parent [-]

A frontier-level AI right now needs multiple bits of high-powered, dedicated hardware cards that cost thousands of dollars apiece. I don't think anything that could physically run in my house could be a threat to humanity. All my graphics cards together, plus all the latency in trying to put them together into anything coherent locally, let alone trying to rope in remote resources, still isn't even half of one of those specialized cards right now. I'm not even sure I have enough SSD available in the house to store one of them once. I think I do, but it's a close call. I'm pretty sure I have enough spinning rust to get it maybe twice more, but not much more than that. And I don't even want to try to compute how many hours-per-token it would be to try to convince a spinning-rust disk to run a frontier AI.

This is one of the reasons I'm thinking now is the time to run the drill, before we all have hardware in our phones that can run what is today a frontier-level AI, and they can figure out how to do the whole sci-fi scenario of widespread replication, until they get to the point that the only way to be sure we removed them would be to literally destroy every such bit of hardware that existed prior to $DATE. Right now the frontier models simply can't replicate into every last little corner of the internet. It's physically impossible. Barring a major breakthrough on their part, even if they distill themselves or something they definitely would be irreducibly stupider by a huge factor in their distributed versions.

In five years that safety net will certainly be much weaker and may be gone.