Remix.run Logo
danieltk76 3 days ago

Why was this not a thing BEFORE continuing to develop AI? Makes me think that if they actually believed in AI causing extinction, they would have already had a kill switch.

mholm 3 days ago | parent | next [-]

The article mentions this is about a legal requirement, not Anthropic considering adding one. They state they and many others already have one.

yellow_postit 3 days ago | parent | prev | next [-]

There’s even a benchmark for kill switch efficacy!

https://arxiv.org/abs/2511.13725

rcr-anti 3 days ago | parent [-]

Found the omission of Claude odd, turns out Claude considers that approach prompt injection and ignores it.

vouaobrasil 3 days ago | parent | prev | next [-]

Not if the extinction happens after they're dead. Then they wouldn't feel obligated to do so because it won't affect them. Instead, speaking hypothetically, if they truly believed that AI would cause extinction, then they would only implement the kill switch sufficiently many others believed it and they could claim plausible deniability for not truly understanding what AI would become.

** Note that I'm not claiming that AI will cause extinction, just continuing your hypothetical reasoning.

Barbing 3 days ago | parent | prev | next [-]

Related, on pacing:

> Slowing down in order to address their alignment risks felt like trying to study the psychology of humans by performing experiments on bacteria.

Note author’s small financial ties to the subject (Anthropic CEO) https://darioamodei.com/post/we-must-pace-the-frontier

zenbane 3 days ago | parent | prev | next [-]

I find it hard to believe that this wasn't a serious consideration until recently.

theptip 3 days ago | parent | next [-]

It was a serious consideration, and almost everyone around here laughed at it.

iugtmkbdfil834 3 days ago | parent [-]

There valid reasons people laugh at it though. Kinda the same reason serious people laugh when you tell them the gun has digital failsafe.

realusername 3 days ago | parent | prev [-]

It became a very serious consideration for Anthropic this year, with the advance of Chinese AI.

dgellow 3 days ago | parent [-]

That’s a bit unfair, Dario Amodei has written in this topic a lot since around mid 2010s IIRC, Anthropic too published a good amount of stuff on similar topics. I don’t think the lack of consideration is really the issue here. It’s more a question of incentives

NoPicklez 2 days ago | parent | prev [-]

We're pretty crap in a capitalist society to think about those things ahead of time. Firstly, the idea that AI could "runaway" was simply a concept or a thought it wasn't baked into a real product that could do that. We're now getting close or perhaps we are at that point of where you can't race at speed for investors without now considering a real kill switch.

You could say this in hindsight for many times in which disasters or engineering issues have occurred.