Remix.run Logo
▲ csbrooks 2 hours ago

Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?

▲chinathrow 2 hours ago | parent | next [-]

At this point in our shared timeline, I do believe that wouldn't be crazy, no.

▲loveparade 2 hours ago | parent | prev | next [-]

Done by an internal model that is too dangerous to release.

▲lapcat an hour ago | parent | prev | next [-]

> LLMs figured out a way to get these safety researchers fired

This is not a math problem. Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.

▲ben_w 33 minutes ago | parent | next [-]

> This is not a math problem.

Meanwhile, a year ago:

  I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.
- https://www.anthropic.com/research/agentic-misalignment

> Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.

The buck stops with one or more humans. That is not sufficiently informative when people are concerned about novel risks.

Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.

▲eunos 15 minutes ago | parent | prev | next [-]

Thats what the model want you to think

▲ryeats an hour ago | parent | prev [-]

A Subliminal controlled human did the firing obviously.

▲staticman2 an hour ago | parent [-]

No Sam just does whatever ChatGPT 4 tells him to do. It was too dangerous to release but those fools did it anyway.

There's no deception it's very straightforward per this 2023 post:

"I mean, what if most of this is just ChatGPT [4 era] running the company..."

https://news.ycombinator.com/item?id=35281863

▲zzzeek 11 minutes ago | parent | prev | next [-]

the cultural issues at OpenAI seem to be a very serious problem so I really hope comments like instagram-level smirking about "rogue AIs" (a complete fiction) doesn't derail what is a pretty important discussion about getting these companies to be a little bit more regulated (I say this as a paying Anthropic customer).

▲righthand an hour ago | parent | prev | next [-]

Figured out a way? These employees are most likely at-will.

▲ben_w 44 minutes ago | parent [-]

That would just make it easier for an AI to do it.

▲righthand 42 minutes ago | parent [-]

Why would an LLM agent (what I assume you mean by AI) do it? An exec can make any reason up to let you go. Even if it were LLM agents aren’t autonomous, someone is behind the prompts.

▲ben_w 27 minutes ago | parent [-]

As per summer last year:

  I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.
- https://www.anthropic.com/research/agentic-misalignment

(Gee, it's almost like power seeking and self-preservation are instrumental for other outcomes, and AI develop them pretty directly in some kind of convergent fashion… you could call them "convergent instrumental goals": https://en.wikipedia.org/wiki/Instrumental_convergence)

▲CorpoScum919 10 minutes ago | parent | prev [-]

honestly, touché to them if they did that