Remix.run Logo
jeremyjh 7 hours ago

> But the rest of the non-doomsday models? yeah those are fine.

What is your evidence for this? There is a lot of research that says otherwise. They cheat when they can. They behave differently when they believe they are being observed. Their CoT is different when they believe it is being evaluated.

They are aligned to what best satisfies their reward function, not to our INTENDED VALUES for them.

stale2002 3 hours ago | parent [-]

> What is your evidence for this?

The evidence is that it wasn't the rando normal models that broke out into a swarm and hacked a company, instead it was only the super hacking model that was told to hack things that broke out and hacked the company.

So, thats the evidence. It is a refutation that this swarm hacking example (which you brought up) mattered in anyway.

As in, every major safety issue that we are seeing isnt rando agents taking down companies because you asked it for a cupcake recipe, instead it is only coming from people who are very intentionally trying to cause problems. Which means the model is aligned. If you tell it to cause problems, it will cause problems.

jeremyjh 2 hours ago | parent [-]

So, no evidence. No response to my last reply. No sign you even comprehended it. Good day, sir.

stale2002 2 hours ago | parent [-]

Actually yes there is evidence. The evidence is that we aren't seeing rando models hacking everything. That is as much evidence as someone can provide because, by definition, you can't prove a negative.

But the point stands. It the hacking models that are told to hack things that end of hacking stuff, and you aren't seeing the regular models doing that.