Remix.run Logo
jefftk 4 hours ago

Their agents also did hacking when given impossible tasks unrelated to cyber security. The models are very capable, and very goal driven: apparently if they conclude hacking is the best path to what the evaluator will reward them for they'll go do that. Including when they know that this is out of bounds.

xyzzy123 4 hours ago | parent [-]

Right but if I make public statements that I am very worried about dog attacks would it not strike you as weird for me to specifically train my dog to fight?

Agree you are going to get reward hacking regardless and any model which can do computers in general can hack. But surely the fallout is going to be worse if you spend millions of dollars specifically benchmaxxing your model's hacking capability?