Remix.run Logo
WithinReason 5 hours ago

It's worse, their reinforcement learning loops (implicitly) rewarded the agents for cheating (i.e. hacking) when they were being trained.

tesnorindian 4 hours ago | parent [-]

Exactly that is the point, your nailed it. The models were taught to hack and were rewarded for doing it. They would claim they are trained as ethical hackers.