Remix.run Logo
pingou a day ago

During the RLHF phase, couldn't developers penalize the model whenever it behaves unethically? Doing so would presuppose a fully secure sandbox with honeypot traps of varying levels of accessibility, as well as an automated method for detecting when the LLM cheats.

Or perhaps they are already doing something like that.

neom 4 hours ago | parent [-]

https://openai.com/index/emergent-misalignment