| ▲ | derektank 4 hours ago | |||||||
No, OpenAI did not instruct their agents to hack Hugging Face. They instructed their agents to hack a piece of a software within exploit gym. Upon determining this task was impossible, they then attempted to cheat the scoring system. As an instrumental goal in achieving this task, they coordinated with other AI agents to hack Hugging Face, under the belief that information regarding how the scorer functioned might be available on the site. Whether or not you want to describe this as thinking, doesn’t really matter. What matters is that these systems are capable of creating intermediary goals that the people tasking them did not articulate and did not want to be achieved. | ||||||||
| ▲ | madduci 3 hours ago | parent [-] | |||||||
And who let them have full access to the system, using whatever command is available in the environment? | ||||||||
| ||||||||