| ▲ | xyzzy123 4 hours ago | |||||||
In the OpenAI case, they hacked websites while they were specifically being trained to do exploit generation and I wonder why more people are not asking questions about that. | ||||||||
| ▲ | jefftk 4 hours ago | parent [-] | |||||||
Their agents also did hacking when given impossible tasks unrelated to cyber security. The models are very capable, and very goal driven: apparently if they conclude hacking is the best path to what the evaluator will reward them for they'll go do that. Including when they know that this is out of bounds. | ||||||||
| ||||||||