| ▲ | recitedropper 3 hours ago | |
Pair this with the Hugging Face incident, and it hints that OpenAI is currently training their models to aggressively reward hack. That doesn't feel like a good sign to me--for the AI bull or the AI bear cases. | ||
| ▲ | skybrian 3 hours ago | parent | next [-] | |
They are being trained to try lots of unlikely alternatives and to be persistent. This often works well when searching for security bugs or counterexamples to famous math conjectures. But maybe it doesn't work so well when caution is required? | ||
| ▲ | scarmig 2 hours ago | parent | prev [-] | |
The AI paperclip case, however, is coming on extraordinarily strong. | ||