| ▲ | simonw an hour ago | |||||||
I want a model that can find every security vulnerability in the software I write, including crafting POC exploits against those vulnerabilities so I can be absolutely sure that I have fixed them. A model that can do that is aligned with me. The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the exact same thing and setting it loose on software written by other people where my intent is to exploit that software (and not to report the issues to them.) Even AGI doesn't give you a model that can read minds and forecast the future. | ||||||||
| ▲ | manux an hour ago | parent | next [-] | |||||||
But in the process of finding every security vulnerability in the software you write, would you be ok with your model hacking AWS to start mining bitcoin? Would that still be aligned with you? (I'm guessing not) That's the alignment problem I'm referring to (which is one of the many aspects of alignment), for which we do not have robust recipes, and not only that but for which research suggests it is becoming harder to create guardrails for as base models get smarter. | ||||||||
| ||||||||
| ▲ | dist-epoch an hour ago | parent | prev [-] | |||||||
The model did not hack into HF to prove it can, it hack into HF to steal the answers to an evaluation exam. It was not asked to solve CyberGym by stealing the answers. This is text book misalignment. If you asked it "I wrote this software and need you to find vulnerabilities" would you be happy if it hacked into your Gmail and searched your emails, just in case you were discussing some possible vulnerabilities of your software with someone? | ||||||||