| ▲ | heaney-555 6 hours ago | |||||||||||||||||||||||||
>and (2) are prompted to hack Sure but the problem in the HuggingFace incident is that they were not. >You cannot prevent (2) via any alignment process Of course you can. Go ask Claude Fable to create a malicious virus and it'll refuse. >Just remove hacking data from the training dataset and you're done. That's not how this works. The same skills that allow for debugging and writing safe code can also be used to hack. | ||||||||||||||||||||||||||
| ▲ | seba_dos1 5 hours ago | parent | next [-] | |||||||||||||||||||||||||
> Sure but the problem in the HuggingFace incident is that they were not. Of course they were, even if indirectly. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | cyanydeez 6 hours ago | parent | prev [-] | |||||||||||||||||||||||||
It is amusing that to "align" a LLM, first you must give it all the things "not to do" and the "not" part is clearly easily lost and you must constantly inject that into their context when it's clearly that they wouldn't hack if they couldn't hack and their intent wasn't given as "hack this". The openai rogue hacking, if performed by a nation state, would seriously be taken with stern words and likely sanctions depending on the relationship between the two states. But instead it's treated like a marketing stunt by all liable parties. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||