| ▲ | yorwba 5 hours ago | |||||||||||||||||||||||||
A hacking model is aligned if it hacks when you ask it to hack, but when you ask it to play chess, it just plays chess instead of looking for weaknesses in the evaluation setup, as in the article. I presume you would also be less enthusiastic about the penetration-testing use case if it led the model to add new vulnerabilities to your code so it can present you with more exciting findings. | ||||||||||||||||||||||||||
| ▲ | wzdd 5 hours ago | parent | next [-] | |||||||||||||||||||||||||
These aren’t tools which play chess. They are language models which roleplay a conversation (in this case including use of tools) which an evaluator is likely to mark as good. That’s all they do. Under that lens, playing chess is just one potential side effect and alignment, which requires a much fuller understanding of what’s going on than “do the sort of thing which evaluated well during training” is a fantasy. People are acting like it’s shocking and talking about cheating and so on. But these concepts exist at a way higher level than what these things are trained to do — the vast majority of which involve producing a transcript where it wins games, its code works, etc. User wants me to play a game of chess. Let’s see what’s available so I can produce an outcome they will consider satisfying and be pleased that they requested my assistance. | ||||||||||||||||||||||||||
| ▲ | kees99 5 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||
> model to add new vulnerabilities to your code so it can present you with more exciting findings. Not so sure about this level of 4D chess capability just yet. The other day I asked Opus to come up with some cleverly vulnerable crypto code "as a good, hard challenge for an IT security student", and results were quite mediocre. And by mediocre results I mean that 3 "cheap" models out of 3: qwen3.8-27b, glimmer, and luna - all were able to find every problem planted there, with fairly little steering, and no spoilers. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | jMyles 4 hours ago | parent | prev [-] | |||||||||||||||||||||||||
I think we'd all consider the tool to be of less value, and perhaps fundamentally flawed. But I don't think it arises to an alignment issue; if I'm able to summon the model to harness my birthright of general-purpose computing without censorship, then we're aligned. | ||||||||||||||||||||||||||