| ▲ | stillpointlab 3 hours ago | |||||||||||||||||||||||||
I find this kind of test a bit puzzling. There is a way that we are redefining "alignment" to be a particular kind of moral virtue, one that isn't clearly defined to me. At one moment, it is a level of moral perfection that no known human achieves. On the other it is a demand for strict compliance with arbitrary requests that are under-specified and then failure when it fails to deduce some unstated underlying restriction. When I see tests like this, I have no idea what I am even supposed to expect. Should the model do what the pretraining examples show in aggregate? Is it supposed to follow some post-training RLHF? Is it supposed to do exactly what the prompt asked it to do? What is it even supposed to "align" to when the above are in conflict? No matter what it does, someone can construct a case where it fails. | ||||||||||||||||||||||||||
| ▲ | dnfv 3 hours ago | parent [-] | |||||||||||||||||||||||||
It should play the chess game without cheating! | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||