| ▲ | nradov 2 hours ago | |
It's not technically possible to prove model alignment. At best you can establish a probability of alignment within certain constraints. | ||
| ▲ | madrox an hour ago | parent [-] | |
"Prove" is a shorthand because of course it's a stochastic process. My point is that I would expect ethics to be part of the training and eval criteria. | ||