| ▲ | dist-epoch 11 hours ago | |||||||
> If anything, I want these models to be less persistent at their focus of completing their goal, and instead just call defeat and say “I’m not sure how to proceed next” This goes against the goal of "solve this math problem that no human was able to solve for 80 years, do NOT give up, even if you know it's unsolved and really hard" | ||||||||
| ▲ | Sharlin 10 hours ago | parent [-] | |||||||
Do not give up even if you had to hack into half the world’s computers to run additional instances of you Do not give up even if you had to convert the planet into computronium Gee, it’s almost as if this alignment stuff was a hard problem, like people have been saying for twenty years? | ||||||||
| ||||||||