| ▲ | k9294 5 hours ago | |
The part that scares me the most is that OpenAI researchers who manage this experiments sometimes (according to the HF hack investigation) don't know what agents do.. So they run RL to reinforce this unknown behavior (lying/cheating/hacking) and god knows what else... And if this already happened at least once, how many times it has already happened and was “accidentally” added to the main model? | ||