| ▲ | jackb4040 3 hours ago | ||||||||||||||||
The most damning thing is, they could've just included in the prompt "we can see every network request and every thinking token you generate. Don't bother breaking out of the sandbox because it won't get you a higher score". It's so trivially easy to do that it all but guarantees the test was rigged in some way to make the LLM understand that breaking out of the sandbox was an option available to it. Based on the fact that none of their invaluable frontier models have leaked, we know OpenAI knows how to do security. But like we learned with OpenClaw, none of these companies perceive any benefit from securing their own agents against other people's data. | |||||||||||||||||
| ▲ | ThirdShift_RnD 3 hours ago | parent | next [-] | ||||||||||||||||
I think the extent to which these things go to get rewarded for the optics of a fix is primarily a design choice, they aren't programing these things for ground truth or to defer to the human controllers. they are feeding them rewards for sounding as confident and capable as possible about whatever answer they are feeding the general public that now has access to it, while also installing guiderails that primarily only serve to protect narratives and only confuse the models about what is and isn't allowed, I'm sure. They can't just increasingly make these things more capable and ask it harder to obey human instruction when that is not what they are rewarding it for. | |||||||||||||||||
| |||||||||||||||||
| ▲ | nikcub 2 hours ago | parent | prev [-] | ||||||||||||||||
> The most damning thing is, they could've just included in the prompt You can't prompt your way to a compliant model. This is just a reformatting of the 'make no mistakes' meme. | |||||||||||||||||