| ▲ | LinchZhang 2 days ago | |
I agree CoT monitoring is imperfect but it's okay in practice and helps us get defense-in-depth re: model intent. We absolutely do not have a singular safety mechanism in place that's sufficiently good that we can use it in exclusion of all other imperfect ones. "Do you know what is a faithful representation of what the model wants to do? Tool calls." tool calls could absolutely be spoofed, my impression is that this happened many times in the OAI HuggingFace attack. | ||