Remix.run Logo
bigglebear 3 days ago

The issue with all of these is that we already know theres an incentive for labs to lie and make up fanciful stories (and Anthropic already does exactly that and has been doing that for a long time), and there's no way to verify any of their claims as being genuine. Even if we want to assume good faith, it doesn't mean we're gauranteed accurate reporting or accurate analysis. There are no repercussions for security incidents so no reason for them not to misuse this process if it benefits their agenda. There's no government agency (unbiased third party - which is why we can't rely on companies like METR) validating claims or providing confirmation of accurate reporting and that they are not misleadingly framing or representing an incident.

What were the system prompts? The full chat log? What was the model trained on? How was it RL'd and with what data? How was this incident uncovered, and what triggered it? You can't make any useful conclusions at all without the full picture.

They say "we investigated X and found no case of Y" - okay, and we're to just trust your judgement? How about you provide us with the data and we can assess for ourselves.

This is all quite pointless and achieves very little.