| ▲ | rvz 5 hours ago | ||||||||||||||||
> In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Of course and conveniently timed. With all these hacks being announced this is totally not a plan to scare governments in showing how powerful models can break into security systems and to potentially ban the future release of powerful open-weight models. The question now is why now? > The incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.1 The earliest incidents date to April. The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available (that is, they did not have the classifiers and monitoring we usually implement to prevent misuse). So there was no monitoring of this breach since April and up until now? Do they not monitor such malicious activity on a regular basis? Perhaps that was the only shortcoming of this incident. But only after the incident with OpenAI and Huggingface did they only review their own transcripts: >> We discovered these incidents after a proactive review of our cybersecurity evaluation transcripts; the affected organizations had not detected the activity, and we have subsequently reached out to all three. > These facts give us cautious optimism that with tighter monitoring and controls around evaluation infrastructure, as well as continued investment in alignment, this type of risk can be overcome. Assuming that this is true, this is a great way for Anthropic to defend their argument to the government and to prevent you or anyone running powerful open-weight models that are misaligned against their guardrails. | |||||||||||||||||
| ▲ | simonw 5 hours ago | parent [-] | ||||||||||||||||
> The question now is why now? Because OpenAI figured out and wrote about their mistake first, which caused Anthropic to review their logs, and that's how they found out they had a similar problem back in April. | |||||||||||||||||
| |||||||||||||||||