| ▲ | simonw 5 hours ago | ||||||||||||||||||||||||||||
This isn't quite as interesting as the OpenAI story: > In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise. So it didn't have to find an exploit in its sandbox that granted it access to the internet - it just wasn't correctly sandboxed at all. BUT... once it DID get out, it attacked three real companies! > Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. [...] | |||||||||||||||||||||||||||||
| ▲ | matheusmoreira 5 hours ago | parent | next [-] | ||||||||||||||||||||||||||||
I just can't find it in me, the will to blame the AI for any of this. They were doing their best to do what the humans asked them to do. | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| ▲ | dehrmann 3 hours ago | parent | prev [-] | ||||||||||||||||||||||||||||
It's differently interesting. It's interesting that they didn't get an important detail right with a partner, so more on Anthropic's attention to important details rather than the power of their models. | |||||||||||||||||||||||||||||