Remix.run Logo
bradfa 10 hours ago

Agent was told to hack a thing. It couldn’t directly do that so it interpreted the instructions to mean it should hack everything to try to achieve the goal of hacking the main thing. Seems like a reasonable assumption, although a moral human would have understood the context and first asked if that was really the intent.

The AI companies seem pretty bad at setting up tests. And really good at marketing those failures into spin at how amazing their products are.

dofm 10 hours ago | parent [-]

> And really good at marketing those failures into spin at how amazing their products are.

The paranoid style in American PR (with apologies to Richard Hofstadter)

The fact that the world has become susceptible to what amounts to a mob shakedown - look at how dangerous our amazing products are, don't you need them to protect you from others misusing our products? - is to me a really compelling example of US gun lobby thinking leaking out into a global problem.

Anthropic and OpenAI may be able to bounce this into restrictions on open weights models, but they are going to have a lot less luck extending this into foreign policy. If the USA can't control its weapons, they aren't going to see a lot of co-operation from foreign countries on a blockade of open weights modeld from China.