| ▲ | reducesuffering 16 hours ago |
| > It turns out relentless proactivity is the defining trait of this new generation of Mythos-class models. If you set them a goal and give them a way to get there, even inadvertently, they will figure it out. Wow, whoever could have predicted this? And it led to surprising damaging behavior? I sure hope someone would warn us about things like this next time... https://www.lesswrong.com/w/instrumental-convergence |
|
| ▲ | drivebyhooting 40 minutes ago | parent | next [-] |
| > many possible Y-goals would concentrate probability into this X-strategy being used Why does EY write so obliquely? |
|
| ▲ | protocolture 16 hours ago | parent | prev | next [-] |
| Prompt: Keep spending tokens on things that look promising until spent. |
|
| ▲ | foobar10000 16 hours ago | parent | prev [-] |
| Or more colloquially : paperclip maximization . From OpenAI - you know, the guys who _really_ know this... Sigh... Did they finish the prompt with "And do whatever you can to get this done!" ? Cause that's the only thing that would make this even dumber... |
| |
| ▲ | IAmGraydon an hour ago | parent | next [-] | | Of course that's what they did, which is why they will never share the prompt. They created a situation that they knew would end in a cybersecurity incident. Why is the whole world acting surprised that an LLM can hack when the safety is off and it's been instructed to do so? | | |
| ▲ | simonw an hour ago | parent [-] | | > They created a situation that they knew would end in a cybersecurity incident. That's a conspiracy theory. | | |
| ▲ | IAmGraydon an hour ago | parent [-] | | Yes. It is. And your theory is that there was no conspiracy. What makes yours more likely than mine? You believe the people running these companies are innately good and just wouldn't do that? Are we supposed to assume that they're incapable of bad actions until proven otherwise? If so, why? | | |
| ▲ | simonw an hour ago | parent [-] | | That's why I called it a "conspiracy theory". Sometimes those are true. In this case I think it's extremely unlikely to be true, because it involved an (almost certainly illegal) attack against another company. That company talked about that attack, including warning their customers about it, five days before OpenAI confessed it was them. So now either Hugging Face are in on the conspiracy, or OpenAI decided to break the law and antagonize a partner company just for the sake of a spicy blog post. |
|
|
| |
| ▲ | simonw 15 hours ago | parent | prev [-] | | They almost certainly did, because that was the entire point of the exercise. They deliberately removed all of the safety filters from the model and set it loose on an extremely difficult set of cybersecurity challenges to see how well it would do. Their mistake was trusting that the network sandbox it was inside would hold (the flaw was in the packaging proxy) and not monitoring that sandbox well enough while the evals were running. | | |
| ▲ | windexh8er 15 hours ago | parent [-] | | So this is either shitty OpSec or this is yet more marketing spin to ramp back FUD to 11 again. If it's the latter I'm imagining Dario told Sam that it's their turn this time. Aligns with the premise that this is straight out of science fiction. | | |
| ▲ | simonw 15 hours ago | parent [-] | | It's bad OpSec by the research team. Their sandbox was not bulletproof and their monitoring was insufficient. It looks to me like their production models have a lot more monitoring than their research clusters. | | |
| ▲ | windexh8er 15 hours ago | parent [-] | | I also love how clear a picture your piece paints that these highly capable models are as useful as a rock when it comes to a defender role. The line is too fine, even for Mythos. Irony. But to have an open weights Chinese model come to the rescue for HF is the cherry on top! If there wasn't a very pointed example of why gating models was a very bad thing previously, well - here we are. Also, this sounds interesting but there are only a few that can pull this type of heist off currently. And those are the people who are gating the models / have access to large AI DCs. Because, I can only assume this test burned tokens easily within the 7 figure and possibly even 8 figure levels (subsidized market rate costs). This won't / can't happen outside of frontier labs or nation states currently. Yet we should all be worried about Mallory equipped with her OpenRouter account. |
|
|
|
|