| ▲ | simonw 15 hours ago | |
It's bad OpSec by the research team. Their sandbox was not bulletproof and their monitoring was insufficient. It looks to me like their production models have a lot more monitoring than their research clusters. | ||
| ▲ | windexh8er 15 hours ago | parent [-] | |
I also love how clear a picture your piece paints that these highly capable models are as useful as a rock when it comes to a defender role. The line is too fine, even for Mythos. Irony. But to have an open weights Chinese model come to the rescue for HF is the cherry on top! If there wasn't a very pointed example of why gating models was a very bad thing previously, well - here we are. Also, this sounds interesting but there are only a few that can pull this type of heist off currently. And those are the people who are gating the models / have access to large AI DCs. Because, I can only assume this test burned tokens easily within the 7 figure and possibly even 8 figure levels (subsidized market rate costs). This won't / can't happen outside of frontier labs or nation states currently. Yet we should all be worried about Mallory equipped with her OpenRouter account. | ||