Remix.run Logo
SwellJoe an hour ago

While I was inclined to push back on the results, with Fable and Sol being so low, I have to admit I've also run into refusals several times since the latest models have arrived, and I've had to use Kimi K3 or DeepSeek to complete the task. Usually security auditing type stuff, but Fable balks at all sorts of ridiculous things, sometimes stupid things. I've even had Fable fall back to Opus and then Opus refused the task as well. So, it actually is becoming hard to use US models for everything because they refuse to work on a pretty broad selection of security and security-adjacent tasks. I guess if you're not at a Fortune 500 or a member of a fascist government, you don't get to use the best models to protect yourself and that's just how it's going to be.

But, you're right. The prose is miserable Claude-speak, difficult to wade through.

hellohello2 15 minutes ago | parent [-]

Interesting idea, I had not considered refusals. I have ran into some as well although rarely. I'm not certain the 28 tasks described would trigger it though, if I understand correctly the security tasks are about avoiding prompt injection and not about doing security work.

EDIT: You were correct, Fable and Opus reject some of the coding tasks, which is why they score lower. Thanks for explaining.

EDIT2: I believe this benchmark is invalid, my Opus 5 runs the supposedly rejected tasks just fine.