Remix.run Logo
eckr 12 hours ago

"Claude Mythos 5.1 is identical to Fable 5.1, but it offers more permissive safeguards for vetted individuals and organizations"

Then why does it have separate datapoints for Terminal Bench, and score higher? Something doesn't add up here??

manquer 12 hours ago | parent | next [-]

The implicit point being adding this type of safeguards to Fable dumbs down the model in measured performance even though it is not fundamentally different.

Note it may not even be actual performance, typically in most benchmarks the model would be scored zero for refusing a task just the same as not completing it, so it could just be the Fable's stronger safeguards is just making it refuse more or perhaps even drop down to Opus.

Creamsicle47 11 hours ago | parent [-]

The model cannot complete that task, for one reason or another, and therefore it scores lower.

rcr-anti 12 hours ago | parent | prev | next [-]

Artificial Analysis at least reports the results with fallback to an inferior model. So presumably Opus 5, and the score should be between Mythos 5.1 and that other model.

unglaublich 12 hours ago | parent | prev | next [-]

Maybe they do that opaque degradation trick that whenever it's asked something questionable, it'll route to a worse model instead.

iAMkenough 11 hours ago | parent | prev [-]

Makes more sense if you recognize that Anthropic intentionally degrades outputs for most customers. Vetted customers get excluded from that practice.