Remix.run Logo
6thbit 4 hours ago

Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed.

Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.

HarHarVeryFunny 3 hours ago | parent | next [-]

It seems they are trying to thread a needle here - they want to say it's very strong, but apparently this time do not want to invite extra government scrutiny.

They do say that (implicitly unlike Mythos) Opus 5 was not trained to exploit software vulnerabilities, which would certainly make it safer in that regard.

"As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats."

square_usual 4 hours ago | parent | prev | next [-]

Easy enough to explain: they're benchmaxxing. Fable is intelligent but not benchmaxxed. Opus is less intelligent but benchmaxxed.

llelouch 3 hours ago | parent | next [-]

Yep , same with 5.6. Fable is still the best.

lifty 2 hours ago | parent [-]

But still nerfed compared to the initial release.

solenoid0937 11 minutes ago | parent [-]

Only when you hit the cyber classifiers.

6thbit 2 hours ago | parent | prev [-]

Honestly that's the simplest explanation and thus likely the correct one.

gallerdude 4 hours ago | parent | prev | next [-]

Capable in term of AI R&D, not capable in terms of hacking (which caused all the Fable drama.) But agree, confusing wording.

flakiness 3 hours ago | parent | prev [-]

Maybe they don't want to say that to avoid the government scrutiny.