Remix.run Logo
andxor 5 hours ago

> Performance is significantly higher than Fable 5.1

That's not clear. Need to see independent benchmarks first.

forgot-my-pw 5 hours ago | parent | next [-]

We need them pelicans on bikes.

bwat49 5 hours ago | parent [-]

Its time to move on to the flamingo on a unicycle bench

andxor 5 hours ago | parent | prev | next [-]

Artificial Analysis just published their aggregate score (61).

Still below Fable 5, let alone Fable 5.1.

EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.

timpera 5 hours ago | parent | next [-]

I agree, Opus 5 scoring higher than Fable 5 on Artificial Analysis really makes me question the relevance of these scores.

CamperBob2 3 hours ago | parent [-]

There is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy.

If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.

natsucks 4 hours ago | parent | prev [-]

I saw this too and I'm really confused.

forgot-my-pw 5 hours ago | parent | prev [-]

AA benchmark: https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...

TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.