Remix.run Logo
ramoz 2 days ago

It crushes Fable on benchmarks and even in the blogs "real-world" studies. But... they are communicating like it ~sometimes~ provides Fable intelligence?

A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??

bitexploder 2 days ago | parent [-]

What if the recent Fable intelligence regression was basically just them serving Opus 5.5 until they got it working well?

leothetechguy 21 hours ago | parent [-]

I have never been a part of the group that believes in frontier labs downgrading models. But this theory seems plausible to me for the first time.

bitexploder 17 hours ago | parent [-]

I guarantee they are at least tweaking quants, caching systems, and finding ways to move serving costs down. This definitely impacts the model’s intelligence at times. There are also a lot of model tweaks, RLHF rollouts etc. I don’t think it means they are doing anything malicious or deceptive. And if Opus 5.5 is literally smarter than Fable? Ehh, it is plausible :)