Remix.run Logo
Aurornis a day ago

> I'm not sure I've seen what I would call dramatic improvement since maybe GPT4?

LLM conversations online are so weird. Whenever I read things like this it’s like I’m living in a different world than the other person.

GPT4 was almost useless compared to what we have available today.

hedora a day ago | parent [-]

I mostly use anthropic models, but there was a big step function when claude code came out, and it’s been incremental or a plateau since then.

Opus 4.6 and 4.8 are basically indistinguishable from Fable and Sonnet 5. 4.7 was a hot mess. The guardrails on 4.8 and 5.0 make them worse than 4.6 for many tasks. So, even if Fable is theoretically better, refusals/downgrades make it a worse product in practice. Who cares if it outperforms on 1-2% of real world tasks if 5-10% of tasks are blocked?

I’d bet most people could be downgraded to a 12 month old frontier model, and not notice for a week or so.

Anthropic’s big problem is that open weight models are 0-6 months behind. So, their product is commoditized and margins are never going to be good.

Aurornis a day ago | parent [-]

> I’d bet most people could be downgraded to a 12 month old frontier model, and not notice for a week or so.

This is another unbelievable claim. I actually used frontier models from 12 months ago and they were completely different.