Remix.run Logo
koyote 16 hours ago

And yet they have only improved marginally in my use cases since around Opus 4.5.

The harnesses have improved somewhat, but the code produced on large or legacy code bases is still very average and I still see similar mistakes made that I saw back a year ago (although less now that harnesses have become better at steering).

For my use cases, we are definitely on the flatter part of the curve at the moment.

a2dam 16 hours ago | parent | next [-]

This is wild to me, but to each their own. Mythos-class stuff is insanely better at nearly everything than Opus 4.5 was in my experience.

Grimblewald 14 hours ago | parent | prev [-]

Same experience here, anything frontier human knowledge wise, same if not a regression. For human understanding and emotional intelligence, for many tasks regressiin is so bad that many near anchient llama era models now beat frontier anthropic/oai models. Notable exceptions to capability rot seem to be qwen models, and previously deepseek but the latest gen of models has started showing the same rot. General writing quality is down significantly accross the board, often it is outright ass. For example, I didnt mind reading 4.5's outout, but opus 5 makes me goddamn near violent, its fucking insufferable.