Remix.run Logo
▲ lnkl 3 hours ago

>This article would be 100% correct if it came out 1 year ago, 75% correct 9 months ago, 50% correct 3 months ago and it's probably 25% correct now if not less.

Feels like I read comment similar to this one each year since 2023.

▲throwawayffffas 3 hours ago | parent [-]

I can tell you it was not correct in January of 2026. The real swift started happening with the latest models opus 4.7, fable 5, kimi k3, glm 5.2.

That's when the models started to be coherent enough for real work.

They still fuck up, but it does not feel the code was written by drunk interns anymore.

▲CoolestBeans 2 hours ago | parent | next [-]

A month ago I was told January of this year was the inflection point. Month before that the inflection point was December of last year. I'm not saying the tech isn't getting better but are the fundamental limitations being surpassed or are the long tail failures just being pushed further away? Because if its the latter this game of "well models really got good nine months ago" won't stop.

▲throwawayffffas 41 minutes ago | parent [-]

It looks to me that this generation of models has reached a level of competence where work can be assigned to them and completed satisfactory for various degrees of satisfactory.

This reflects my own experience with these models. It's not a matter of inflection point if you ask me, it's a matter of accruing capabilities last year the output was not up to my standards 98% of the time, now it looks more like 30% of the time.

I am sure next year models will be better, but the point where the models begun being good enough to start using seriously for my use cases has now passed.

▲its_k1r4 18 minutes ago | parent | prev [-]

[flagged]