Remix.run Logo
▲ throwawayffffas 4 hours ago

I can tell you it was not correct in January of 2026. The real swift started happening with the latest models opus 4.7, fable 5, kimi k3, glm 5.2.

That's when the models started to be coherent enough for real work.

They still fuck up, but it does not feel the code was written by drunk interns anymore.

▲CoolestBeans 3 hours ago | parent | next [-]

A month ago I was told January of this year was the inflection point. Month before that the inflection point was December of last year. I'm not saying the tech isn't getting better but are the fundamental limitations being surpassed or are the long tail failures just being pushed further away? Because if its the latter this game of "well models really got good nine months ago" won't stop.

▲throwawayffffas 2 hours ago | parent [-]

It looks to me that this generation of models has reached a level of competence where work can be assigned to them and completed satisfactory for various degrees of satisfactory.

This reflects my own experience with these models. It's not a matter of inflection point if you ask me, it's a matter of accruing capabilities last year the output was not up to my standards 98% of the time, now it looks more like 30% of the time.

I am sure next year models will be better, but the point where the models begun being good enough to start using seriously for my use cases has now passed.

▲its_k1r4 an hour ago | parent | prev [-]

[flagged]