| ▲ | Aurornis an hour ago | |
Today's frontier models can't even handle 100% of the SWE benchmarks that have been around for longer than they've been training the models. Companies that are benchmaxxing their models against the benchmarks haven't even been able to get them to 100%. I've been using Astra and Opus 5.5 all week and I still have to intervene and tell them to do something different all the time. I think the people whose SWE jobs consisted of simple features, bug fixes, or tweaking .toml until the server works are on their way out. Someone needs to steer the LLMs around for now, though. There is a very long tail of non-trivial work and expertise that can't yet be replaced by a CEO telling the LLM to make the product work. There are many CEOs trying to do that right now and, outside of very simple CRUD apps, it doesn't work yet. | ||