Remix.run Logo
▲ badRNG an hour ago

Do you think that even today's frontier models are unable to handle this task?

▲Aurornis 18 minutes ago | parent | next [-]

Today's frontier models can't even handle 100% of the SWE benchmarks that have been around for longer than they've been training the models. Companies that are benchmaxxing their models against the benchmarks haven't even been able to get them to 100%.

I've been using Astra and Opus 5.5 all week and I still have to intervene and tell them to do something different all the time.

I think the people whose SWE jobs consisted of simple features, bug fixes, or tweaking .toml until the server works are on their way out.

Someone needs to steer the LLMs around for now, though. There is a very long tail of non-trivial work and expertise that can't yet be replaced by a CEO telling the LLM to make the product work. There are many CEOs trying to do that right now and, outside of very simple CRUD apps, it doesn't work yet.

▲sroussey an hour ago | parent | prev [-]

Today’s frontier models create the problem.

▲falcor84 an hour ago | parent | next [-]

That isn't actually relevant. They create problems, but when confronted with them and asked to fix them, my experience is that they can fix them very effectively. I haven't had an AI session get "stuck" on me, not being able to solve a problem in over a year.

The challenge as I see it is in shifting checks further left, but actual mitigation is just not really an issue any more.

▲WithinReason an hour ago | parent | prev | next [-]

What problem? They keep finding security bugs that humans couldn't for decades.

▲knowaveragejoe 29 minutes ago | parent | prev [-]

So ask them to fix the problems in a targeted way.