Remix.run Logo
▲ Olscore an hour ago

About 25-30% of the things shipped by agents need fixes.

▲eru an hour ago | parent [-]

How does that compare to human shipping?

▲necovek 30 minutes ago | parent [-]

I'd instead say that for both it's really at 100%, and not any number in between.

Problem is that we can't define precisely what is good enough or when software is "finished": I mean, we are trying to do that with human languages, so it should not be a surprise.

Yes, LLMs are now similarly aware of the average context a human would be aware of, but for anything specific to the situation a human has better chances of resolving the ambiguity.