| ▲ | dhorthy 10 hours ago | |||||||
i guess to clarify my contention: generic prompting like "review this code" or "make the architecture better" will raise the floor but cannot come close to human-quality code without humans understanding what code exists and what to ask for. Obviously cursor and other labs have a stake in this going one direction, but I don't buy it, and I don't think you should either. In fact, "Rebuild sqlite from spec" has all the problems with every other benchmark that I cited - model knows the whole problem up front and never has to iterate on the pile of slop it created cheating its way to a solution. In any case, politely, I think we're mostly arguing vibes here and I'm not sure its going to get anywhere. | ||||||||
| ▲ | monkpit 4 hours ago | parent [-] | |||||||
And what percentage of the software in the world needs to come close to human-quality code? I’d argue the percentage of the whole is VERY low. | ||||||||
| ||||||||