Remix.run Logo
ajspig1 a day ago

agreed. I'd also love to see more comparisons like this.

I do think that a well speced out prompt can be finished faster by an agent then if prompted lazily. I've also found measuring this to be challenging, since issues that come from lazy prompting can hide until its too late to pinpoint exactly what prompt introduced them. But maybe this should be the new "git blame"... gears are already turning for what evals could look like for this.