Remix.run Logo
kzrdude 3 hours ago

There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback.

That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user.

Regardless what you call it, it's a divergence between what the power user wants and what the model developers want, I think.