Remix.run Logo
polotics 4 hours ago

Yep and I just ran a simple enterprisey "ambiguity" bench on the big three (US) model providers: same ambiguous initial-prompt with same clarifications and pushback prompt sequence afterwards.

The edge of correct/better when facing ambiguity is very fuzzy, all models from the past 6 month or so have similar random ways of spinning between too-literal avenues and oddly misplaced misled fixations. Taking the right initiatives in face of uncertainty is definitely AGI, and its not there, and perceptrons + attention layers just ain't got what it takes no matter how hard you push.

organsnyder 3 hours ago | parent [-]

The best I've come up with is to ask an agent to come up with questions to help get to a focused design. But it still needs me in the loop for that.