| ▲ | wongarsu 3 hours ago | |||||||||||||||||||||||||
Any model with a good score on this benchmark would have a good claim on superhuman abilities. Humans are pretty terrible at being thrown a long policy document and being expected to follow it And while we shouldn't anthropomorphize these models too much, I wouldn't be surprised if many of the core reasons for failures are similar. Working memory is a limited resource; you can only focus on so many things at once; reasoning depth is limited; and many real-world policies are not actually meant to be implemented in the same way they are written and have insufficient specification of edge cases With humans, we usually do the equivalent of RLHF, both via "training" with simulated cases, and via feedback while on the job. You would never hand a newbie a 124 page policy document and expect them to correctly apply it on the first task, or to do it reliably in the first month | ||||||||||||||||||||||||||
| ▲ | batshit_beaver an hour ago | parent | next [-] | |||||||||||||||||||||||||
The challenge with comparing these things to humans, is that humans learn. A newbie might not respect your organization’s set of policies on day one, but what about 3 months in? Or 3 years? Meanwhile there’s still no reasonable mechanism for automatically fine tuning LLMs or adjusting their harnesses to make them better at completing your organization’s objectives more successfully. They’re still overwhelmingly governed by the shared weights and harness policies found to be successful for the average case. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | loremium 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||
isn't it because there are too many contradictions and ambiguity? the reason it works for humans is because we don't apply everything at once either. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | ActionHank 3 hours ago | parent | prev [-] | |||||||||||||||||||||||||
That's a great comparison, human vs ai on a wall of text. The problem is that it doesn't fit the sales pitch of LLMs and agents - humanlike or better, repeatably, 24/7, for a fraction of the price, you just need to make sure that you give it all the rules. Unfortunately we can't really have a meaningful conversation until the money vampires have left so we will need to reschedule this until after the bubble. | ||||||||||||||||||||||||||