Remix.run Logo
smallmancontrov 3 hours ago

These are ~1000-word stories (prompts at [1]), squarely in the sweet spot where planning logistics are minimal, context windows comfortably fit everything, and the appropriate level of abstraction omits any details which could glaringly reveal weak world-model knowledge.

SOTA LLMs from 2 years ago would saturate this study, and while modern LLMs are amazing and I suspect they would do quite well on a similar but more challenging study that forces them out of said comfort zone, this isn't that.

[1] https://www.cambridge.org/core/journals/judgment-and-decisio...