| ▲ | tcp_handshaker 5 hours ago | |||||||||||||||||||||||||
You can easily test this yourself with the SOTA models....or read the corroborating literature... "General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks" https://arxiv.org/abs/2604.11778 "...General365, a benchmark specifically designed to assess general reasoning in LLMs. By restricting background knowledge to a K-12 level, General365 explicitly decouples reasoning from specialized expertise. The benchmark comprises 365 seed problems and 1,095 variant problems across eight categories, ensuring both high difficulty and diversity. Evaluations across 26 leading LLMs reveal that even the top-performing model achieves only 62.8% accuracy, in stark contrast to the near-perfect performances of LLMs in math and physics benchmarks..." | ||||||||||||||||||||||||||
| ▲ | xnorswap 5 hours ago | parent | next [-] | |||||||||||||||||||||||||
What proportion of the human population could answer the example from that paper?
What does "The person in green shoes does not wear multicolored clothing" even mean?Nowhere is "multicoloured" defined, are we to assume it should be treated as a colour and implied that someone else must be wearing "multicoloured clothing"? Because strictly that doesn't logically follow, and it ought to be phrased as "The person in green shoes is not the person wearing multicoloured clothing" if that is the case. This is an extremely hard logic puzzle, especially since it's revealed at the end that there are multiple solutions. I'd expect anyone to struggle unless armed with prolog. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | fn-mote 4 hours ago | parent | prev [-] | |||||||||||||||||||||||||
Super interesting, thanks for the reference. I particularly found the note about “local collapse” helpful (near the end of section 3). The idea is that even though benchmarks contain a wide variety of different reasoning tasks, each individual problem requires only a few skills - unlike this benchmark where they deliberately construct tasks that span many categories. | ||||||||||||||||||||||||||