| ▲ | user43928 2 hours ago | ||||||||||||||||||||||
Let's get concrete about the GPT improvements you were talking about: On the MMLU benchmark, GPT-2 had an accuracy of 32.4%. GPT-3 improved this to 43.9%. GPT-4 scored 86.4%. On GPQA Diamond, it scored 31%, vs GPT-5 at 86%. GPT-6 scores 96%. If we take ARG-AGI-2, it would be 9.9% for GPT-5 vs 95% with GPT-6. The benchmarks do not corroborate the picture you were drawing about improvements slowing down between GPT major versions. | |||||||||||||||||||||||
| ▲ | abalashov 2 hours ago | parent [-] | ||||||||||||||||||||||
I don't much care about benchmarks. Talk obvious, commonsensical increases in utility to me. | |||||||||||||||||||||||
| |||||||||||||||||||||||