Remix.run Logo
▲ user43928 2 hours ago

Let's get concrete about the GPT improvements you were talking about:

On the MMLU benchmark, GPT-2 had an accuracy of 32.4%. GPT-3 improved this to 43.9%.

GPT-4 scored 86.4%. On GPQA Diamond, it scored 31%, vs GPT-5 at 86%. GPT-6 scores 96%.

If we take ARG-AGI-2, it would be 9.9% for GPT-5 vs 95% with GPT-6.

The benchmarks do not corroborate the picture you were drawing about improvements slowing down between GPT major versions.

▲abalashov 2 hours ago | parent [-]

I don't much care about benchmarks. Talk obvious, commonsensical increases in utility to me.

▲user43928 2 hours ago | parent | next [-]

There wasn't that much utility in old, unreliable AI.

I remember their capabilities like so:

GPT-3 could produce convincing looking texts, sometimes.

GPT-4 was somewhat smarter and would give more accurate answers. At that time image understanding was released, wasn't it?

Then with GPT-5 we have a reasoning model, another jump in capabilities.

Now, compare GPT-6 Astra with its ability to implement software, work on long horizon tasks, and visual understanding.

▲ragequittah an hour ago | parent | prev [-]

If you haven't seen the (massive) increases in utility I'd wager you haven't been using the tools much.

▲abalashov 20 minutes ago | parent [-]

So I'm holding it wrong?