Remix.run Logo
▲ abalashov 2 hours ago

There seems to be two common fallacies in SV technocrats' AI discourse:

(1) A far-reaching tendency to overextrapolate from the low-hanging fruit of the last few years of pretraining progress. GPT-2 to GPT-3 may have been a quantum leap, but GPT-3 to GPT-4 was not, and GPT-4 to 5 even less so.

The party has been kept going by RL and agents, but still, there is indeed a point of diminishing returns, not just relative to available compute but to how much training is possible when the entire intellectual output of humanity, plus a raft of synthetic data, has already been inhaled by the training process.

If one is to internalise the things that are said here on HN with regularity about model progress, and sentences ending with "yet" or "for now", then it would be easy to conclude that my 10 year-old son, who gained 3 inches of height last year, will be tallest structure on the planet by age 17.

(2) Inability to distinguish between technological, computational, and energetic limits of LLM capabilities vs. ontological / conceptual ones. There are some things LLMs cannot do, or at least do well, at any size, at infinite size and with infinite compute, simply due to the very nature of what LLMs are to begin with.

This latter topic receives almost no attention, except maybe from Gary Marcus and Yann LeCun. In that respect, this article is a breath of fresh air, insofar as it highlights that LLMs aren't "AI" at all, as we have traditionally understood the concept.

They really _are_ stochastic parrots. The relevant questions are about how much that matters for some domain or set of applications, not whether they are an emerging alien intelligence with civilisation-threatening capabilities.

▲jmoggr an hour ago | parent | next [-]

> There are some things LLMs cannot do, or at least do well, at any size, at infinite size and with infinite compute, simply due to the very nature of what LLMs are to begin with.

How can we be so confident that there is anything LLMs fundamentally cannot do? Proving that impossibility seems hard.

Drawing trend lines far into the future is foolish, but the recent trend is clear and its not obvious how much more progress is needed before they start having a meaningful impact on more aspects of life.

> not whether they are an emerging alien intelligence with civilisation-threatening capabilities.

The people building them are explicitly attempting to do this. They may not succeed, but what if they do? Seems worth considering that scenario.

▲abalashov an hour ago | parent | next [-]

> How can we be so confident that there is anything LLMs fundamentally cannot do? Proving that impossibility seems hard.

This is where applied business programmers part ways with philosophers of mind and cognitive scientists. However, the lack of an inner model of the world is a formidable limitation, while the reliability of purely statistical-inferential processes will never be adequate for some basic building blocks of modernity.

▲Sevii an hour ago | parent | prev [-]

LLM utility is exponential with increased LLM capability. Even if improvements to LLMs slow down their impact will continue to increase. And there is no evidence LLMs are slowing down. In fact improvements still seem to be accelerating.

▲abalashov an hour ago | parent [-]

You like these evocative words that imbue one with a feeling of "vroom" and "whoosh", like "exponential" and "accelerating", don't you?

But arithmetically speaking, is any of that true? Is model progress truly "accelerating"? How can you compare the delta from GPT-2 to GPT-3 with the delta from GPT-5 to GPT-6 with no sense of irony? And, at the risk of being quite blunt, do you know the meaning of the term "exponential"?

▲user43928 an hour ago | parent [-]

Let's get concrete about the GPT improvements you were talking about:

On the MMLU benchmark, GPT-2 had an accuracy of 32.4%. GPT-3 improved this to 43.9%.

GPT-4 scored 86.4%. On GPQA Diamond, it scored 31%, vs GPT-5 at 86%. GPT-6 scores 96%.

If we take ARG-AGI-2, it would be 9.9% for GPT-5 vs 95% with GPT-6.

The benchmarks do not corroborate the picture you were drawing about improvements slowing down between GPT major versions.

▲abalashov an hour ago | parent [-]

I don't much care about benchmarks. Talk obvious, commonsensical increases in utility to me.

▲user43928 37 minutes ago | parent [-]

There wasn't that much utility in old, unreliable AI.

I remember their capabilities like so:

GPT-3 could produce convincing looking texts, sometimes.

GPT-4 was somewhat smarter and would give more accurate answers. At that time image understanding was released, wasn't it?

Then with GPT-5 we have a reasoning model, another jump in capabilities.

Now, compare GPT-6 Astra with its ability to implement software, work on long horizon tasks, and visual understanding.

▲card_zero an hour ago | parent | prev | next [-]

Necessary pedantry: if he's 4 foot 6 inches now, 3 inches is 5.555 (recurring) percent of that. If he grows at 105.555% per year then by 17 he is a fairly plausible 6 foot 6. (I think he then passes the forty foot mark soon after age 50.)

▲abalashov an hour ago | parent [-]

Ah, but he only grew about half an inch from 9 to 10, so the "rate" of "runaway progress" is "accelerating" "exponentially" (wait, how many derivatives is that?), and he will soon EsCapE CoNTaiNmEnT and hack HuggingFace...

▲card_zero an hour ago | parent [-]

Ha, of course, add derivatives until desired prediction is met.

▲ an hour ago | parent | prev [-]
[deleted]