Remix.run Logo
Why I'm still bearish on LLMs after Navier-Stokes(dank.systems)
43 points by jaykru 6 hours ago | 10 comments
randomImmigrant 2 minutes ago | parent | next [-]

I think bearish on LLMs for automation, and bullish for LLM+human experts in specific fields, is about the right expectation for current architectures.

Apart from issues with task generalization, or perhaps related to it, is the fact that LLMs have real trouble with timekeeping, and cannot estimate the real world time it will take them to do things very well. This plus the memory issues make dreams of long horizon agents, that could plausibly handle changing specifications, quite implausible with current architectures.

In narrow domains with more deterministic outputs though, this is less of an issue, and we see multiple agents succeed much better.

The fusion of that capacity, with humans in the loop able to better direct such agents and act as their temporal tethers, is where I think the real action will be for a while at least.

ausbah 27 minutes ago | parent | prev | next [-]

> the best alternative to rigorous specification is human review. human review doesn't scale well to the volumes of output produced by language models. to make matters worse

when the business model is selling more tokens you get such per serve ice times that lead to “more” thinking, engagement baiting, fluffy narratives, and straight up dark patterns

robinpie 28 minutes ago | parent | prev | next [-]

I really appreciate seeing a tempered take that's not literally denialist about current capabilities.

an0malous 3 minutes ago | parent | next [-]

I don’t know who you’re talking about, even the most bearish people like Gary Marcus and Ed Zitron acknowledge that LLMs are useful in these same cases the OP admits. Gary Marcus is even still a long term AI advocate, he just doesn’t think LLMs are enough and we need more foundational breakthroughs. Zitron says it’s valuable technology but not worth the trillion dollar valuations the frontier labs are targeting.

The lack of temperament is very skewed towards the bulls who have been saying AGI is here, software engineering is solved, mathematics is solved, it’s going to destroy the white collar job market, and it’s going to kill us all for like 5 years now.

jaykru 18 minutes ago | parent | prev | next [-]

Thanks :) I do enjoy and use these things every day and the current capabilities are indeed amazing, just ludicrously overpriced at the frontier.

brindleth 9 minutes ago | parent | prev [-]

> current frontier models need laborious oversight and guardrails on even the simplest tasks

It is literally denialist about current capabilities

jaykru 5 minutes ago | parent [-]

why don't anthropic and openai ship yolo mode by default?

pfdietz 27 minutes ago | parent | prev | next [-]

Specifically: bearish on LLMs generally, not bearish on LLMs for pure math.

jaykru 17 minutes ago | parent [-]

yes, huge for pure math and activities that look like it.

jaykru 6 hours ago | parent | prev [-]

archive link in case i get hugged lol https://archive.ph/Z4gxF