Remix.run Logo
maxnevermind a day ago

> Someday, such as in 2040, there may be available, for every human being, the performance equivalent of 'a B300 GPU for contemporary LLMs'. What would this world be like?

If we talk about just LLMs, given how things have been going since ChatGPT, my bet it would not change that much. LLMs are not foundational technology such as Internet or Steam engine or Rail roads were. There are very few products that can build upon them because of reliability issues which are completely unresolvable for LLMs, chatbots is a decent product that came out of it, coding harnesses is another one, this is not even close to the impact Internet or Steam engine had. LLMs gave us nice productivity tools for highly motivated expert knowledge workers, that is all. LLMs are getting better and will get better, but it is impossible to describe the universe and compress it into few terabytes and that is what they are doing atm effectively, so all serious LLMs's issues will still be there in 2040: the lack on continues learning, hallucinations, terrible sample ratio, agent's failures on long horizon tasks, instruction following failures.

setopt a day ago | parent | next [-]

I disagree with this take. While LLMs themselves are currently unreliable, the work done in the math community on hooking up creative LLMs to reliable verifiers like Lean show that it’s possible to construct systems where the unreliability is suppressed. For now, that still requires experts to set up and monitor, but I do believe that in a couple of decades we’ll make progress on how to do more mundane tasks in a reliable way without expert supervision, where an LLM still sits as the translation layer between humans and machines. And that universal human-to-machine translator, I would certainly consider a foundational technology.

EDIT: If you asked people on the street in 1990, they’d probably not consider the Internet to be in the same category as the Steam engine either. I mean, you already had phone and fax, so it wasn’t that ground breaking. And I’ve even read articles from the mid-90s declaring the Internet a temporary fad.

maxnevermind a day ago | parent | next [-]

> I disagree with this take. While LLMs themselves are currently unreliable, the work done in the math community on hooking up creative LLMs to reliable verifiers like Lean show that it’s possible to construct systems where the unreliability is suppressed. For now, that still requires experts to set up and monitor, but I do believe that in a couple of decades we’ll make progress on how to do more mundane tasks in a reliable way without expert supervision, where an LLM still sits as the translation layer between humans and machines. And that universal human-to-machine translator, I would certainly consider a foundational technology.

Yes, for a verifiable domains you can set up a harness and brute-force a search space if you have enough money for compute. Why do you think they keep coming up with those examples of impressive achievements like solving math puzzles? Why not focus on something with economic value to it? My answer is they can't, those are hard problems, those require building an actual product, those require reliability.

> EDIT: If you asked people on the street in 1990, they’d probably not consider the Internet to be in the same category as the Steam engine either. I mean, you already had phone and fax, so it wasn’t that ground breaking. And I’ve even read articles from the mid-90s declaring the Internet a temporary fad.

We are not people people on the street we are people who are directly involved in application of the technology, we posses a higher level of insight.

xerlait a day ago | parent [-]

I generally side with your skepticism, but there are verifiable domains with economic value such as drug discovery.

flerovium114 a day ago | parent [-]

Are LLMs doing drug discovery? To my knowledge, that’s all classic “machine learning”

setopt a day ago | parent [-]

Not yet, as far as I know, but perhaps someone will find a way to do that in the future. We're talking about the next two decades here, a lot can happen in that time.

DNA itself can be thought of as a language which describes proteins, and I personally don't know enough about the limits of LLMs as a technology to claim that it can never be adapted to, say, do reverse translation from desired protein shapes into DNA sequences.

pixl97 21 hours ago | parent [-]

This is why it's better to call the underlying technology transformers rather than 'language model'. It can make a model from anything you can digitize and is informational (random noise would not be informational for example). It's the relations between the bits of data that matters.

wjnc a day ago | parent | prev | next [-]

I share this sentiment, while deploring the current AI board room sentiment. I know some people that design (safe) buildings. They use software all the time for their load-bearing work (lol). Think about the creativity that could be unleashed if that (like LEAN for math) becomes a commodity.

The same thing for my job: creating insurance premiums is somewhat hard but not stellar. A combination of skills, data, tools and people. I can imagine a future where you can post a 'have good weather on holiday or money back' bond on a platform. (I can think of more serious applications...) The sheer diversity and amount of liquidity AI's can create is enormous. (Switching to a very general outlook here.) And with liquidity hopefully comes more specificity in the ROI on saving our planet. (Or the disproving of the necessity thereof, if that is your outlook.)

knollimar a day ago | parent [-]

Construction isn't using LLMs any time soon like that. We don't even send our files in structured data because everyone's playing nose goes for liability and sending pdfs

wjnc a day ago | parent [-]

I know! And financials and actuaries aren't either. Excel, Python, R and don't know what we can use ducktape and spit for to keep together. But as the LEAN-example quite nicely shows: with the right scaffolding I am convinced the GPU World /could/ look rosy.

An economic problem would be - how could the builder of a scaffold capture _some_ value without right away be copied. The fact that LLMs threw away every protection of intellectual property at their inception makes it a lot harder to invest in (with the goal of capturing some) value. I am at least /somewhat/ influenced by RMS on the 'seductive mirage' of IP but I have my doubts on how copyleft could bring more than breadcrumbs to those who build scaffolds.

knollimar a day ago | parent [-]

The LEAN example is the ideal case. It's like a baby POC to me.

I can only assume the only realistic way to get all the federated data will be a ton of massivr vertical consolidation of industry by hubris laden tech execs crossing industry domains.

wjnc a day ago | parent [-]

In spirit of this conversation - Terence Tao is awesome right? He was already one of the best mathematicians on the planet when he happened to start dabbling in computer assisted proofs (obviously as a giant standing on the shoulders of giants) at exactly the right time for the computer revolution to find a perfect use case (as you point out).

alexpotato an hour ago | parent | prev | next [-]

> hooking up creative LLMs to reliable verifiers like Lean show that it’s possible to construct systems where the unreliability is suppressed.

You can also have LLMs create these things called "programs" that are written in "code" that result in deterministic outputs when given inputs.

I say this with a bit of snark to highlight the point that both humans and LLMs can write code that is cheaper to run and easy to verify. I'm not saying Lean being used like this is a bad idea, just that it's just one end of the spectrum.

_superposition_ a day ago | parent | prev | next [-]

Most notably Paul Krugman Nobel winning economist in 98: "By 2005 or so, it will become clear that the Internet’s impact on the economy has been no greater than the fax machine’s."

Predicting the future is hard.

voncheese a day ago | parent | next [-]

Hadn't seen this quote before, amazing.

Also goes to show how hard it is to predict anything that is a massive change - these levels of changes are so infrequent that we can't rely on priors to predict the future.

megagpt1 21 hours ago | parent | prev [-]

What will be LLMs' social media?

setopt 9 hours ago | parent [-]

More or less the same as social media but for agents, perhaps?

When I want to book an airline ticket in the future, I expect my LLM agent to contact the airline LLM agent, and the two to discuss options and negotiate a deal on my behalf.

_superposition_ 2 hours ago | parent [-]

Negotiate? For an airline ticket? When's the last time a human did that and why would an agent? There's a price, and if you don't like it you go to a competitor you don't call customer service and negotiate? Who in their right mind would let an LLM set prices?

dclowd9901 a day ago | parent | prev [-]

I don't think you've said anything different than the person you're responding to. The difference primarily seems to be whether you think AI is transformative vs simply being another tool (which is really just a matter of perspective).

If you're not an expert software engineer, it's transformative. If you are, it's just another tool.

stevepotter 19 hours ago | parent [-]

I’m an expert software engineer. I would categorize something that multiplies my productivity by > 2x is transformative. AI has certainly done that and more for me

roel_v a day ago | parent | prev | next [-]

This take is unfathomable to me. How are LLMs not massively revolutionary, despite them not being infallible? This comment reminds me of there being competitions for finding useful purposes for electricity when it was first discovered. Imagine talking about the usefulness of LLMs while thinking that LLMs in 2026 just compress a few TB of data into a set of weights and that's all they work with lol.

entropy47 a day ago | parent | next [-]

Whichever side of the divide you're on, I think it's really weird that on a site like HN (where people have at least a passing familiarity with technology) there are such polarised takes on whether it's the next industrial revolution, or a technological parlour trick.

I'm giving away which side I'm on now, but most of the TAM have experienced the product at this point, and trillions of dollars have been invested. I would wager most other revolutionary technologies found their game changing, broad reaching applications by this point.

david-gpu a day ago | parent | next [-]

Did you experience the popularization of the WWW? There were massive disagreements even among smart tech-savvy people. We even had a stock market boom and just cycle associated with it.

Funnily, none of that mattered. What mattered is that enough ordinary people on the street found it useful, both for personal and business reasons. Ask yourself: are regular people using this tech, either as producers or consumers? And remember that this is the least powerful that it will ever be.

entropy47 a day ago | parent | next [-]

It's very hard to disagree with the point that it will get more powerful, perhaps in uenxpected ways.

I don't know that the average regular person is using this tech as a producer, and while it's hard not to consume it - I don't know if anybody consuming it really wants to. I'm in a bubble of people where we give it a lot more mindshare, but e.g my electrician is aware of the tech, knows what it does, and doesn't really care. I think you're right that if it could manage his books, help him rewire a switchboard, order inventory for him etc that would be very different.

ipsod a day ago | parent | next [-]

> I don't know if anybody consuming it really wants to.

You don't know of people who ask LLMs questions? Or just chat with it for fun and to learn?

Maybe your electrician isn't interested. A pharmacist and a nurse I talk to are. Virtually everybody I've talked to uses it to some degree or another. I think it'll soon replace Alexa/Siri for most people, as well as google search.

That's pretty far from nobody.

HarlequinHair a day ago | parent [-]

I think the first parent post all this conversation originated from, pointed that LLMs are not groundbreaking inventions.

If people use it to replace voice assistants or search engines, you can see how this new invention is not really adding anything you already hadn't.

AIs in general are very useful, but apart from the campaign they did to change everything's name into AI (also well known algorithms and "old" machine learning), if you look at all the possible applications, LLMs are possibly the least useful of them.

Especially if you take into account how much money they spent, and how much effort they put to shove it in your throat.

If a technology is convenient for people, you shouldn't need to insist to make people use it. If it is convenient, people will use it in a natural way.

ipsod a day ago | parent | next [-]

> If a technology is convenient for people, you shouldn't need to insist to make people use it.

This is like the railroads, and the tracks have been laid. Because of the enormous cost, it can't be left to "hope they'll find it useful," and has to be pushed like crazy, or we're going to be left with a train system with no goods or passengers, and no path into the future.

Personally, having AI as an intelligence I can call like it's a library from apps I build... and to help build those apps... I feel that the future is as bright as the sun. I've always believed in software, but I feel like a chef that's just discovered salt and butter.

I don't know what it means for the world stage, but I believe that software is the highest art, and it's just had an incalculably massive buff. As an artist, I am stoked. The possibilities for creation of good things are limitless, and if you can't see that, I can't help you.

However, I wouldn't necessarily invest in AI companies. I don't know how this all fits into the rest of the world, except that I believe in my own ability to make products, same as I ever did, but now with AI.

HarlequinHair 21 hours ago | parent [-]

> This is like the railroads, and the tracks have been laid.

Except that for railroads I don't need people to convince me about how/when/why it can be useful.

Railroads came out to solve an issue we had (and then it evolved). LLMs are not giving us any new solution, just different speed (better) and quality (worse) to solve things we co&ld already solve.

It's basically the same concept as fast fashion. Worse quality, more speed, possibly more economic. You can change your wardrobe possibly every year, but you will never wear good quality clothes.

Now you can write code at impressive speed, but it is worse in quality, and you need to review it more often than "normal" code.

AI code should always be fully reviewed and put into the correct context to make it useful, otherwise long term problems will void short term benefits

ipsod 19 hours ago | parent [-]

Nobody convinced me AI was useful. At this point, nobody could convince me it wasn't - it's like going from whittling stuff out of wood to having a kitted-out fully-manned machine shop.

It's clear from your words you lack experience and skill in this area - you've only got opinions.

HarlequinHair 18 hours ago | parent [-]

My "only" experience happen to be coding 5 days a week with agentic AI, plus reviews.

Only because I do not share your point of view it doesn't mean I lack experience or skills.

You are pretty fast at jumping at conclusions.

ipsod 13 hours ago | parent | next [-]

I guess we might actually have similar feelings, except, you say "only speed", and "only economics". Only time and only money. As if those aren't two of the most valuable currencies on earth.

ipsod 17 hours ago | parent | prev [-]

> You are pretty fast at jumping at conclusions.

Yeah, because I can see some of my own thoughts and experience from a long time ago in your words, and my experience tells me loud and clear that they're outdated and wrong.

david-gpu 21 hours ago | parent | prev [-]

If people use [the Internet] to replace [Sears catalogs] or [paper mail], you can see how this new invention is not really adding anything you already hadn't.

HarlequinHair 20 hours ago | parent [-]

It would have been a good counter-example, if only internet wasn't: - adding encryption - removing processing time - removing geography - removing logistic issues ... and much more at a convenient price.

They basically invented teleportation for sear catalogs and paper mail, can you see how this is different? Also, Internet was built to solve a problem we had, and it solved it beautifully and at reasonable prices.

I for sure can understand it introduced different issues, buy this is another topic.

david-gpu 19 hours ago | parent [-]

You can ask LLMs questions about an immense variety of subjects, and they can give you customized explanations to those questions. It is as if an acquaintance of yours had digested Wikipedia and tens of thousands of books and had infinite patience to help you understand whatever you are curious about. And while the answers are not going to be picture perfect every time, just like Wikipedia or a well-educated acquaintance, they will be generally very useful.

I understand exactly what the Internet of the 90s did and did not do, because I was there to witness it myself. I had these same types of conversations with naysayers back then. Guess who was right?

HarlequinHair 18 hours ago | parent [-]

Maybe you have all the means to avoid sycophancy bias and other flaws LLMs have by design, this doesn't mean that everyone can use it the same way.

The fact the you were right in a previous topic, doesn't grant you objective truth. You have your strong opinions on AI, and that's ok. Fortunately not everyone has to share the same idea.

If you think LLMs can do better than you, feel free to use them in production and bring them to work with you, it's your choice. Do not force others do the same, tho.

calgoo a day ago | parent | prev [-]

From what i hear in a small subset of local small businesses is that they are using it for backoffice stuff, like taxes, finding legal codes, organization, planning etc.

maxnevermind a day ago | parent | prev | next [-]

> Funnily, none of that mattered. What mattered is that enough ordinary people on the street found it useful, both for personal and business reasons.

Broad economic impact due to Internet boom in the the US came from 2 main categories: 1) New huge internet enabled tech companies appeared(Amazon, Google, Meta etc.), all created their own products, all were build on reliable foundational technology. LLMs is not a reliable foundational technology. 2) Existing legacy companies could utilize internet to connect their teams/departments/offices though a bunch of new software which were build on reliable foundational technology. LLMs is not a reliable foundational technology.

> And remember that this is the least powerful that it will ever be.

This mantra is being repeated but "serious LLMs's issues" I described are still there and will be there because they are part of how LLMs are built and work. Without addressing those you can't build products but LLMs's capability as a personal productivity tool will keep rising, yes.

The "best" possible outcome of LLMs to a broader economy might be that a smaller group of experts will be able to do the job in companies due to boost of their personal productivity and a bunch of people will be freed up(laid off) and they will have to go and work on something else and thus a boost of productivity in the economy. But it seems it won't be any low hanging fruits(problems to solve) this time as it was during the internet era. It is not like we out of problems: new cancer treatments, self driving cars, nuclear fusion, modular nuclear reactors etc. There is a huge value to capture there, here is the thing though, those are hard problems, it is not a new TikTok, gmail, netflix etc. Yes, people will have LLMs now but can we actually start solving hard problems with them? Because it might be the case that a gravy train of the last few decades for Silicon Valley is over, no more useless internet enabled services, no more billion dollar companies built on just applying internet to yet another thing and producing another digital product. People's free time is limited, its redistribution across digital services will not grow the economy, global internet penetration is already pretty high and won't grow that much, so there might not be much value to capture there.

aa-jv a day ago | parent | prev | next [-]

>Did you experience the popularization of the WWW? There were massive disagreements even among smart tech-savvy people. We even had a stock market boom and just cycle associated with it.

In the early phase of the Internet envelope, we all wanted to build the real Internet and destroy the AOL's and Compuserves' .. but now the Internet has turned into a mass of AOL's and Compuserves.

I think the same thing is happening with the silo'ization of LLM's. People don't want to have to use the "AOL of LLM's" for things - they want to build out their own infrastructure and have a pool of local LLM's (where once it was the company's own modem pool, etc.)

It is, oddly enough, parallel with the original arguments in the 70's and 80's regarding supercomputers versus personal computers, or the cloud versus locally managed resources, etc. We are still fighting this battle, just with higher bandwidth.

Do we operate a terminal with a fixed window into a bigger computer, or do we operate a heavier local computer with a proper interface?

kakacik a day ago | parent | prev [-]

Most people I know, blue and white collar up to and including surgeons and IT folks, use it as better search.

My wife is a GP, and all she would ever need is a reliable speech to voice recognition of medical jargon, in french. Not 98%, not almost there, just there.

Most people don't live in IT echo chambers like HN. For them, chatgpts are just another online gimmick that saves a bit of time vs googling things, or puts together some paperwork quicker, and thats it.

For me, I can see potential beyond but are largely unmoved by it. The price to pay, in fucking up our environment, losing skills and/or losing skills to build skills, making large swaths of population unemployed with no replacement in sight, seems much more like devolution and net loss. Extremely few winners and big bunch of losers. Very happy to be wrong here of course, but only time will tell and neither you nor me know for sure.

roel_v a day ago | parent [-]

How can it be at the same time "a better search" and also "making large swaths of population unemployed with no replacement in sight" ?

stevepotter a day ago | parent [-]

Because it’s so powerful and versatile. The other day my wife took a picture of our living room and chathgpt picked out the perfect rug. That’s both a better search and avoided, say, an interior decorator. In that scenario it’s kinda both.

roel_v a day ago | parent [-]

But only a small part of AI applications are that search-like. I mean we could even argue if this is still only search, but that's semantics and boring. But even if we would, there are a lot more applications than pure search. So the GP's original claim that it's just an improved form of something that doesn't put people out of a job, and also puts large amounts of people out of a job, doesn't make sense. Framing tax advise as "yeah it's just a better way to search tax codes" is in the same boat. It's such a massive improvement of both searching and synthesizing that it's a bit weird to call it that. Like saying a 2026 car is just an improvement of a Model T.

stevepotter 19 hours ago | parent [-]

I think we agree. I think we’ve barely touched the surface on its applications. It’s digitized reasoning. That’s hard to wrap one’s head around. My hope is that the vertical applications that gain traction aren’t from a handful of companies.

cemkum a day ago | parent | prev | next [-]

The TAM of electricity is 99.999% of humanity, yet there was 82 years between Alessandro Volta's Voltaic Pile experiments and the first electrical grid in the world, Pearl Street Station, New York City.

A farmer may have seen an LLM answering their questions about how to make a less dry chicken sandwich, but he hasn't seen a specifically designed and tested harness that autonomously manages a crop harvester. There is not a lot of work in this space right now, because programming and math are infinitely more verifiable, and the capabilities of never models would probably invalidate any usecases, but if the model progress slows down, all attention will go towards making them work for any aspect of production.

entropy47 a day ago | parent | next [-]

Fair points (and I appreciate very much that they read like a human wrote them). As a counter point - and I know it's a weak one - a few years ago people were telling me that crypto was going to replace currency, and soon every person on the planet would be using it for every transaction. Many people believe that will still happen, but I don't think many serious people do.

I want to believe there is big bucks in this technology - my employer is tied up in it and so is my compensation. I just find it personally very unremarkable in a way that I never did with phones or the internet. When I got my first phone, or first started using Dogpile for research - I saw immediate value, and would have paid big bucks to keep it in my life. AI has had a few years and a huge amount of investment (time and money) and if it disappeared overnight I wouldn't pay $10 to get it back. I keep trying it, there's just something missing in my brain where it should be clicking for me :(

munksbeer a day ago | parent [-]

> a few years ago people were telling me that crypto was going to replace currency, and soon every person on the planet would be using it for every transaction. Many people believe that will still happen, but I don't think many serious people do.

As you say, that is a weak counter point. I do get what you're saying, but crypto flaws were built in from the start:

- Deflationary currency isn't going to be spent, it is going to be hoarded, if it would have any value at all. The design was self defeating from the very start.

- Technical limits of PoW. Obvious right from the start, which led to the block wars. PoS "solves" that but with other costs. And it still doesn't solve the first point.

Despite these, crypto did get adopted, just not in any way that crypto maximalists said it would. It got hoarded, as expected.

But the real test is: If crypto disappeared tomorrow, would the world notice or really care? I would argue, apart from people losing money, most of the world wouldn't even notice it was gone.

There is no way you could say that for LLMs.

TimMurnaghan a day ago | parent | prev | next [-]

Nitpick or two. Holborn in London was a little earlier. Also neither was really a grid. That took another ~45 years. Maybe that's better for your argument - but I'm not quite sure that the analogy fully holds - as it's a different class of infrastructure problem.

pixl97 21 hours ago | parent | prev [-]

>but he hasn't seen a specifically designed and tested harness that autonomously manages a crop harvester.

Eh, I'd say there is a lot more stuff like this being developed these days. Autonomous seeders are becoming more common and (semi-)autonomous tractors are in use.

I'd say a very large part of the in field work can be done now, and is using enhancements like vision models to identify potential problems that crop up to notify humans. It's when the machines have to interact in tight busy spaces or with non-farm traffic that control is handed over to a human.

bsenftner a day ago | parent | prev | next [-]

That polarizing take is magnified by our failure to teach how to manage discussions that contain controversy and disagreement. Then there is the fact that most people's idea of debate is some game of dominance and not to find solutions.

titzer a day ago | parent | prev | next [-]

We're still in early stages, but in just 2 years AI has presented higher education with the single most disruptive change in decades. Now a computer can do your entire coursework for you, is available to converse on a broad range of topics, and can accomplish unaided huge open-ended problems.

> most other revolutionary technologies found their game changing, broad reaching applications by this point.

I think your timeline expectations are miscalibrated. The steam engine goes back centuries. Even in the mid 1700s to 1800s, advances in engine design were decades apart, and the logistics of building and deploying steam engines mandated an uptake in years, even in perfectly-fitted applications.

Like these applications where there were existing manual-labor solutions that not only fought the technology but were still competitive, AI and LLMs weren't born in a vacuum. In order to revolutionize, it must displace, and people are rightly finding ways to do their jobs in more productive ways that moderate adoption speed.

Can you name some other technologies that have exploded from 0 to billions of users in just 3 years?

_superposition_ a day ago | parent [-]

Let's not forget the electric car... Still yet to be realized at scale but looking more likely.

madaxe_again a day ago | parent | prev [-]

It took the better part of a century for people to go from “hey look a steam engine” to “why don’t we put it on wheels?”.

Electricity took about half that time, 50 years, to go from “hey look a motor” to “why don’t we put them everywhere and build power grids?”

TV took 25 years from patent to mass market adoption.

The internet, 15 years from lynx to iPhone.

Google published the transformer architecture in 2017. We are currently at “oh hey a talking computer”.

pixl97 21 hours ago | parent [-]

Oh, based on the METR report we're beyond a talking computer. Like most technology it's just not spread evenly.

madaxe_again 18 hours ago | parent [-]

We absolutely are - but we’re still at the “weirdo factory owner installs experimental electric motor” stage. You look at what’s going on in robotics, look at GROOT/Isaac, and the circle closes quite neatly. World’s going to change in ways we can barely even imagine, no more than someone in 1720 could imagine the buffet car on an intercity express train.

Roark66 a day ago | parent | prev | next [-]

I'm with you on this. I remember when the Internet became a thing, how revolutionary it was to all areas of my life.

This is comparable.

While I dislike the bonkers valuations, and "were building Agi so it tells us how to be profitable" is onion worthy statement LLMs are an absolutely revolutionary technology.

Even with all the hallucinations.

Imagine telling someone 15 years ago Internet/Google is useless because people sometimes do not tell the truth online... It's like that.

I also hate the fact AI is being used as justification to grab all compute in the world and lock it up in one country's datacenters.

I hope this endeavour fails, but sadly being realistic I have to admit I think the so called "AI bubble" will not pop. Instead the companies will be bailed out with printed money.

I recently looked up there are approximately $20Trillion in circulation. Even if they print extra $2T it will dilute existing money supply by 10%. So everyone that uses USD as a currency will pay for the datacenter build out wether we want it or not.

All the compute will be slurped from the market for this and companies like nvidia that get used to 50bln deals will never go back to making consumer stuff.

I never thought I'll live to see the PC revolution reverse, but that is what seems to be happening.

Gareth321 a day ago | parent | prev [-]

I strongly agree. I cannot understand how anyone could seriously use this technology and come away thinking it's mostly useless. I suspect there is a LOT of bias baked into their opinion.

onion2k a day ago | parent | prev | next [-]

agent's failures on long horizon tasks

We've moved from LLMs being able to work on a task for about 2 minutes to about 2 hours in the last 18 months, and that's mostly limited by the context window size filling up. In 14 years time I don't really see a reason why that wouldn't have extended a time frame that's effectively continuous forever, or at least a ceiling that's indistinguishable from that.

The question really becomes "why would we want that?". The main reason you'd want an AI that can focus on a task forever is to completely remove the human from the loop. That's something we should be cautious about.

FinnKuhn a day ago | parent | next [-]

  I do not think agents need to work continuously. They "just"
  need to work on a project longer than an employee is able to
  work on it for them to be commercially useful. Although this
  does not consider that agents might take a shorter or longer
  amount of time for the same task. Therefore, we should begin
  to compare these in tasks completed during that time instead.
onion2k a day ago | parent [-]

Not sure about that. The nature of work changes if you have some[one|thing] that can work on it continuously forever. The goal of using AI shouldn't be to do the things a person does; it should be to do the things a person would be able to do if they had (practically) unlimited time.

This is one of the inflection points around working with AI. When your thinking shifts from "AI does what the person used to do" to "AI does something different that leads to the same outcome as before, but with all that cool stuff I'd love to have the time for", then the equation changes. For example, I will never write another app that doesn't have 100% test coverage again. I stress- and soak-test everything these days. That's changing how I write code - I need to write things in ways that have deterministic harnesses for things mutable state outside of my immediate control (like random numbers or datetimes or state loaded from a save) so that I can task AI with building a fuzzer that does tens of thousands of random tests on every push to see if anything broke. I couldn't do that in the past because it was always effort that a client wouldn't get enough value from to pay me to do it. I can do those things now. It's ace.

(Everywhere I put 'I' you can replace with 'AI of the Week')

bbmatryoshka a day ago | parent | prev | next [-]

2 hours... of agent work, usually for the same output the average worker has to use much more time

knollimar a day ago | parent [-]

Throw understanding images in there and it becomes less true in my experience. Sure they shit out code and search text well, but one infographic and they circle trying to rasterize it.

Gareth321 a day ago | parent | prev [-]

My codex tasks regularly cross 8 hours, and I'm only using Sol High. It's not unusual for tasks to span much longer. It just requires instructions to continue working until the spec is complete.

Anthropic and OpenAI are currently obsessed with getting humans out of the training and improvement loops. It's going to happen very soon and when it does I think we see staggering improvements in a very short space of time. Basically, the Singularity.

munksbeer 21 hours ago | parent | next [-]

Are you breaking up your tasks and spawning new sessions for each, or are you just yoloing and letting it auto compact when it blows the context many times on such a long running task?

Gareth321 21 hours ago | parent [-]

Yolo. I keep a large project markdown document + incident/logs/feature documents which its instructed to review and update as necessary. It still occasionally misses stuff but it's surprisingly effective. Disclaimer: this is for hobby software. For work I'm more cautious - usually.

bwfan123 a day ago | parent | prev [-]

> It just requires instructions to continue working until the spec is complete

Try putting an LLM agent in a deterministic workflow without humans in the loop. My experience with this is not encouraging. Getting it to work requires sprinkling some context magic and hoping and praying the LLM does the right thing. More astrology or religion and less science. Great for use cases with humans-in-the-loop, but less than impressive when you need determinism and reliable operation.

Gareth321 21 hours ago | parent [-]

> Try putting an LLM agent in a deterministic workflow without humans in the loop.

I would not use a sewing machine to repair my deck :) LLMs are non-deterministic, by design.

onion2k 21 hours ago | parent [-]

LLMs are non-deterministic, by design.

They don't have to be though. You can give them a temperature of zero so they pick the highest-probability token every time to give you a deterministic output. I imagine this doesn't work on frontier models because there's a lot going on, but you can definitely do it with a small local model.

pazimzadeh a day ago | parent | prev | next [-]

that's a lot of opinion without much reasoning/support

even if you're right, all fields are progressing at the same time, including biology, and the synergies are going to be significant

embedding-shape a day ago | parent | prev | next [-]

> LLMs gave us nice productivity tools for highly motivated expert knowledge workers, that is all.

I feel like this is incredibly simplistic, and if you changed "LLM" to "computer" or "internet" and went back decades, it's highly likely you'd read the exact same takes in the newspaper back then about those things.

It's also at the same time downplaying just how valuable "nice productivity tools for highly motivated expert knowledge workers" could be, just like computers did for us "highly motivated expert knowledge workers" in the first place. Or computers isn't a "foundational technology" either?

maxnevermind a day ago | parent [-]

> I feel like this is incredibly simplistic, and if you changed "LLM" to "computer" or "internet" and went back decades, it's highly likely you'd read the exact same takes in the newspaper back then about those things.

We are not people who write newspapers we are people who are directly involved in application of the technology, we posses a higher level of insight.

> I feel like this is incredibly simplistic, and if you changed "LLM" to "computer" or "internet" and went back decades, it's highly likely you'd read the exact same takes in the newspaper back then about those things.

Computers, the Internet, Steam engine, Rail roads are much more reliable. If you were a problem solver and entrepreneur you could go and apply those on small scale, get profits, reinvest, set up a flywheel, build a fortune on the top of those, there were a lot of low hanging fruit due to the fact that the foundation was laid out, a break-even period of James Watt steam engine was like 2-3 years I believe until everyone saw it and margins flattened. Can you do that with LLMs? It doesn't look that way.

pixl97 20 hours ago | parent | next [-]

>Computers, the Internet, Steam engine, Rail roads are much more reliable.

Because mostly they are doing a limit number of tasks in a non-generic manner.

Railroads had all kinds of problems at first, from iron rails bending and skewering the riders to engines exploding like high pressure bombs cutting down swaths of people like grain. People still used them with those risks because taking a wagon was worse.

The internet itself is a pretty simple concept. It's "just" packet routing. Now, to make that reliable at scale we've added an enormous amount of complexity to it. If you talk an old TCP stack and dump it on todays internet it you'd find yourself hacked pretty quickly. The internet is a generic packet router, but not a generic anything router.

Computers is where the reliability starts to fall apart. Once you go from specialized purpose to general purpose the size of the problem space explodes. You wouldn't get companies called things like Micro$lop (and these things were from before the AI days).

Then when you get to AI/LLMs and especially when you push into AGI you're creating a do anything now machine. And when you can do anything it takes a lot of training to eliminate the choices that end badly. You are the endpoint of 500 million years of training, for example.

>Can you do that with LLMs? It doesn't look that way.

Can anyone do it now? You're talking about low hanging fruit 200 years ago, but how much of that is left? How much harder are we going to have to fight for the next apple up?

embedding-shape a day ago | parent | prev [-]

> We are not people who write newspapers we are people who are directly involved in application of the technology, we posses a higher level of insight.

Cool, you're still a person who share their thoughts, just like the people who had similar reactions to when computers and the internet first came around. So many people were railing against the internet with "You'll never make payments over the internet", "it's too slow" and whatever. To the surprise of none of us nerds, the internet eventually ate the whole world and now large swaths of the population can't imagine life without it, for better or worse.

> Computers, the Internet, Steam engine, Rail roads are much more reliable.

You're comparing the reliability of huge ecosystems after they've gone mainstream, after decades, with something that is a couple of years old, hardly a fair comparison. Instead, compare where LLMs are today with the first years of computers, the internet and all that, and maybe you'll gain a new perspective on where LLMs are? None of those things were reliable initially and in many cases, involved actual human deaths before safety actually caught up.

Even when I got started with using the internet, which isn't that long time ago but around the modem-era, things were not so reliable or simple as they are today.

jstummbillig a day ago | parent | prev | next [-]

The thing that would change it a lot is that people would start trading away their GPUs, because some people would get vastly more economic use out of them.

RajT88 a day ago | parent [-]

This. Look at what happened when the government gave out free ATSC tuners for people with old TV's. eBay filled up with ATSC tuners. lol

lukan a day ago | parent | prev | next [-]

"If we talk about just LLMs"

In general we don't, as the "chatbots" already have way more capability than "just" being able to work with text, which "just" means encoding and processing real world information and lot's of it.

For example I just fixed the gearing of my bicycle by dropping some pictures into claude and it gave me a step by step guide to fix it.

"LLMs gave us nice productivity tools for highly motivated expert knowledge workers, that is all"

And I cannot understand how that can be "all" as it means giving this power to every human, not just the few in the position to hire expert knowledge workers. Steam engines gave us the capability to let machines do the hard labour. Now with better engines and tools, AI will give robots the capability to do allmost anything humans can do. If that is not revolutionary then I don't know what is.

The main reason we humans got to the top of the food chain was because we are expert knowledge workers, able to transform the world around us to our needs.

"but it is impossible to describe the universe and compress it into few terabytes and that is what they are doing atm effectively, so all serious LLMs's issues will still be there in 2040: the lack on continues learning, hallucinations, terrible sample ratio, agent's failures on long horizon tasks, instruction following failures."

So again, just limiting this debate to a simple fixed model - then no, not too useful - but a model that can take notes and at some point reliable clear up it's context - that is a whole different story.

And it has been repeated often enough, but AI does not need to be perfect - it just has to be as or more reliable than us humans. Because yes - AI does make misstakes - but so do humans.

falcor84 a day ago | parent | prev | next [-]

> it is impossible to describe the universe and compress it into few terabytes

Why? I'm not an expert, but from my understanding, there say species that function well in their ecological niches with significantly smaller cognitive capacities. And we already have autonomous cars and drones navigate and function reasonably well. So I don't see any reason to believe it's impossible. And even if terabytes aren't sufficient, there's no real barrier to increasing capacity.

maxnevermind a day ago | parent [-]

> there say species that function well in their ecological niches with significantly smaller cognitive capacities

Yes, but LLMs are not animals, animals learn from experience and LLMs don't.

Btw I meant bigger and bigger amount of data of extracted reasoning chains when you go deeper and generate more and more of them in your attempt to describe the universe, the amount of permutations explodes. And it seems LLMs can't workaround that because they don't build world model inside so they can't deduct it from pre-built world/object model, they must memorize it and look it up later.

pixl97 20 hours ago | parent [-]

>but LLMs are not animals, animals learn from experience and LLMs don't.

I mean they kind of do by distilling said experience and putting it in the next model.

Current LLMs can't do it because building said world model is super expensive, anything that lowers that requirement brings us closer to continuous learning.

Also the models we train these days are typically generalized human text models. Animal models are a bit different because they won't be "word" models and they aren't going to be generalized text models like us humans use. In fact there is a story just today on HN about a user creating some rather simple transformer models to solve a number of the Arc-AGI problems not using (human) words at all. The reason you don't see more of this is most people aren't dumping compute into these kinds of issues, but instead going for the AGI prize.

smusamashah a day ago | parent | prev | next [-]

Computers were made and automated lots of work, even though it was still us operating those computers and making and running those automations. But eventually computers took over and running so much of the world.

Now LLMs are automating the computers themselves. This is a whole new layer that we never had before, not in this amount or with this much availability. It is not "not foundational".

bsenftner a day ago | parent | prev | next [-]

Wow, you really do not understand them. The 'problem' with LLM AI is the weak educations of the people that cannot grasp their subtle multidimensional nature, and how that creates a requirement on the users' behalf to understand the language they use to communicate in a similar manner as numbers are to algebra. LLM AIs are literally literature calculus but that truthful phrase flies right over the majority's heads.

dgellow a day ago | parent | prev | next [-]

Also, way more centralization, with all that logic pushed to a few AI vendors

holoduke a day ago | parent | prev | next [-]

Absolutely untrue. LLM will be the automation force of everything. It will connect and orchestrate our entire society. It will be a much bigger impact that any other revolution in the past. Our lives will be unrecognizable in 10 years from now.

lacedeconstruct a day ago | parent [-]

In 10 years from now we will have another AI winter until a new thing arrives that pushes further.

ajjahs a day ago | parent | prev [-]

[dead]