| ▲ | baxtr 5 hours ago |
| I think it might have accelerated things but on a much more basic level, there seems to be no real moat in synthesizing the world’s knowledge into LLMs. |
|
| ▲ | rpozarickij 4 hours ago | parent | next [-] |
| There's no question that training leading LLMs requires some serious expertise and know-how, but surely already having advanced LLMs/agents must be helping tremendously not only for software engineers but also for those working on LLMs themselves. |
| |
| ▲ | joe_the_user 3 hours ago | parent [-] | | I think one could describe LLM optimization as "hard but not a moat". Years ago, optimizing neural nets was described "graduate student descent" - it's tricky but throw enough conventionally smart people at it and it will happen. It's like tuning a hot rod and finding a reproducible bug in a large code base. It's hard and there are tricks but not absolute hurdles, no problems waiting for a conceptual breakthrough (and at today's scales, are there any problems waiting for an Einstein to solve? That's an open (AI) question). |
|
|
| ▲ | cyanydeez 5 hours ago | parent | prev | next [-] |
| I also think we're seeing the sigmoid approaching. |
| |
| ▲ | ben_w 23 minutes ago | parent | next [-] | | I want that to be true (assuming you mean specifically the second half of the sigmoid) just to give me room to adapt to the changes we've already seen; but I've seen comments saying things are slowing down since around when GPT-4 came out. | |
| ▲ | spwa4 4 hours ago | parent | prev [-] | | What is really going on: all the AI labs are doing panicked model releases (and panicked training of new ones) because Qwen4 is rumored to come out end of October and is rumored to be very nice. Question is: is it another "Deepseek-moment" nice? Or just nice? Btw: with Qwen4 I mean the next large Qwen model that is based on the Qwen4 architecture (Qwen 3.8 flash next was "almost" based on the new arch but obviously was a small model) | | |
| ▲ | ejeje12a 4 hours ago | parent | next [-] | | It doesn’t matter. What matters more is if firm’s start using a bundle of American and Chinese models and when they find their feet - how large is the market for frontier? Frontier has to displace labour one for one at some point or it’s over. | |
| ▲ | cyanydeez an hour ago | parent | prev | next [-] | | I'm pretty sure the point of Qwen3.8-Flash-Next was to get the open source engines to integrate the qwen4 architecture. The fact that it basically broke open the local model supremacy was just a nice side effect. I'm running: https://github.com/peonist-ai/halogen-server with a quant4, PLE offloaded, and it's resident VRAM is 36GB at 265k context. Shave 10 more GB off and the TAM openai and anthropic are targeting is a lost cause. Local models are what 90% of people will need. If the world governments can get a handle on the memory cartel, then there's no more moat for most normal humans. | |
| ▲ | verdverm 4 hours ago | parent | prev | next [-] | | I wondered if there would be a Qwen 4 or we would go straight to 5, re: tetraphobia, but perhaps it's more like an uno reverse card in this case https://en.wikipedia.org/wiki/Tetraphobia | | |
| ▲ | satvikpendem 3 hours ago | parent [-] | | DeepSeek being Chinese also has 4 so I don't think it's a big deal for model makers. | | |
| ▲ | verdverm 3 hours ago | parent [-] | | GLM moved through their 4-series without consequence I wonder if the next DS models will also graduate to 5.x, I think I saw they are training up a 10T model, and just raised $12B too |
|
| |
| ▲ | christkv 4 hours ago | parent | prev [-] | | Awesome 3.8 next runs great on my Framework Desktop so I'm loving more local models. | | |
| ▲ | hypercube33 3 hours ago | parent [-] | | I have similar hardware - what specific version of the 3.8 Next model are you running and how many tokens/s are you seeing? 3.6 35B A3B gets about 68t/s for me so I've been sticking with that model for the mean time. | | |
|
|
|
|
| ▲ | scotty79 5 hours ago | parent | prev [-] |
| I think the moat is going to be compute. So far compute needed to push the frontier is still extremely cheap so the capital can afford to spread its bets. But when further improvement is going to cost in trillions, capital will have to pick a winner and bet only on him. It won't be a matter of finding the best bet, it will be a matter of survival. This will cause the picked winner to get massively ahead with sheer compute alone used both for training and inference dedicated to recursive self improvement. |
| |
| ▲ | flir 4 hours ago | parent | next [-] | | I guess all predictions age like milk, but here's one: There's a law of diminishing returns at play here, and doubling the energy cost of training to wring 2% more performance out of the technology isn't going to be very useful, because most of the problems it is capable of solving will be solvable with the previous-gen 98%-as-good model. ("there's a law of diminishing returns at play here" is an article of faith. But then, so is the belief that these models will keep getting better). | | |
| ▲ | londons_explore 4 hours ago | parent [-] | | As soon as you can demonstrate decent financial returns (ie. the AI can run a company better than humans can), suddenly it makes sense to put a lot more $$$ in even if returns are diminishing - since whoever runs companies the best gets control of a big chunk of the world economy. | | |
| ▲ | ejeje12a 4 hours ago | parent | next [-] | | “ ie. the AI can run a company better than humans can” lol
You can always tell who has never ran a business before with comments like this | | |
| ▲ | satvikpendem 3 hours ago | parent [-] | | I understood their point, if we truly get AGI then no reason to think AI would be worse than a human. | | |
| ▲ | asplake an hour ago | parent | next [-] | | AGI with people skills? | |
| ▲ | ejeje12a 3 hours ago | parent | prev [-] | | [flagged] | | |
| ▲ | cheevly 2 hours ago | parent | next [-] | | You are literally an AI spambot account, that's the hilarious irony here. | |
| ▲ | satvikpendem 3 hours ago | parent | prev [-] | | Well yes that's the worry isn't it? That people will soon all be unemployed? Just because it sounds farfetched doesn't mean you should stick your head in the sand, seeing the pace of development these days, that is "reasoning properly" and it just seems you are trying to block out what seems inconvenient to hear. | | |
|
|
| |
| ▲ | flir 3 hours ago | parent | prev [-] | | If that works... why not just jump straight to a planned economy run by LLM? Skip the whole messy "free market" thing altogether? (I don't think it will work). | | |
| ▲ | allendoerfer 3 hours ago | parent [-] | | What do you have against multi-agent reinforcement learning systems and why do you think they are not AI? | | |
| ▲ | flir 3 hours ago | parent [-] | | londons_explore is arguing for a winner-takes-all scenario, with an early advantage locking everyone and everything else out. |
|
|
|
| |
| ▲ | pennomi 5 hours ago | parent | prev | next [-] | | Surely there is a point where algorithmic improvements will be more cost effective than buying more hardware. | | |
| ▲ | Zigurd 4 hours ago | parent | next [-] | | If you're actually applying LLMs, all of the things around the LLM that adapt it to coding, for example, that enable it to use existing validation tools for code, and enable it to diagnose and fix tool chain issues that aren't directly coding problems, are what makes the difference between a model that that scores a little higher on a coding benchmark and a model that's useful in a particular code base on a particular platform. Are there any use cases that have enabled one customer of a frontier LLM to outperform a competitor using a different frontier LLM? Or is this why we are seeing confected points of comparison like solving challenge problems in mathematics? | |
| ▲ | 4 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | scotty79 5 hours ago | parent | prev [-] | | I'm afraid it might be the other way around. RSI might pick all of the low hanging fruit soon. There must be a physical limit of how much intelligence you can squeeze out of some amount of parameters and compute. There are going to still be worthwhile improvements but they are going to be more like not how to make transformers 10x cheaper but how to make next training run cost 9 trillions instead of 10 with a very particular optimization designed at the cost of hundreds of millions for this one specific run. | | |
| ▲ | LPisGood 4 hours ago | parent | next [-] | | I think we’re no where near a physical information theoretic limit. | | |
| ▲ | thfuran 3 hours ago | parent [-] | | The hardware also is, so there ought to be a whole lot more room for improvement. |
| |
| ▲ | bee_rider 3 hours ago | parent | prev | next [-] | | It is just occurring to me that “RSI” expands to recursive self improvement. Thought people were talking about repetitive stress injuries; either in regards to programmers writing too much code/not having to write code anymore, or the frontier AI companies and their tendency to applaud themselves. | |
| ▲ | ejeje12a 4 hours ago | parent | prev [-] | | [flagged] |
|
| |
| ▲ | jayd16 5 hours ago | parent | prev | next [-] | | But the old models still exist at trivial marginal cost. The frontier models would need to dominate every price point to really take all and so far they haven't been. | | |
| ▲ | ambicapter 4 hours ago | parent [-] | | Don't worry, AI boosters will be in here soon denigrating anyone that uses anything but the latest and greatest models as irrelevant. |
| |
| ▲ | bushbaba 5 hours ago | parent | prev [-] | | Not just compute but energy. Most of Europe has no access to the cost effective power generation needed | | |
| ▲ | Danox 4 hours ago | parent | next [-] | | Build Nuclear, Build Thorium the Chinese are building whatever they can. They’re not locked in by special interest. Is that because they have lots of engineers on the job in government? | |
| ▲ | jsw97 4 hours ago | parent | prev | next [-] | | Training location is flexible. Iceland? | |
| ▲ | pyrale 4 hours ago | parent | prev | next [-] | | Europe has lots of zero-cost windows for electricity, and areas with cheap prices. The real issue is access to oil and gas. | | |
| ▲ | randomNumber7 3 hours ago | parent [-] | | Maybe for training big models one can wait for times when the wind is blowing. | | |
| ▲ | flir 3 hours ago | parent | next [-] | | There are worse ideas. I could imagine a belt of data centres around the equator, that hand off their computational loads as the sun sets. Good scifi-esque premise. | |
| ▲ | kergonath 2 hours ago | parent | prev [-] | | That does not really sound practical to let datacenters costing billions idle half the time. |
|
| |
| ▲ | oakesm9 4 hours ago | parent | prev [-] | | France is actually pretty cheap in Europe. About 15% more than average USA electric prices (but I know that varies a lot across the states so still likely much more than the cheaper areas) |
|
|