| ▲ | abixb 3 days ago | |||||||||||||||||||||||||||||||
I'd wager on the second scenario. Anyone who's been paying attention to the industry knows that most of the 'gains' have come from test-time compute and architecting harnesses in novel ways. In my estimation, capability increases from "pre-training" alone died early last year, and we're now probably seeing test-time and other benchmark hacks approaching their limit as well. If you zoomed back to late-2024, people in the industry were predicting how we'd have AGI by now and the economy would've already 'taken off' with massive productivity growth and ushering in of great prosperity ('deflationary spiral'). Where is it? Where is the productivity growth? Where is the deflationary spiral? To be fair, models have gotten better in jagged ways, but reliability is far from usable, especially in long duration tasks, and there has been no effort by the AI companies to address the human brain's bandwidth bottleneck -- they hit the gas like there's no tomorrow and we have enormously capable but jaggedly intelligent multi-modal models with agentic capabilities that are only as effective as the human using it. This whole thing has become a giant mess. | ||||||||||||||||||||||||||||||||
| ▲ | huurtehoog 2 days ago | parent | next [-] | |||||||||||||||||||||||||||||||
I have been doing some research on the 'productivity paradox'. Fascinating stuff. Turns out productivity growth stalled after the 1960s and has been very low ever since. No one can properly explain why. One thing stands out: investment in compute drives economic growth, but crucially doesn't demonstrably increase productivity either of labor or capital, or total factor productivity. What I suspect is happening is: computers and software drive wealth concentration. 60 years on from the first general commercial computer, IBM 360, the entire industry has been driven by redistributing wealth towards an ever diminishing number of public companies. With that perspective, what is happening with LLMs seems to fall right into the 6 decade pattern. I'm still in the middle of reading the 2 dozen papers or so I found about it so far but it has been fascinating. | ||||||||||||||||||||||||||||||||
| ▲ | pu_pe 2 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
How do you explain the fact that Qwen3.8 27B performs vastly better than any open model from even one year ago, if using the same test-time compute and harness? | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||
| ▲ | BobbyTables2 3 days ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||
I even wonder if the frontier AI models are really as capable as they claim or if the companies behind them have just special cases all the “hard” questions. For example, the earlier generative LLMs couldn’t correctly answer ‘how many r’s in “strawberry”?’ due to the underlying nature of the tokens. If they get it correct today, how do they do it? It feels like we’re being deceived by the Wizard of Oz… | ||||||||||||||||||||||||||||||||
| ▲ | NichoPaolucci 2 days ago | parent | prev [-] | |||||||||||||||||||||||||||||||
I'm with this. My company was absolute chaos at the end of last year. AI this, AI that. On Productivity: There ARE improvements. But, these improvements were also essentially people stopping their normal workload, giving that to someone else, and focusing on building an AI tool. Our sales are way down because the VP of sales is now a tech bro. On Model capabilities: I think they could substantially improve and still have a similar level of usefulness for my company. I'm not working on Navier-Stokes. I'm building software for a business. The better models ARE more accurate, handle more of the workload than they could a year ago, and are still incredibly useful tools. But there's only so much I can gain from letting a model work for longer periods of time and doing X+n reviews with X+n subagents. At the end of the day, I need to maintain my personal understanding of the BUSINESS use cases + decisions so that I can make judgements that AI would never be able to make. And, as SOON as an open model has similar capability to the current SOTA we're probably going to cut ties with our subscriptions. I guess it would kinda be like driving an F1 car to the grocery store. I think I'd rather have the Toyota Camry of models. | ||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||