Remix.run Logo
xutopia a day ago

I can't help but feel that every 1% improvement over a benchmark from new releases of flagship LLMs isn't really that impressive anymore. To some extent it feels like LLM improvements is slowing down significantly despite increasing resources being spent on them.

The people promising AGI just 6 months ago are now more and more quiet about the possibility and rightfully so. LLMs alone are most likely not the path to AGI.

Dlemlo a day ago | parent | next [-]

My personal experience diaagrees.

The Opus moment in November felt very different but despite that, a handful things did not woork well with Opus in November and I tried them last week and they now just work.

In parallel GPT-6 and Fable 5.1 show significant skill improvements in 3D modeling and Astra shows significant progress in computer use.

Do we live in a parallel universe? I mean it, this week was crazy from a progress point of view.

ChickeNES a day ago | parent [-]

> Do we live in a parallel universe? I mean it, this week was crazy from a progress point of view.

This is how I feel too, wandering into some of these threads. Sure the damn things aren't perfect yet, but people act as if they are incapable of doing anything, or no more worth the hype than Bootstrap?? (the latter comes from this very thread!)

Dlemlo a day ago | parent [-]

When Anthropic gave out some usage, I was able to do a lot more and it felt differently than when you always have to think about limits.

I now spend $90 this month and do even more now (privat). It might really be a big bias of these people either not spending any time on it or just doing very small experiments and immediadly being put of by some random issue.

ChickeNES a day ago | parent [-]

Oh I've been maxing out my twin Pro/Max $200 accounts for the past 12 months or so :P

ChickeNES a day ago | parent | prev [-]

No, the people promising AGI 6 months ago are now saying it's here. (and I agree)

nozzlegear 17 hours ago | parent [-]

Is the AGI in the room with us?